Systems and methods for audio processing - Patents.com
The method employs non-autoregressive neural networks to enhance speech recognition and language understanding, achieving significant computational efficiency and accuracy improvements in speech processing systems.
Patent Information
- Application Number
- JP2023208671
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-01-09
- Filing Date
- 2023-12-11
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2043-12-11
AI Technical Summary
Existing speech recognition systems face challenges in efficiently combining speech recognition and speech language understanding, particularly in handling low-reliability tokens and requiring high computational resources.
A computer-implemented method using a non-autoregressive encoder neural network and a bidirectional decoder neural network to process audio signals, generating embeddings and initial transcripts, and iteratively improving token predictions to enhance speech recognition and language understanding.
This approach reduces the real-time factor (RTF) by up to 6 times compared to autoregressive baselines, improves the accuracy of speech recognition and language understanding, and is suitable for resource-constrained devices.
Smart Images

Figure 0007673166000005 
Figure 0007673166000006 
Figure 0007673166000007
Abstract
Description
[Technical field]
[0001] FIELD OF THE DISCLOSURE The embodiments described herein relate to systems and methods for speech processing, and in particular for speech recognition and spoken language understanding. [Background technology]
[0002] Speech recognition methods and systems receive speech audio and recognize the content of such speech audio, e.g., textual content of such speech audio. Speech recognition systems, including hybrid systems, may include an acoustic model (AM), a pronunciation dictionary, and a language model (LM) to determine the content of the speech audio, e.g., to decode the speech. Early hybrid systems utilized Hidden Markov Models (HMMs) or similar statistical methods for the acoustic model and / or the language model. Later hybrid systems utilize neural networks for at least one of the acoustic model and / or the language model. These systems are sometimes referred to as deep speech recognition systems.
[0003] Speech recognition systems with end-to-end architectures have also been introduced. In these systems, a single neural network is used, which can be considered to have an implicit integration of an acoustic model, a pronunciation lexicon, and a language model. The single neural network can be a recurrent neural network. More recently, Transformer models have been used in speech recognition systems. Transformer models may perform speech recognition using a self-attention mechanism where dependencies are captured regardless of their distance. Transformer models may utilize an encoder-decoder framework.
[0004] Systems and methods, by way of non-limiting examples, will now be described with reference to the accompanying drawings, in which: [Brief description of the drawings]
[0005] [Figure 1A] FIG. 1A is a diagram illustrating a voice assistant system according to an exemplary embodiment. [Figure 1B] FIG. 1B is a diagram illustrating a speech transcription system according to an exemplary embodiment. [Figure 1C] FIG. 1C is a flow diagram of a method for performing a voice assistant according to an exemplary embodiment. [Figure 1D] FIG. 1D is a flow diagram of a method for performing audio transcription according to an exemplary embodiment. [Diagram 2] FIG. 2 is a flow diagram of a method for audio processing in accordance with an exemplary embodiment. [Diagram 3] FIG. 3 is a diagram illustrating an example audio processing output. [Figure 4] FIG. 4 is a diagram illustrating an iterative refinement process for generating a transcript. [Diagram 5] FIG. 5 is a schematic diagram of a voice processing system in accordance with an exemplary embodiment. [Figure 6] FIG. 6 is a schematic diagram of an exemplary decoder block. [Figure 7] FIG. 7 is a schematic diagram of an example encoder block. [Figure 8] FIG. 8 is a schematic diagram of an encoder neural network according to an example embodiment. [Figure 9] FIG. 9 is a schematic diagram of the processing in the hidden layer of the encoder neural network. [Figure 10] FIG. 10 is a flow diagram of a method for training a speech processing system in accordance with an exemplary embodiment. [Figure 11] FIG. 11 is a flow diagram of the processing performed in the hidden layers of the encoder neural network during training. [Figure 12] FIG. 12 is a schematic diagram of hardware for implementing the method and system according to an exemplary embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0006] In the drawings, like reference numbers refer to like elements.
[0007] According to one aspect, a computer-implemented method for speech processing is provided. The method comprises receiving an audio signal capturing speech by a user. The method further comprises processing the audio signal by an encoder comprising a non-autoregressive encoder neural network to generate an embedding of the audio signal and an initial transcription of the speech, the initial transcription comprising one or more tokens, each of the one or more tokens being associated with a first confidence score of being a correct token for its position in the initial transcript. The method further comprises modifying the initial transcript to generate a masked token sequence, the modifying the initial transcript comprising masking one or more tokens in the initial transcript having a first confidence score below a threshold confidence level. The method further comprises processing the embedding of the audio signal and the masked token sequence by an output decoder comprising a non-autoregressive bidirectional decoder neural network to generate a speech processing output, the speech processing output comprising a predicted token for each of a plurality of masked tokens in the masked token sequence generating an output transcript of the speech, a label indicative of a classification of the utterance, and one or more labels indicative of a plurality of parameter types associated with a plurality of respective first words in the output transcript, the plurality of parameter types being associated with a classification of the speech.
[0008] The disclosed method jointly performs speech recognition and spoken language understanding on an input audio signal (or acoustic features of an audio signal). In particular, the predicted tokens and output transcript elements of the speech processing output relate to speech recognition, while the label indicating the classification of the utterance and one or more labels indicating parameter types relate to speech language understanding. For example, the label indicating the classification of the utterance may correspond to an action to be taken in response to the utterance by an appropriate device or system. This type of speech processing output is sometimes known as "intent classification". With respect to the parameter type labels, the action may require determining certain parameter values to fully form and execute the action. Words (or sub-word units) in the output transcript may be labeled with corresponding parameter types to fill the required parameters. This type of speech processing output is sometimes known as "slot filling". The method may then further comprise causing the device or system to perform the action, for example, by sending an appropriate command.
[0009] The method uses a non-autoregressive neural network. Compared to an autoregressive neural network that generates outputs over multiple time steps, typically with previously generated outputs fed back as inputs to generate one element at a time, a non-autoregressive neural network generates all of its outputs simultaneously in one time step. Non-autoregressive approaches typically generate their outputs faster than autoregressive approaches and may be used in computing systems with more limited computing resources, such as mobile devices or embedded systems, or when low latency is desired. For example, the inference speed of a spoken language understanding system may be measured using the real-time factor (RTF), which is the ratio of execution time to the length of the input utterance. Using a non-autoregressive system according to the method described herein, a six-fold reduction in RTF can be achieved compared to a comparable autoregressive baseline.
[0010] The decoder neural network is bidirectional. That is, when processing a particular element of the input sequence, the bidirectional decoder neural network processes the particular element based on both the elements preceding the particular element in the input sequence and the elements following the particular element in the sequence. In comparison, a unidirectional decoder neural network typically only considers the elements preceding the particular element. In this way, the decoder neural network can generate an output by considering the entirety of the input. This is particularly advantageous because the tasks performed by the decoder, namely predicting masked tokens and generating labels for spoken language understanding, benefit from being able to condition on the entirety of the input to the decoder. The encoder neural network may also be bidirectional.
[0011] Furthermore, by splitting the tasks performed by the encoder and decoder, joint implementation of ASR and Spoken Language Understanding (SLU) is further facilitated. In particular, the encoder is used to provide an initial ASR transcript, and any low confidence tokens in the transcript are masked. The decoder performs the task of predicting the low confidence masked tokens from the available high confidence tokens, and the SLU tasks of intent classification and slot filling, which are more directly related tasks than pure ASR and SLU. Thus, the decoder can operate more effectively since it is performing more related tasks, and can learn more effective joint features to perform the tasks.
[0012] The embeddings of the audio signal generated by the encoder neural network provide a latent representation of the audio signal that contains important information for use in generating an initial transcription and for the decoder in generating the speech processing output. The embeddings of the audio signal may be vector or tensor representations of the audio signal.
[0013] The tokens may correspond to words or sub-word units, such as word pieces, phonemes, characters, or other suitable units as deemed appropriate by one of skill in the art. For example, the set or vocabulary of tokens may be generated by applying byte pair encoding (BPE) to a suitable corpus, such as a training data set.
[0014] An initial transcription of the utterance may be generated based on a Connectionist Temporal Classification (CTC) algorithm. A greedy search may be used to independently determine the most likely token at each position in the sequence. Alternatively, a beam search may be used.
[0015] The decoder neural network of the output decoder may comprise a first output head for generating predicted tokens for each of the masked tokens in the masked token sequence, and a second output head for generating one or more labels indicative of a first associated parameter type for each word in the output transcript, the first output head and the second output head being different output heads. In this way, one portion of the decoder neural network may be dedicated to generating predictions for the masked tokens, and another portion of the decoder neural network may be dedicated to generating speech level labels and parameter type labels. Furthermore, two separate output sequences may be generated, one from each output head. This may be advantageous as the length of each output sequence is reduced, since traditionally encoder / decoder architectures have difficulty modeling long sequences and long-term dependencies.
[0016] Producing an output transcript of the utterance may comprise iteratively refining token predictions for the masked tokens in the masked token sequence. The iterative refinement may comprise generating a predicted token for each of a plurality of masked tokens in the masked token sequence, where each predicted token is associated with a first confidence score of being a correct token, and for each predicted token, replacing the corresponding masked token in the masked token sequence with the predicted token when the first confidence score is above a threshold confidence level, and retaining the corresponding masked token in the masked token sequence when the first confidence score is below the threshold confidence level. That is, in subsequent iterations, the new masked token sequence may again be processed by the decoder to generate a set of predictions for the remaining masked tokens until there are no masked tokens remaining or until a fixed number of iterations have been performed. The iterative refinement allows predictions to be conditioned on higher confidence tokens coming first. As higher confidence tokens are gradually filled in, predictions become easier to make for the positions of previously lower confidence tokens.
[0017] The processing by the encoder includes generating an intermediate layer embedding of the audio signal and an intermediate layer transcription of the speech in one or more intermediate layers of the encoder neural network, where the intermediate layer transcription comprises one or more tokens, each of the one or more tokens being associated with a second confidence score of being a correct token for its position within the intermediate layer transcript, and modifying the intermediate layer transcript to generate an intermediate layer masked token sequence, where modifying the intermediate layer transcript includes modifying one or more tokens in the intermediate layer transcript that have a second confidence score below the intermediate layer threshold confidence level. processing the intermediate layer embedding of the audio signal and the intermediate layer masked token sequence by an intermediate layer decoder comprising a non-autoregressive bidirectional decoder neural network to generate an intermediate layer decoded output; combining the intermediate layer embedding of the audio signal and the intermediate layer decoded output to provide input for one or more subsequent neural network layers of the encoder neural network following a respective intermediate layer of the encoder neural network; and processing the input by the one or more subsequent neural network layers to generate an embedding of the audio signal and an initial transcription of the speech.
[0018] That is, the encoder neural network may comprise an intermediate layer decoder that may be configured to perform the same task as the output decoder (may have the same architecture and may share parameters). The output of the intermediate layer decoder is provided to a subsequent layer of the encoder neural network to process and generate an embedding of the audio signal. In this way, the encoder can generate a representation of the audio signal that is more useful for the processing performed by the output decoder. Furthermore, certain algorithms for generating a transcript, such as the CTC algorithm, have a strong conditional independence assumption. That is, each token in the transcript is assumed to be conditionally independent. By using an intermediate layer decoder and processing the intermediate layer decoded output in a subsequent layer of the encoder neural network, the conditional independence assumption can be relaxed when generating an initial transcript.
[0019] The encoder neural network may comprise multiple hidden layer decoders at different hidden layers of the encoder neural network.
[0020] The intermediate layer decoding output may comprise a masked token probability distribution comprising a probability distribution over a plurality of candidate tokens for predicting a token for each of a plurality of masked tokens in the intermediate layer masked token sequence, an utterance classification probability distribution comprising a probability distribution over a plurality of labels indicating a classification of the utterance, and a parameter type probability distribution comprising a probability distribution over a plurality of labels indicating a parameter type for a plurality of respective second words in the intermediate layer transcript.
[0021] Combining the intermediate-layer embedding of the audio signal with the intermediate-layer decoded output may comprise combining the masked token probability distributions, the speech classification probability distributions, and the parameter type probability distributions, performing a linear projection of the multiple combined probability distributions such that the dimensionality matches the intermediate-layer embedding of the audio signal, and summing the linearly projected combined probability distributions with the intermediate-layer embedding of the audio signal. The masked token probability distributions, the speech classification probability distributions, and the parameter type probability distributions may be combined based on a concatenation or other suitable combining operation.
[0022] The audio signal may comprise a plurality of audio segments. Generating an intermediate layer transcription of the speech may comprise generating a probability distribution over a plurality of candidate tokens for transcribing each audio segment and generating the intermediate layer transcription based on the combined probability distribution for each audio segment. Combining the intermediate layer embedding of the audio signal and the intermediate layer decoded output may comprise determining an audio segment location corresponding to each token in the intermediate layer transcription, mapping a masked token probability distribution and a parameter type probability distribution to a plurality of audio segment locations based on the corresponding token in the intermediate layer transcription, and associating a speech classification probability distribution with each of the plurality of audio segment locations in the mapping. Combining the masked token probability distribution, the speech classification probability distribution, and the parameter type probability distribution may comprise combining the probability distributions over a plurality of candidate tokens, the masked token probability distribution, the speech classification probability distribution, and the parameter type probability distribution with a plurality of corresponding audio segment locations according to the mapping.
[0023] The method may further comprise prepending a prefix token to the masked token sequence prior to processing by the output decoder to generate a label indicative of the utterance classification. If the encoder neural network comprises an intermediate layer decoder, the method may further comprise prepending a prefix token to the intermediate layer masked token sequence prior to processing by the intermediate layer decoder to generate the utterance classification probability distribution. The prefix token prepended to the masked token sequence may facilitate generation of a label indicative of the utterance classification. The decoder may use the first element position to convey speech level information, and the first element in the output sequence may correspond to the speech level label. The prefix token may be " <cls>" can be shown as:
[0024] The decoder neural network of the output decoder and / or the intermediate layer decoder may be based on a Transformer architecture. In this regard, the decoder neural network(s) may comprise multiple stacked Transformer decoder blocks.
[0025] The encoder neural network may be based on a conformer architecture and may comprise multiple stacked conformer blocks.
[0026] The encoder may comprise an audio processing front end configured to extract acoustic features from the audio signal. The audio processing front end may comprise a convolutional neural network. Alternatively, the audio signal may be pre-processed externally and the received audio signal may be provided with acoustic features.
[0027] According to another aspect, a computer-implemented method for training a speech processing system is provided. The speech processing system includes an encoder comprising a non-autoregressive encoder neural network having a plurality of trainable parameters, and an output decoder comprising a non-autoregressive bidirectional decoder neural network having a plurality of trainable parameters. The method includes receiving an audio signal capturing speech by a user to train the speech processing system. The method further includes processing the audio signal by the encoder to generate an embedding of the audio signal and a conditional probability distribution over a plurality of candidate transcriptions of the utterance given the input audio signal. The method further includes receiving a ground-truth transcription of the utterance, the ground-truth transcription comprising a sequence of one or more tokens. The method further includes calculating an encoder loss value based on the conditional probability distribution over the plurality of candidate transcriptions of the utterance and the ground-truth transcription of the utterance. The method further includes modifying the ground-truth transcript to generate a masked token sequence, and modifying the ground-truth transcript comprises masking one or more tokens in the ground-truth transcript. The method further comprises processing, by an output decoder, the embedding of the audio signal and the masked token sequence to generate a speech processing output, the speech processing output comprising a probability distribution over a plurality of candidate tokens for predicting a token for each of a plurality of masked tokens in the masked token sequence, a probability distribution over a plurality of labels indicative of a speech classification, and a probability distribution over a plurality of labels indicative of a parameter type associated with each word in the output transcript, the plurality of parameter types being associated with a speech classification. The method further comprises receiving a plurality of truth labels indicative of the speech classification and the plurality of parameter types, respectively. The method further comprises calculating a decoder loss value based on the speech processing output, the truth transcript, and the plurality of truth labels.The method further comprises adjusting values of trainable parameters of a decoder neural network and an encoder neural network of the output decoder based on the decoder loss value. The method further comprises adjusting values of trainable parameters of the encoder neural network based on the encoder loss value.
[0028] In the training method, the ground truth transcript is used to generate a masked token sequence that is provided as input to the decoder to avoid the decoder learning erroneous relationships due to errors in the initial transcript generated by the encoder as expected during the first stage of training. The decoder is still provided with the embedding of the audio signal from the encoder as input.
[0029] Adjusting the values of the trainable parameters of the encoder / decoder neural network may be based on backpropagation and stochastic gradient descent. The training process may be repeated for a certain number of iterations and / or until a stopping criterion is reached.
[0030] The processing by the encoder may comprise generating, in one or more hidden layers of an encoder neural network, a hidden layer embedding of the audio signal and a hidden layer conditional probability distribution over a plurality of candidate transcriptions of the utterance given an output of a layer preceding the hidden layer. The method may further comprise calculating a hidden layer loss value based on the hidden layer conditional probability distribution and the true value transcription, the encoder loss value being further based on the hidden layer loss value. The method may further comprise generating a hidden layer transcription of the utterance based on the hidden layer conditional probability distribution, the hidden layer transcription comprising a sequence of one or more tokens. The method may further comprise modifying the hidden layer transcription to generate a hidden layer masked token sequence, the modifying the hidden layer transcription comprising masking one or more tokens in the hidden layer transcript having a confidence score below a hidden layer threshold confidence level. The method may further comprise processing the hidden layer embedding of the audio signal and the hidden layer masked token sequence by a hidden layer decoder comprising a non-autoregressive bidirectional decoder neural network to generate a hidden layer decoded output. The method may further comprise combining the intermediate layer embedding of the audio signal and the intermediate layer decoded output to provide input for one or more subsequent neural network layers of the encoder neural network following a respective intermediate layer of the encoder neural network. The method may further comprise processing the input, by the one or more subsequent neural network layers, to generate an embedding of the audio signal and a conditional probability distribution over multiple candidate transcriptions of the utterance given the input audio signal.
[0031] If the encoder neural network includes an intermediate layer decoder, the intermediate layer loss value is not based directly on any of the intermediate layer decoded outputs, which take into account processing by the rest of the encoder neural network to generate the encoder transcript (the initial transcript) and any further intermediate intermediate layer transcripts, and thus are only indirectly incorporated into the encoder loss value.
[0032] The encoder and decoder loss values may indicate errors caused by the system. The encoder loss value and / or the intermediate layer loss value may be calculated based on the CTC loss function. The decoder loss value may be based on the negative log-likelihood of the true masked token obtained from a probability distribution over a plurality of candidate tokens for predicting a token for each of a plurality of masked tokens in the masked token sequence. More precisely, the probability distribution is over all potential tokens in the vocabulary, conditional on the plurality of observed tokens and embeddings input to the decoder. The decoder loss value may be based on a classification loss value determined based on the negative log-likelihood of the plurality of true labels obtained from a plurality of probability distributions over a plurality of labels indicating a classification of the utterance and a plurality of labels indicating the parameter type, respectively.
[0033] Any of the above probability distributions may be parameterized by the encoder and / or decoder neural networks. A softmax layer may be used to generate the probability distributions from a particular layer of a neural network.
[0034] The encoder and decoder may be configured to perform the operations described in relation to the first method aspect.
[0035] These methods are computer-implemented methods. Some methods according to the embodiments can be implemented by software, so some embodiments encompass computer code provided to a general-purpose computer on any suitable carrier medium. The carrier medium can comprise any temporary medium, such as any storage medium, such as a floppy disk, a CD-ROM, a magnetic device or a programmable memory device, or any signal, such as an electrical signal, an optical signal, or a microwave signal. The carrier medium can comprise a non-transitory computer-readable storage medium. According to a further aspect, a carrier medium is provided that comprises computer-readable code configured to cause a computer to implement any of the above-mentioned methods.
[0036] According to another aspect, there is provided a system comprising a memory having processor-readable instructions stored thereon and a processor arranged to read and execute instructions in the memory, the processor-readable instructions arranged to control the system to perform operations according to any of the method aspects described above.
[0037] It will be readily understood that the embodiments may be combined and that features described in the context of one embodiment may be combined with other embodiments.
[0038] For illustrative purposes, example contexts in which the subject innovation may be applied are described with reference to Figures 1A-1D, although it should be understood that these are exemplary and that the subject innovation may be applied in any suitable context, e.g., any context in which speech recognition / processing is applicable.
[0039] Voice Assistant System FIG. 1A is a diagram illustrating a voice assistant system 120 according to an example embodiment.
[0040] A user 110 may speak commands 112, 114, 116 to a voice assistant system 120. In response to the user 110 speaking a command 112, 114, 116, the voice assistant system implements the command, which may include outputting an audible response.
[0041] To receive the spoken commands 112, 114, 116, the voice assistant system 120 includes or is connected to a microphone. To output an audible response, the voice assistant system 120 includes or is connected to a speaker. The voice assistant system 120 may include functionality, e.g., software and / or hardware, suitable for recognizing the spoken command, implementing or causing the command to be implemented, and / or outputting a suitable audible response. Alternatively or additionally, the voice assistant system 120 may be connected via a network, e.g., via the Internet and / or a local area network, to one or more other systems, e.g., cloud computing systems and / or local servers, suitable for recognizing the spoken command and causing the command to be implemented. A first part of the functionality may be implemented by the hardware and / or software of the voice assistant system 120, and a second part of the functionality may be implemented by one or more other systems. In some examples, functionality, or a substantial portion thereof, may be provided by one or more other systems, where the one or more other systems are accessible over a network, but may be provided by voice assistant system 120 when they are not accessible due, for example, to disconnection of voice assistant system 120 from the network and / or failure of one or more other systems. In these examples, voice assistant system 120 may still be capable of operating without connection to the one or more other systems, but may be able to take advantage of the greater computational resources and data availability of the one or more other systems, for example, to be able to implement a wider range of commands, to improve the quality of voice recognition, and / or to improve the quality of the audible output.
[0042] For example, in command 112, user 110 asks, "What is X?" This command 112 may be interpreted by voice assistant system 120 as a spoken command to provide a definition of term X. In response to the command, voice assistant system 120 may query a knowledge source, such as, for example, a local database, a remote database, or another type of local or remote index, to obtain a definition of term X. Term X may be any term for which a definition can be obtained. For example, term X may be a dictionary term, such as a noun, verb, or adjective, or may be the name of an entity, such as the name of a person or company. Once a definition is obtained from a knowledge source, the definition may be synthesized into a sentence, such as, for example, a sentence of the form "X is [definition]." The sentence may then be converted into an audible output 122, for example, using text-to-speech functionality of voice assistant system 120, and output using a speaker included in or connected to voice assistant system 120.
[0043] As another example, in command 114, user 110 says, "Turn off the lights." Command 114 may be interpreted by the voice assistant system as a spoken command to turn off one or more lights. Command 114 may be interpreted by voice assistant system 120 in a context-dependent manner. For example, voice assistant system 120 may know the room in which it is located and specifically turn off the lights in that room. In response to the command, voice assistant system 120 may cause one or more lights to turn off, e.g., cause one or more smart light bulbs to no longer emit light. Voice assistant system 120 may cause one or more lights to turn off by interacting directly with the one or more lights, e.g., via a wireless connection, such as a Bluetooth connection, between the voice assistant system and the one or more lights, or by interacting indirectly with the lights, e.g., by sending one or more messages to a smart home hub or a cloud smart home control server to turn off the lights. The voice assistant system 120 may generate an audible response 124 that confirms to the user that the voice assistant system 120 heard and understood the command, such as a voice saying "turn off the lights."
[0044] As a further example, in command 116, user 110 says, "Play some music." Command 116 may be interpreted by the voice assistant system as a spoken command to play music. In response to the command, voice assistant system 120 may access a music source, such as a local music file or a music streaming service, stream music from the music source, and output the streamed music 126 from a speaker included in or connected to voice assistant system 120. The music 126 output by voice assistant system 120 may be personalized for user 110. For example, voice assistant system 120 may recognize user 110, for example, by the nature of the user's 110's voice, or may be statically associated with user 110, and then resume music previously played by user 110 or play a playlist personalized for user 110.
[0045] Audio transcription system FIG. 1B is a diagram illustrating a speech transcription system in accordance with an exemplary embodiment.
[0046] A user 130 may speak to a computer 140. In response to the user speaking, the computer 140 generates a text output 142 that represents the content of the speech 132.
[0047] To receive the voice, computer 140 includes or is connected to a microphone. Computer 140 may include software suitable for recognizing the content of the voice audio and outputting text representative of the content of the voice, e.g., transcribing the content of the voice. Alternatively or additionally, computer 140 may be connected via a network, such as via the Internet and / or a local area network, to one or more other systems suitable for recognizing the content of the voice audio and outputting text representative of the content of the voice. A first portion of the functionality may be implemented by the hardware and / or software of computer 140, and a second portion of the functionality may be implemented by one or more other systems. In some examples, the functionality, or a majority of the functionality, may be provided by one or more other systems, where these one or more other systems are accessible via a network, but when they are not accessible due to, e.g., disconnection of computer 140 from the network and / or failure of one or more other systems, the functionality may be provided by computer 140. In these examples, computer 140 may still be able to operate without connection to one or more other systems, but may be able to take advantage of the greater computational resources and data availability of the one or more other systems, e.g., to improve the quality of the voice transcription.
[0048] The output text 142 may be displayed on a display included in or connected to computer 140. The output text may be input to one or more computer programs executing on computer 140, such as a word processing computer program or a web browser computer program.
[0049] Voice Assistant Method 1C is a flow diagram of a method 150 for performing voice assistance according to an exemplary embodiment. Optional steps are indicated by dashed lines. The exemplary method 150 may be implemented as one or more computer-executable instructions executed by one or more computing devices, such as, for example, hardware 1200 described in connection with FIG. 12. The one or more computing devices may be or include a voice assistant system, such as, for example, voice assistant system 120, and / or may be integrated into a general-purpose computing device, such as a desktop computer, a laptop computer, a smartphone, a smart television, or a game console.
[0050] In step 152, voice audio is received using a microphone, such as a microphone of the voice assistant system or a microphone integrated into or connected to the general-purpose computing device. As the voice audio is received, it may be buffered in a memory, such as a memory of the voice assistant system or the general-purpose computing device.
[0051] In step 154, the content of the speech audio is recognized. The content of the speech audio may be recognized using methods described herein, such as method 200 of FIG. 2. Prior to use of such methods, the audio may be preprocessed as described in further detail below. The recognized content of the speech audio may be text, syntactic content, and / or semantic content. The recognized content may be represented using one or more vectors. Additionally or alternatively, for example after further processing, the recognized content may be represented using one or more tokens. If the recognized content is text, each token and / or vector may represent a character, a phoneme, a morpheme or other morphological unit, a word part, or a word.
[0052] In step 156, a command is implemented based on the content of the voice audio. The command implemented may be, but is not limited to, any of the commands 112, 114, 116 described in connection with FIG. 1A and may be implemented in the manner described. The command implemented may be determined by matching the recognized content with one or more command phrases or patterns. The match may be approximate. For example, for a command 114 to turn off the lights, the command may be matched with a phrase that includes the words "light" and "off", such as "turn off the lights" or "turn off the lights". The command 114 may be matched with phrases that approximately correspond semantically to "turn off the lights", such as "turn off the lights" or "turn off the lights".
[0053] At step 158, an audible response is output based on the content of the voice audio, for example using a speaker included in or connected to the voice assistant system or general-purpose computing device. The audible response may be any of the audible responses 122, 124, 126 described in connection with FIG. 1A and may be generated in the same or similar manner as described. The audible response may be a spoken sentence, word, or phrase, music, or another sound, such as, for example, a sound effect or alarm. The audible response may be based on the content of the voice audio itself and / or may be indirectly based on the content of the voice audio, for example, based on an performed command that is itself based on the content of the voice audio.
[0054] If the audible response is a spoken sentence, phrase, or word, outputting the audible response may include using a text-to-speech function to convert a text, vector, or token representation of the sentence, phrase, or word into spoken audio corresponding to the sentence, phrase, or word. The representation of the sentence or phrase may be synthesized based on the content of the speech audio itself and / or the executed command. For example, if the command is a definition retrieval command of the form "What is X?", the content of the speech audio includes X and the command causes the definition [def] to be retrieved from a knowledge source. A sentence of the form "X is [def]" is synthesized, where X is from the content of the speech audio and [def] is the content retrieved from the knowledge source by executing the command.
[0055] As another example, if the command is a command to cause the smart device to perform a function, such as a turn off the lights command to turn off one or more smart light bulbs, the audible response may be a sound effect indicating that the function has been or is being performed.
[0056] As indicated by the dashed lines in the figure, the step of generating an audible response is optional and may not occur for some commands and / or some implementations. For example, for a command that causes the smart device to perform a function, the function may be performed without an audible response being output. An audible response may not be output because the user has other feedback that the command has been successfully completed, such as a light going off.
[0057] How to transcribe audio 1D is a flow diagram of a method 160 for performing audio transcription according to an exemplary embodiment. The exemplary method 160 may be implemented as one or more computer-executable instructions executed by one or more computing devices, such as, for example, the hardware 1200 described in connection with FIG. 12. The one or more computing devices may be computing devices such as a desktop computer, a laptop computer, a smartphone, a smart television, or a game console.
[0058] At step 162, voice audio is received using a microphone, such as a microphone integrated into or connected to the computing device. As the voice audio is received, it may be buffered in a memory, such as the memory of the computing device.
[0059] In step 164, the content of the speech audio is recognized. The content of the speech audio may be recognized using methods described herein, such as method 200 of FIG. 2. Prior to use of such methods, the audio may be preprocessed as described in more detail below. The recognized content of the speech audio may be textual, syntactic, and / or semantic content. The recognized content may be represented using one or more vectors. Additionally or alternatively, for example after further processing, the recognized content may be represented using one or more tokens. If the recognized content is text, each token and / or vector may represent a character, a phoneme, a morpheme or other morphological unit, a word part, or a word.
[0060] In step 166, text is output based on the content of the speech audio. If the recognized content of the speech audio is text content, the text to be output may be the text content or may be derived from the recognized text content. For example, the text content may be represented using one or more tokens, and the text to be output may be derived by converting the tokens to the characters, phonemes, morphemes or other morphological units, word parts, or words that they represent. If the recognized content of the speech audio is or includes semantic content, an output text having a meaning corresponding to the semantic content may be derived. If the recognized content of the speech audio is or includes syntactic content, an output text having a structure, such as a grammatical structure, corresponding to the syntactic content may be derived.
[0061] The output text may be displayed. The output text may be input into one or more computer programs, such as a word processor or a web browser. Further processing may be performed on the output text. For example, spelling and grammar errors in the output text may be highlighted or corrected. In another example, the output text may be translated, such as using a machine translation system.
[0062] Figure 2 - Method flow diagram 2 is a flow diagram of a speech processing method 200. In particular, the speech processing method combines speech recognition and spoken language understanding. In step 210, an audio signal capturing an utterance by a user is received. The audio signal may be obtained from a sound capture device, such as a microphone, on a user device. The utterance may represent a command or action that the user wants the device to perform.
[0063] In step 220, the audio signal is processed by an encoder to generate an embedding of the audio signal and an initial transcription of the speech. The encoder comprises a non-autoregressive encoder neural network configured to generate an embedding of the audio signal. Compared to an autoregressive neural network that generates outputs over multiple time steps, typically with previously generated outputs fed back as inputs to generate one element at a time, the non-autoregressive neural network generates all of its outputs simultaneously in one time step. Thus, the non-autoregressive encoder neural network generates an embedding of the audio signal, all in one time step. The embedding of the audio signal provides a latent representation of the audio signal that contains important information for use in generating the initial transcription and for the decoder in generating the speech processing output. The embedding of the audio signal may be a vector or tensor representation of the audio signal.
[0064] The encoder may perform pre-processing on the audio signal before processing by the encoder neural network to generate an embedding of the audio signal. For example, the encoder may perform a frequency transformation of the audio signal to generate an acoustic feature vector. Alternatively, the audio signal may be pre-processed before receiving it at step 210. Further details regarding the encoder are detailed below.
[0065] The initial transcription of the utterance comprises one or more tokens. A token may correspond to a word or sub-word unit, such as a word piece, a phoneme, a character, or other suitable unit as deemed appropriate by one of skill in the art. In one example, a set or vocabulary of tokens is generated by applying byte pair encoding (BPE) to a suitable corpus, such as a training data set. For example, the vocabulary may comprise 500 word pieces generated by BPE.
[0066] Each of the one or more tokens in the initial transcript is associated with a confidence score of being the correct token for its position in the transcript. For example, the audio signal may be pre-processed into temporal audio segments, and for each audio segment, the encoder may use an encoder neural network to generate a probability distribution over the token vocabulary that represents the likelihood that an audio segment comprises a voice corresponding to a particular token. Any suitable speech recognition decoding method may be used to determine a transcription of the utterance from the probability distribution. For example, a connectionist time series classification (CTC) method may be used. The confidence score associated with the one or more tokens may be generated based on the probability distribution generated by the encoder and / or the particular decoding method used. Further details regarding the generation of the initial transcript are described below.
[0067] At step 230, the initial transcript is modified to generate a masked token sequence. The modification comprises masking one or more tokens in the initial transcript that have a confidence score below a threshold confidence level. For example, a special " <mask>" tokens may be used to replace tokens having confidence scores below a threshold confidence level. Thus, only tokens in the initial transcript that have a high confidence of being correct are retained, and low confidence tokens are masked out. In one example, the confidence score threshold is 0.999, although it will be appreciated that any other value deemed appropriate by one of ordinary skill in the art may be used. Additionally, a prefix token may be prepended to the masked token sequence to facilitate generation of a label indicating the classification of the utterance. The prefix token may be " <cls>" can be shown as:
[0068] In step 240, the embedding of the audio signal generated by the encoder in step 220 and the masked token sequence generated in step 230 may be processed by an output decoder to generate a speech processing output. The output decoder comprises a non-autoregressive bidirectional decoder neural network. As described above, the non-autoregressive neural network generates its output simultaneously in one time step. When processing a particular element of the input sequence, the bidirectional decoder neural network processes the particular element based on both the elements preceding the particular element in the input sequence and the subsequent elements following the particular element in the sequence. In comparison, a unidirectional decoder neural network typically only considers the elements preceding the particular element. Further details regarding the decoder are provided below.
[0069] The speech processing output comprises a predicted token for each masked token in the masked token sequence to generate an output transcript of the utterance. That is, the output decoder uses the embedding of the audio signal and the high confidence tokens from the initial transcript generated by the encoder to help derive the low confidence tokens to be masked and filtered out. The decoder may be trained using a token prediction task specifically for this purpose. The training of the decoder is described in more detail below. The predicted tokens generated by the decoder may be used to replace the corresponding masked tokens in the masked token sequence to generate an output transcript of the utterance.
[0070] The speech processing output further comprises a label indicating a classification of the utterance. For example, the label may correspond to an action to be taken in response to the utterance by an appropriate device or system. This type of speech processing output is sometimes known as "intent classification." The method may further comprise causing a device or system to perform an action, for example, by sending an appropriate command.
[0071] The speech processing output further comprises one or more labels indicating a parameter type associated with each word in the output transcript. The parameter type is associated with a classification of the utterance. For example, an action may require a particular parameter value to be determined in order to fully form and execute the action. Words (or sub-word units) in the output transcript may be labeled with a corresponding parameter type for filling the required parameters. This type of speech processing output is sometimes known as "slot filling."
[0072] The audio processing output may be provided as two separate columns, the first column comprising the output transcription of the utterance and the second column comprising the utterance level labels and parameter type labels.
[0073] FIG. 3 provides an example output transcript 310, intent labels 320 that classify the utterance, and slot labels 330 that indicate corresponding parameter types for words in the transcribed utterance.
[0074] 2, the method 200 begins with an input audio signal capturing an utterance by a user and provides a speech processing output comprising predicted tokens for each of the masked tokens in the masked token sequence to generate an output transcript of the utterance, a label indicating a classification of the utterance, and one or more labels indicating a parameter type associated with each word in the output transcript. Thus, the method provides end-to-end joint speech recognition and spoken language understanding.
[0075] 4, generating an output transcript of an utterance may comprise iteratively refining token predictions for masked tokens in a token sequence. For example, in FIG. 4, a masked token sequence 410 received by a decoder may include "H", <mask> 」、「 <mask>In the first iteration, the decoder may make predictions 420 for each of the masked tokens in the masked token sequence. For example, the decoder may predict that the first masked token is an "A" with a confidence score of 0.5. The decoder may predict that the second masked token is an "L" with a confidence score of 0.95. The decoder may then, for each predicted token, replace the corresponding masked token in the masked token sequence with the predicted token when the confidence score is above a threshold level. Otherwise, if the confidence score is not above the threshold level, the masked token is kept, i.e., is not replaced. For example, in FIG. 4, the prediction of "A" with a confidence score of 0.5 is below the threshold level of 0.9, so the first masked token is kept and not replaced with the predicted token. However, the prediction of “L” for the second masked token has a confidence score greater than the threshold level, and therefore the prediction “L” is the second “ <mask>" token. Thus, the modified masked token sequence 430 is <mask>", "L", "L", and "O".
[0076] A second iteration may be performed to determine the remaining masked tokens in the masked token sequence by again processing the masked token sequence with the newly filled tokens by the decoder. For example, based on the newly filled tokens, the decoder now makes a prediction 440 that the masked token is "E" with a confidence score of 0.95. This is greater than the confidence threshold level, so the masked token is replaced with the predicted token "E". The new masked token sequence 450 is "H", "E", "L", "L", "O". Since there are no more masked tokens in the masked token sequence, the transcription may be considered complete. In general, iterative refinement may be performed until all masked tokens are replaced or until a fixed number of iterations have been performed. In the latter case, the masked tokens are replaced with their predictions in the last iteration, regardless of the confidence score. In one example, a maximum of 10 iterations are performed, but it will be understood that any suitable maximum number may be selected as deemed appropriate by those skilled in the art.
[0077] Iterative refinement allows predictions to be conditioned on higher confidence tokens coming first: as higher confidence tokens are gradually filled in, predictions become easier to make about the positions of previously lower confidence tokens.
[0078] An exemplary method for generating an initial transcript is now described. The exemplary method uses a CTC-based algorithm. As described above, the audio signal may be pre-processed to provide a plurality of temporal audio segments. The audio segments may be processed by an encoder neural network to provide a probability distribution over a token vocabulary indicating the likelihood that an audio segment contains speech corresponding to a particular token in the vocabulary. The probability distribution over the token vocabulary for each audio segment may be used to determine the tokens for that audio segment. The transcript may be generated using a greedy selection process, i.e., the token with the highest probability in each audio segment may be selected individually. Alternatively, a beam search may be used to consider multiple potential candidate token sequences, since the selection of the individual most likely tokens may not necessarily result in the most likely sequence overall. Dynamic programming may be used to efficiently calculate the probabilities of the token sequences.
[0079] It may occur that multiple audio segments correspond to a single token, and there may be audio segments that correspond to silence. Repeated consecutive tokens may be collapsed into a single token instance for transcription. In the CTC algorithm, the vocabulary includes special blank tokens to allow audio segments to be labeled as silence and to provide a separator to prevent legitimate consecutive tokens, such as the double "L" in "HELLO", from being collapsed. Thus, repeating tokens are first collapsed, and then blank tokens are removed to provide the transcription. Further details regarding training and the CTC loss function are detailed further below.
[0080] Figure 5 - Architecture Figure 5 is a schematic diagram of a system for performing audio processing according to an exemplary embodiment. In particular, system 500 may implement the processes of Figures 1-4. System 500 may be implemented using one or more computer-executable instructions on one or more computing devices, such as, for example, hardware described below.
[0081] The system 500 is configured to receive an audio signal 502 capturing speech by a user. As mentioned above, the audio signal may be obtained from a sound capture device, such as a microphone, on a user device.
[0082] The system 500 comprises an encoder 504 configured to process an audio signal 502 to generate an embedding of the audio signal 506 and an initial transcription 508 of the speech. The encoder 504 comprises a non-autoregressive encoder neural network 510 configured to generate the embedding of the audio signal 506. The encoder 504 may optionally comprise a transcription engine 512 configured to generate the initial transcription 508 of the speech. For example, the transcription engine 512 may be configured to implement the CTC algorithm described above. Alternatively, the encoder neural network 510 may directly output the initial transcription 508.
[0083] The system 500 further comprises a masking engine 514 configured to modify the initial transcript 508 to generate a masked token sequence 516. As described above, the initial transcript 508 comprises one or more tokens, each of which is associated with a confidence score of being the correct token for its position within the transcript. Modifying the initial transcript 508 comprises masking one or more tokens in the initial transcript 508 that have a confidence score below a threshold confidence level. Although the masking engine 514 is shown in FIG. 5 as being part of the encoder 504, the masking engine 514 may be external to the encoder 504.
[0084] The system 500 further comprises an output decoder 518 configured to process the audio signal embedding 506 and the masked token sequence 516 to generate an audio processing output 520 .
[0085] The output decoder 518 comprises a non-autoregressive bidirectional decoder neural network 522. The speech processing output 520 comprises predicted tokens 524 for each of the masked tokens in the masked token sequence 516 to generate an output transcript 526 of the utterance, a label 528 indicating a classification of the utterance, and one or more labels 530 indicating a parameter type associated with each word in the output transcript 526, the parameter type being associated with the classification of the utterance.
[0086] The decoder neural network 522 may be configured to generate the audio processing output 520 directly, or may generate outputs for use in generating the audio processing output, such as probability distributions for making masked token predictions and determining labels 528, 530. The decoder neural network 522 may comprise a first output head 532 for the predicted tokens and a second output head 534 for the labels 528, 530. The first output head 532 is different from the second output head 534. In this way, one portion of the decoder neural network 522 is dedicated to generating the masked token predictions 524, and another portion of the decoder neural network 522 is dedicated to generating the labels 528, 530. Thus, two separate output strings may be generated, one from each output head.
[0087] Prior to processing by the output decoder 518, a prefix token may be prepended to the masked token sequence to facilitate generation of a label indicative of the classification of the utterance. The decoder neural network 522 may use the first element position to convey the utterance level information, and the first element in the output sequence with the labels 528, 530 may correspond to the utterance level label 528. The prefix token is " <cls>" can be shown as:
[0088] The decoder neural network 522 may be based on a transformer architecture. Details regarding the transformer architecture can be found in Vaswani et al., "Attention is all you need", Advances in Neural Information Processing Systems, pp. 5998-6008, 2017, which is incorporated herein by reference in its entirety. However, as mentioned above, unlike the conventional transformer decoder in Vaswani et al., the decoder neural network 522 is non-autoregressive and bidirectional. That is, while a conventional transformer decoder generates one output element at a time conditional on previously generated elements, the decoder neural network 522 simultaneously generates outputs for all positions in the output sequence. Furthermore, when generating output elements, a conventional transformer decoder avoids paying attention to the positions of elements that have not yet been generated by masking those positions. When the decoder neural network 522 is based on a transformer architecture, it is bidirectional because it pays attention to all positions in the output. Therefore, when based on a transformer architecture, the decoder neural network 522 operates without masking of future positions and is non-autoregressive. In this regard, the decoder neural network 522 operates more similarly to the encoder of a transformer. However, the decoder neural network 522 may still utilize the same structure as a conventional transformer decoder block.
[0089] 6 provides a schematic diagram of an example transformer decoder block 600 for use in constructing the decoder neural network 522. The decoder block 600 may comprise an input 602 to the block 600. For example, the input 602 may be the masked token sequence 516 if the decoder block 600 is the first decoder block of the decoder neural network 522, or the input 602 may be the output of a previous decoder block. Additionally, the input 602 of the first decoder block 600 may include a positional encoding that provides relative position information for the input elements.
[0090] The decoder block 600 may further comprise a multi-head self-attention layer 604 configured to process the input 602 and perform a self-attention operation to generate an attention layer output. In general, the self-attention operation associates pairs of elements in the input 602 to determine a new representation of the input 602. The self-attention operation is performed by applying a linear transformation matrix W q The self-attention operation may further comprise generating a query vector Q from the input 602 using a second linear transformation matrix W k The self-attention operation may further comprise generating a key vector K from the input 602 using a third linear transformation matrix W v 6. Generating a value vector V from the input 602 using a key vector. In general, the key vector serves as a means of content-based addressing for the value V. Values may be extracted based on the similarity of the query vector and the key vector.
[0091] For example, to determine a similarity score between an element of the query vector and an element of the key vector, a dot product of the query vector and the key vector may be taken. The similarity score may be scaled by a scaling factor, such as 1 / the square root of the dimension of the key vector. A softmax operation may be applied to the similarity scores to provide a set of attention weights. The attention weights may be used to perform a weighted sum of the value V to provide the output of the self-attention operation and a new representation of the input 602.
[0092] The decoder block 600 may further comprise an “Add&Norm” layer 606 configured to combine the input 602 of the block 600 with the output of the multi-head self-attention layer 604 (via residual connections 608) and perform a layer normalization operation on the combination. The combination operation may be a sum.
[0093] The decoder block 600 may further comprise a multi-head cross-attention layer 610 configured to perform an attention operation between the embedding 506 of the audio signal generated by the encoder neural network 510 and the output of the Add&Norm layer 606. The attention operation in the multi-head cross-attention layer 610 may be the same as the attention operation in the multi-head self-attention layer 604, except that the key vector and the value vector are generated from the embedding 506 of the audio signal from the encoder neural network 510. The query vector may be generated using the output of the Add&Norm layer 606. In this way, the attention operation relates the embedding 506 generated by the encoder neural network 510 to the representation provided by the decoder block 600 to modify the representation based on the information provided by the encoder neural network 510.
[0094] The decoder block 600 may further include a second Add & Norm layer 612 configured to perform the same operations as the first Add & Norm layer 606 using as inputs the output of the multi-head cross attention layer 610 and the output of the first Add & Norm layer 606 (via the residual connection 614).
[0095] The decoder block 600 may further include a feedforward layer 616 configured to process the output of the second Add&Norm layer 612 to generate a layer output. The decoder block 600 may further include a third Add&Norm layer 618 configured to perform the same operations as the first and second Add&Norm layers 606, 612 using the output of the feedforward layer 616 and the output of the second Add&Norm layer 612 (via the residual connection 620) to generate an output 622 of the decoder block 600.
[0096] The decoder neural network 522 may comprise multiple stacked decoder blocks 600. In one example, the decoder neural network 522 comprises six stacked decoder blocks, although it will be appreciated that other numbers of decoder blocks may be used. The architecture parameters for the decoder blocks may be the same as the encoder blocks for the attention type layers and the feedforward layers.
[0097] As mentioned above, the decoder neural network 522 may include a first output head 532 and a second output head 534. Each output head may include a linear layer followed by a softmax layer to generate a probability distribution. The input of the output heads may be the output of the final decoder block 600.
[0098] Referring again to the encoder 504, the encoder 504 may optionally comprise an audio processing front end 536. The audio processing front end 536 may be configured to extract acoustic features from the audio signal 502. For example, the audio signal 502 may be divided into frames of 25 ms with a shift of 10 ms. The acoustic features may be extracted frame by frame using a filter bank. In one example, the acoustic features comprise an 80-dimensional filter bank with three-dimensional pitch-related information. However, the processing for acoustic feature extraction may be performed outside the encoder 504, and the received audio signal 502 may comprise acoustic features to be processed by the encoder neural network 510.
[0099] The audio processing front end 536 may comprise a two-layer convolutional neural network, where the convolutional layer may have a kernel size of 3 with a stride of 2. This has the effect of reducing the frame rate by a factor of four.
[0100] The encoder neural network 510 may be based on a conformer architecture. Details regarding the conformer architecture can be found in Gulati et al., "Conformer: Convolution-augmented Transformer for speech recognition," in Interspeech, 2020, pp. 5036-5040, which is incorporated herein by reference in its entirety. In short, however, the conformer attempts to incorporate convolution into a transformer block to provide a conformer block. FIG. 7 provides a schematic diagram of a conformer encoder block 700.
[0101] The example conformer encoder block 700 of FIG. 7 comprises an input 702 to the encoder block 700. The encoder block 700 may comprise a first half-step feedforward layer 704 configured to process the input 702. The half-step feedforward layer is a feedforward layer that multiplies the output of the feedforward layer by 0.5. The encoder block 700 may be configured to combine 706 the output of the first half-step feedforward layer 704 with the input 702 to the encoder block 700. The combination 706 may be a summation.
[0102] The encoder block 700 may further comprise a multi-head self-attention layer 708 configured to process the output of the combination 706 to perform a self-attention operation to generate an attention layer output. The self-attention operation may be performed as described above in connection with the decoder block 600. The input to the multi-head self-attention layer 708 may be extended to include position information that may utilize a relative sinusoidal position encoding scheme.
[0103] The encoder block 700 may be further configured to combine 710 the output of the multi-head self-attention layer 708 with the output of the previous combination 706 via a residual connection. The combination 710 may also be a sum.
[0104] The encoder block 700 may further comprise a convolution layer 712 configured to process the output of the combination 710. The convolution layer 712 may implement a single convolution operation or may implement multiple convolution operations. For example, the convolution operation may comprise a point-wise convolution and / or a 1D depth-wise convolution. The convolution layer 712 may include additional operations before or after the convolution operation, such as layer normalization, gated linear units, batch normalization, swish activation, and dropout.
[0105] The encoder block 700 may be further configured to combine 714 the output of the convolutional layer 712 with the output of the previous combination 710 via a residual connection. The combination 714 may also be a sum.
[0106] The encoder block 700 may further comprise a second half-step feedforward layer 716 configured to process an output of the combination 714. The encoder block 700 may further be configured to combine 718 the output of the second half-step feedforward layer 716 with the output of the previous combination 714 via a residual connection.
[0107] The encoder block 700 may further comprise a layer normalization operation 720 configured to process the output of the combination 718 to provide an output 722 of the encoder block 700 .
[0108] The encoder neural network 510 may comprise multiple stacked conformer blocks 700. In one example, the encoder neural network 510 comprises 12 stacked conformer blocks, each block having 2048 units in the half-step feedforward layer, 4 attention heads in the multi-head self-attention layer with an attention dimension of 256, and a convolution kernel size of 15 in the convolution layer, although it will be understood that other numbers of blocks and architecture parameters may be used as deemed appropriate by those skilled in the art.
[0109] 8, the encoder neural network 510 may include an intermediate layer decoder 802 at one or more intermediate layers (i.e., hidden layers of the neural network) of the encoder neural network 510. When the encoder neural network 510 includes multiple stacked conformer blocks, the intermediate layer decoder 802 may be provided between two blocks. In one example, the intermediate layer decoders are provided at the end of the third, sixth, and ninth blocks (of the 12 blocks in the encoder neural network).
[0110] The intermediate layer decoder 802 may be configured substantially similarly to the output decoder 518, may have the same architecture, and may share parameters.
[0111] A particular intermediate layer 804 may be configured to generate an intermediate layer embedding 806 of the audio signal, which may simply be the output of the intermediate layer 804. An intermediate layer transcript 808 may be generated from the intermediate layer embedding 806. The encoder 504 may further comprise an intermediate layer transcription engine 810 configured to generate the intermediate layer transcript 808. Alternatively, the transcription engine 512 may be used, or the intermediate layer 804 may directly provide the intermediate layer transcript 808.
[0112] The encoder 504 may further comprise an intermediate layer masking engine 812 configured to modify the intermediate layer transcript 808 to generate an intermediate layer masked token sequence 814, similar to how the masking engine 514 is configured to modify the initial transcript 508 to generate a masked token sequence 516. That is, modifying the intermediate layer transcript 808 comprises masking one or more tokens in the intermediate layer transcript 808 that have a confidence score below an intermediate layer threshold confidence level. The intermediate layer threshold may be set to a different value than the threshold confidence level used to mask the initial transcript 508. If there are multiple intermediate layer decoders, the intermediate layer threshold may be different thereamong and / or may have the same value. For example, if there are intermediate layer decoders at the end of the third, sixth, and ninth blocks, the intermediate layer threshold may be 0.9, 0.99, and 0.999, respectively. Instead of a separate intermediate layer masking engine 812, the masking engine 514 may be used.
[0113] The intermediate layer decoder 802 may be configured to process the intermediate layer embedding 806 of the audio signal and the intermediate layer masked token sequence 814 to generate an intermediate layer decoded output 816. The encoder 504 is configured to combine the intermediate layer embedding 806 of the audio signal and the intermediate layer decoded output 816 to provide an input 818 for one or more subsequent neural network layers of the encoder neural network 510 following the respective intermediate layer 804 of the encoder neural network 510. The combination may be a concatenation or summation or other function deemed appropriate by one of ordinary skill in the art. Exemplary combining operations are described in more detail below.
[0114] One or more subsequent neural network layers may be configured to process the input 818 to generate an embedding 506 of the audio signal and an initial transcription 508 of the speech as described above.
[0115] The intermediate layer decoded output 816 may comprise data regarding prediction of masked tokens of the intermediate layer masked token sequence 814, generation of speech level classifications, and generation of parameter type labels. For example, the intermediate layer decoded output 816 may comprise a masked token probability distribution comprising a probability distribution over candidate tokens for predicting a token for each of the masked tokens in the intermediate layer masked token sequence 814. The intermediate layer decoded output 816 may comprise an utterance classification probability distribution comprising a probability distribution over labels indicating a classification of the utterance. The intermediate layer decoded output 816 may comprise a parameter type probability distribution comprising a probability distribution over labels indicating a parameter type for each word in the intermediate layer transcript 808.
[0116] In this regard, the encoder 504 may be configured to combine the hidden layer embedding 806 of the audio signal and the hidden layer decoded output 816 by combining the masked token probability distribution, the utterance classification probability distribution, and the parameter type probability distribution. The probability distributions may be combined by concatenating the distributions into a single vector or tensor, or by other suitable techniques. The encoder 504 may further be configured to perform a linear projection of the combined probability distribution such that the dimensionality matches the hidden layer embedding 806 of the audio signal. The linear projection may be learned as part of training the encoder, which is described in more detail below.
[0117] The encoder 504 may be further configured to sum the projected combined probability distribution and the intermediate layer embedding 806 of the audio signal to provide an input 818 for one or more subsequent neural network layers. The combination operation may be implemented as a further neural network layer of the encoder neural network 510.
[0118] Further combining operations will now be described with reference to FIG. 9. The audio signal 502 may comprise a plurality of audio segments. The encoder 504 may be configured to generate an intermediate layer transcription 808 of the utterance based on the generated probability distributions over the candidate tokens for transcribing each audio segment and generating an intermediate layer transcription 808 based on the probability distribution for each audio segment. The probability distributions may be generated based on the intermediate layer embedding 806. For example, as shown in FIG. 9, out The hidden layer embeddings 806, labeled Z, are transformed using a linear neural network layer and a softmax layer 902 to generate a probability distribution over the candidate tokens for each audio segment. The probability distribution for each audio segment is represented by a vector labeled Z 904.
[0119] The encoder 504 may then be configured to generate an intermediate layer transcription 808 from the probability distribution 904, for example, using the CTC algorithm described above. In FIG. 9, the intermediate layer transcription 808 is labeled Ŷ 906 and contains the token y 1 , y 2 , y 3 , y 4 The encoder 504 may be configured to determine the audio segment position corresponding to each token in the intermediate layer transcript 808, which is later used in the combination operation.
[0120] As described above, the encoder 504 may be configured to generate the hidden layer masked token sequence 814. In FIG. 2 teeth," <mask>" tokens. Furthermore, to generate utterance-level classification probability distributions / labels, <cls>" token is prefixed to the beginning of the masked token sequence 814 of the intermediate layer.
[0121] As described above, the intermediate layer decoder 802 may be configured to generate a probability distribution of the masked tokens, a speech classification probability distribution, and a parameter type probability distribution. The probability distribution of the masked tokens and the parameter type probability distribution may be mapped to audio segment positions based on corresponding tokens in the intermediate layer transcript 808. This is shown in FIG. ’2 The probability distribution 908 of the masked tokens labeled with <mask>907. Similarly, the parameter type probability distributions are aligned with their corresponding tokens, i.e., o s 1 910 is y 1 It is aligned with 912, s 2 914 is <mask> / y 2 It is aligned with 907. s 3 916 is y 3 It is aligned with 918, s 4 920 is y 4 It is positioned in a row with 922.
[0122] A speech classification probability distribution may be associated with each of the audio segment positions in the mapping, and in FIG. i 924 is shown diagrammatically by being broadcast at each of the audio segment positions.
[0123] The probability distributions at each audio segment position may then be combined. The combination may be the concatenation of the probability distributions mapped to the corresponding audio segment positions. To facilitate the combination, the probability distributions over the candidate tokens 904 may include additional slots in which other probability distributions are placed. Thus, the token vocabulary may include speech level classes and parameter type classes.
[0124] 9, the encoder neural network 510 may further comprise a linear neural network layer 926 configured to apply a linear projection to the combined probability distribution such that its dimensions match those of the hidden layer embedding 806. The encoder neural network 510 may further comprise a layer normalization layer 928 configured to perform a layer normalization operation on the hidden layer embedding 806 before summing with the linearly projected combined probability distribution.
[0125] Figure 10-Training Figure 10 is a flow diagram illustrating a process 1000 for training a speech processing system, such as the system of Figure 5. The speech processing system comprises an encoder and an output decoder. The encoder comprises a non-autoregressive encoder neural network and the output decoder comprises a non-autoregressive bidirectional decoder neural network, both of which comprise a number of trainable parameters. The training aims to obtain optimal values of the trainable parameters for the speech processing task described herein.
[0126] In step 1005, an audio signal capturing an utterance by a user is received to train the speech processing system. The audio signal may be obtained from a training data set. For example, a suitable training data set may be the SLURP data set, which is described in Bastianelli et al. "SLURP: A spoken language understanding resource package," in 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 7252-7262, which is incorporated herein by reference in its entirety. In summary, the SLURP data set comprises 140,000 utterances from the domain of a personal robotic assistant in the home.
[0127] In step 1010, the audio signal is processed by the encoder to generate an embedding of the audio signal and a conditional probability distribution over candidate transcriptions of the utterance given the input audio signal. The conditional probability distribution may be generated based on the particular method used to generate the candidate transcriptions of the utterance. For example, the conditional probability distribution may be generated based on the CTC algorithm. Further details regarding the generation of the conditional probability distribution are described below.
[0128] In step 1015, a truth-based transcription of the utterance is received. The truth-based transcription comprises a sequence of one or more tokens and is one target output for the system. The training data set typically includes a truth-based transcription for the input audio signal. It will be appreciated that the truth-based transcription may be received at an earlier point in time, for example, along with the audio signal in step 1010.
[0129] In step 1020, an encoder loss value is calculated based on the conditional probability distribution over the candidate transcriptions of the utterance and the true-value transcription of the utterance. The encoder loss value may indicate an error caused by the encoder when the conditional probability distribution is used to generate a transcription of the utterance. The encoder loss value may be calculated based on a loss function and may comprise a negative log-likelihood. The loss function may also depend on the particular transcription method described above. For example, the encoder loss value may be based on a CTC loss function. Exemplary loss functions are described in more detail below.
[0130] In step 1025, the truth-based transcript is modified to generate a masked token sequence, i.e., one or more tokens in the truth-based transcript are modified to generate a masked token sequence, i.e., <mask>" token. The number of tokens selected for masking may be randomly selected from a uniform distribution. Alternatively, the number of tokens selected for masking may be determined as a proportion of the length of the true transcript, or a fixed number of tokens may be selected for masking. The token or tokens selected for masking may also be randomly selected from a uniform distribution.
[0131] In step 1030, the embedding of the audio signal and the masked token sequence are processed by an output decoder to generate a speech processing output comprising a probability distribution over candidate tokens for predicting a token for each of the masked tokens in the masked token sequence, a probability distribution over labels indicative of a speech classification, and a probability distribution over labels indicative of a parameter type associated with each word in the output transcript, the parameter type being associated with a speech classification. Each of the probability distributions may be generated as described above.
[0132] In step 1035, a truth label is received, which indicates the classification and parameter type of the utterance, respectively. The truth label comprises another target output of the system. The training data set typically includes the truth label. It will be appreciated that the truth label may be received earlier, for example, with the truth transcript in step 1015 or with the audio signal in step 1005.
[0133] In step 1040, a decoder loss value is calculated based on the speech processing output, the ground truth transcription, and the ground truth label. The decoder loss value may indicate an error caused by the decoder when a probability distribution from the speech processing output is used to generate an output transcription of the speech and a classification of the speech level and parameter type. The decoder loss value is calculated based on a loss function. The loss function may be a joint loss function and may comprise an encoder loss value and a decoder loss value as components. The decoder loss value may comprise a negative log-likelihood. Exemplary loss functions are described in more detail below.
[0134] In step 1045, values of trainable parameters of the decoder neural network and the encoder neural network are adjusted based on the decoder loss value. In step 1050, values of trainable parameters of the encoder neural network are adjusted based on the encoder loss value. Adjustment of the trainable parameters may be performed using standard neural network techniques, for example, using backpropagation and stochastic gradient descent. In one example, stochastic gradient descent with mini-batches of size 64 is used.
[0135] The training process 1000 may be repeated for a number of iterations and / or until a stopping criterion is reached. In one example, the training process is repeated for up to 400 epochs (i.e., 400 passes through the training data set).
[0136] The training data set may be augmented by applying the SpecAugment algorithm. Details regarding SpecAugment can be found in Park et al., "Specaugment: A simple data augmentation method for automatic speech recognition," in Interspeech, 2019, pp. 2613-2617, which is incorporated herein by reference in its entirety. In brief, for each training audio signal, SpecAugment converts the audio signal into a spectrogram format and then applies one or more transforms to the spectrogram. Exemplary transforms include warping in the time dimension, masking blocks of continuous frequency channels, and masking blocks of speech in time. The transformed spectrogram may be converted back to an audio signal, or alternatively, the system may process the transformed spectrogram directly as an acoustic feature or to generate acoustic features.
[0137] Referring now to FIG. 11, the audio processing system comprises an intermediate layer decoder, such as the systems of FIGS. 8 and 9, and the processing by the encoder in step 1010 during training may further comprise the following process 1100.
[0138] In step 1105, an intermediate layer conditional probability distribution is generated over the candidate transcriptions of the utterance given the intermediate layer embedding and the outputs of the layers preceding the intermediate layer. This intermediate layer conditional probability distribution may be generated as described above with reference to Figures 8 and 9.
[0139] In step 1110, an intermediate layer loss value is calculated based on the intermediate layer conditional probability distribution and the true value transcription. The encoder loss value may be further based on the intermediate layer loss value. For example, the intermediate layer loss value may be summed with the encoder loss value described above. The intermediate layer loss value may be calculated similarly to the encoder loss value, for example, using a CTC loss function. Further details regarding the intermediate layer loss value are described below.
[0140] In step 1115, an intermediate-layer transcription of the utterance is generated based on the intermediate-layer conditional probability distribution. The intermediate-layer transcription may be generated, for example, using the CTC algorithm, as described above.
[0141] In step 1120, the intermediate layer transcript is modified to generate an intermediate layer masked token sequence. That is, as described above, one or more tokens in the intermediate layer transcript that have a confidence score below the intermediate layer threshold confidence level are marked as " <mask>" token.
[0142] In step 1125, the intermediate layer embedding of the audio signal and the intermediate layer masked token sequence are processed by an intermediate layer decoder to generate an intermediate layer decoded output. The intermediate layer decoder comprises a non-autoregressive bidirectional decoder neural network having a number of trainable parameters. The trainable parameters of the intermediate layer decoder may be shared with the output decoder and any other intermediate layer decoders present.
[0143] In step 1130, the intermediate layer embedding of the audio signal and the intermediate layer decoded output are combined to provide input for one or more subsequent neural network layers of the encoder neural network following the respective intermediate layer of the encoder neural network. The combining may be performed as described above with reference to Figures 8 and 9.
[0144] In step 1135, the input is processed by one or more subsequent neural network layers to generate an embedding of the audio signal and a conditional probability distribution over candidate transcriptions of the utterance given the input audio signal.
[0145] Next, exemplary loss functions suitable for calculating the encoder loss value and the decoder loss value are detailed. As mentioned above, the encoder loss value may be based on a particular decoding method for generating the transcription. For example, if the CTC algorithm is used, a standard CTC loss function may be used as the encoder loss function for calculating the encoder loss value.
[0146] As described above, the encoder neural network may provide a conditional probability distribution over the token vocabulary for each temporal audio segment. Tokens are selected for each temporal audio segment to provide a token sequence known as an alignment. More formally, the conditional probability of an alignment A given input audio X may be formulated as a product of token probabilities per segment as follows:
[0147]
number
[0148] where T is the number of time audio segments.
[0149] The conditional probability of a transcription Y given an input audio X may be computed by marginalizing over the set of valid alignments for transcription Y.
[0150]
number
[0151] Here, β -1 (Y) returns all possible valid alignments for transcript Y. Given the nature of CTC alignments, the sum can be computed efficiently using dynamic programming.
[0152] The CTC loss function may be based on minimizing the negative log-likelihood. L ctc =-logP ctc (Y|X)
[0153] The encoder loss value may be the negative log-likelihood for an input training audio signal and its corresponding ground truth transcription. Alternatively, depending on the type of update rule used to adjust the trainable parameters, the encoder loss value may be the sum of the negative log-likelihoods over a batch of training data items or the entire training data set.
[0154] If an inter-layer transcription is generated, the encoder loss value may further comprise an inter-layer loss value, which may be calculated similarly to the encoder loss value and combined to generate an overall encoder loss value. For example, the overall encoder loss value may be a weighted sum of the encoder loss value and the inter-layer loss value according to the following formula: L enc = ηL ctc +(1-η)L inter-ctc Here, L enc is the overall encoder loss value, and L ctc is the encoder loss value calculated based on the initial transcription, and L inter-ctc is the intermediate layer loss value, and η is a pre-defined hyper-parameter. In one example, η is set to 0.5, but it will be understood that other values may be used.
[0155] If multiple intermediate layer transcripts are generated, an intermediate layer loss value may be calculated for each one and an average may be taken to provide an overall intermediate layer loss value for use in the above equation. Note that in these examples, the intermediate layer loss value is not directly based on any of the intermediate layer decoded outputs. The intermediate layer decoded outputs will take into account processing by the remainder of the encoder neural network, which generates the encoder transcript (initial transcript) and any further intermediate intermediate layer transcripts, and thus are only indirectly incorporated into the encoder loss value.
[0156] The decoder loss value may comprise a first component based on a probability distribution over the candidate tokens for predicting the masked token and a second component based on a probability distribution over the utterance level labels and the parameter type labels. For example, the conditional probability distribution for predicting the masked token may be expressed as follows:
[0157]
number
[0158] Here, Y mask is the set of masked tokens, and Y obs where X is the set of observed tokens and X is the embedding of the audio signal produced by the encoder neural network. The first component of the decoder loss value can then be calculated as the negative log-likelihood, as shown in the following equation: L dec-mask =-logP dec-m (Y mask |Y obs ,X)
[0159] Like the encoder loss value, the negative log-likelihood may be calculated for a single training audio signal and its corresponding true value, or may be the sum of the negative log-likelihoods over a batch of training audio signals or an entire training data set.
[0160] Probability distribution P dec-mask can be parameterized by the decoder neural network, and in particular by the first output head with a softmax output as described above.
[0161] The probability distribution over the utterance level labels and parameter type labels may be parameterized by the decoder neural network and by the second output head, which also has a softmax output as described above. The probability distribution may be expressed according to the following equation:
[0162]
number
[0163] Here, O=[o i ;o s ] and utterance level label o i Parameter type label o s ={o s 1 ,…,o s n }, Y is a masked token sequence including the prefix token, {h0,…h n } is the final hidden state of the second output head of the decoder neural network, and N is the length of the true transcription.
[0164] The second component of the decoder loss value may be calculated as the negative log-likelihood: L dec-labels =-logP dec-labels (O|Y,X)
[0165] Again, the negative log-likelihood may be calculated for a single training audio signal and its corresponding truth label, or may be the sum of the negative log-likelihoods over a batch of training audio signals or the entire training data set.
[0166] The overall decoder loss value may be a weighted sum of the first and second components above, as shown below: L dec = γL dec-mask +(1-γ)L dec-labels where γ is a pre-defined hyper-parameter. In one example, γ is set to 0.5, although it will be appreciated that other values may be used.
[0167] The encoder and decoder loss values may also be combined to form an overall loss function for training the system. For example, a weighted sum may be calculated as shown below: L sys = μL enc +(1-μ)L dec where μ is a pre-defined hyper-parameter. In one example, μ is set to 0.4, although it will be understood that other values may be used.
[0168] The global loss function may be used to determine error values for the backpropagation and gradient descent methods that determine adjusted values for the trainable parameters of the encoder neural network and the decoder neural network. It will be appreciated that the encoder loss value component of the global loss function is used to determine the adjusted values for the trainable parameters of the encoder neural network, and the decoder loss value component of the global loss function is used for both the decoder neural network and the encoder neural network given the structure of the encoder / decoder architecture of the audio processing system.
[0169] The token vocabulary may be generated from the training data set using the BPE. In one example, the token vocabulary comprises 500 word pieces generated using the BPE from the SLURP data set. The token vocabulary may also be expanded to include utterance level labels and parameter type labels. For example, the token vocabulary may comprise 70 utterance level labels and 56 parameter types from the SLURP data set. The token vocabulary may also include <mask> 、 <cls> 、 <eos> 、 <sos>, and <blank>In one example, the token vocabulary comprises a total of 682 tokens from a combination of word piece, label, and function tokens.
[0170] Computer Hardware 12 is a schematic diagram of hardware that may be used to implement methods and systems according to embodiments described herein. Note that this is just one example and other arrangements may be used.
[0171] The hardware comprises a computing section 1200. In this particular example, the components of this section are depicted in one location, however, it will be understood that they are not necessarily co-located.
[0172] Components of computing system 1200 may include, but are not limited to, a processing unit 1213 (such as a central processing unit (CPU)), a system memory 1201, and a system bus 1211 that couples various system components including the system memory 1201 to the processing unit 1213. The system bus 1211 may be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. The computing section 1200 also includes an external memory 1215 connected to the bus 1211.
[0173] The system memory 1201 includes computer storage media in the form of volatile and / or nonvolatile memory, such as read-only memory. A basic input / output system (BIOS) 1203, containing the routines that help to transfer information between elements within a computer, such as during start-up, is typically stored in the system memory 1201. Additionally, the system memory includes an operating system 1205, application programs 1207, and program data 1209 used by the CPU 1213.
[0174] Also connected to the bus 1211 is an interface 1225. The interface may be a network interface through which the computer system receives information from additional devices. The interface may also be a user interface that allows a user to respond to specific commands and the like.
[0175] In this example, a video interface 1217 is provided. The video interface 1217 comprises a graphics processing unit 1219 connected to a graphics processing memory 1221.
[0176] The graphics processing unit (GPU) 1219 is particularly well suited for training a speech recognition system due to its adaptation to data-parallel operations such as training neural networks, so in one embodiment, processing for training a speech recognition system may be split between the CPU 1213 and the GPU 1219.
[0177] It should be noted that in some embodiments, different hardware may be used for training the speech recognition system and for performing the speech recognition. For example, the training of the speech recognition system may be performed on one or more local desktop or workstation computers, or on a device of a cloud computing system, which may include one or more separate desktop or workstation GPUs, one or more separate desktop or workstation CPUs, such as a processor with a PC-oriented architecture, and a large amount of volatile system memory, such as 16 GB or more. On the other hand, for example, mobile or embedded hardware may be used for performing the speech recognition, which may include a mobile GPU that is part of a system on a chip (SoC) or may not include a GPU, and may further include one or more mobile or embedded CPUs, such as a processor with a mobile-oriented or microcontroller-oriented architecture, and a small amount of volatile memory, such as less than 1 GB. For example, the hardware performing the speech recognition may be a voice assistant system 120, such as a smart speaker or a mobile phone including a virtual assistant. The hardware used to train the speech recognition system may have significantly more computational power, such as being able to perform more operations per second, and have a larger memory than the hardware used to perform tasks using agents. Performing speech recognition, e.g., by performing inference using one or more neural networks, is significantly less computationally resource intensive than, e.g., training a speech recognition system by training one or more neural networks, and thus hardware with fewer resources can be used. Furthermore, techniques can be used to reduce the computational resources used to perform speech recognition, e.g., to perform inference using one or more neural networks. Examples of such techniques include model distillation and, in the case of neural networks, neural network compression techniques such as pruning and quantization.
[0178] Although specific embodiments have been described, these embodiments are presented by way of example only and are not intended to limit the scope of the invention. The novel devices and methods described herein may be embodied in a variety of other forms, and various omissions, substitutions, and changes may be made in the form of the devices, methods, and products described herein without departing from the spirit of the invention. The appended claims and their equivalents are intended to cover any such forms or modifications that fall within the scope and spirit of the invention.< / blank> < / sos> < / eos> < / cls> < / mask> < / mask> < / mask> < / mask> < / mask> < / cls> < / mask> < / cls> < / mask> < / mask> < / mask> < / mask> < / cls> < / mask> < / cls>
Claims
1. 1. A computer-implemented method for audio processing, comprising: Receiving an audio signal capturing speech by a user; processing the audio signal with an encoder comprising a non-autoregressive encoder neural network to generate an embedding of the audio signal and an initial transcription of the speech, wherein the initial transcription comprises one or more tokens, each of the one or more tokens associated with a first confidence score of being a correct token for a position within the initial transcript; modifying the initial transcript to generate a masked sequence of tokens, where modifying the initial transcript comprises masking the one or more tokens in the initial transcript having a first confidence score below a threshold confidence level; processing the embedding of the audio signal and the masked sequence of tokens by an output decoder comprising a non-autoregressive bidirectional decoder neural network to generate a speech processing output; and the audio processing output comprises: a predicted token for each of a plurality of masked tokens in the masked token sequence generating an output transcription of the utterance; and A label indicating a classification of the utterance; one or more labels indicating a plurality of parameter types associated with a plurality of respective first words in the output transcript, where the plurality of parameter types are associated with the classification of the utterance. Equipped with The decoder neural network of the output decoder comprises: a first output head for generating the predicted token for each of the plurality of masked tokens in the masked token sequence; a second output head for generating the one or more labels indicative of a plurality of parameter types associated with the plurality of respective first words in the output transcript; Equipped with The method, wherein the first output head and the second output head are different output heads.
2. A computer-implemented method for speech processing, comprising: Receiving an audio signal capturing speech by a user; processing the audio signal with an encoder comprising a non-autoregressive encoder neural network to generate an embedding of the audio signal and an initial transcription of the speech, wherein the initial transcription comprises one or more tokens, each of the one or more tokens associated with a first confidence score of being a correct token for a position within the initial transcript; modifying the initial transcript to generate a masked sequence of tokens, where modifying the initial transcript comprises masking the one or more tokens in the initial transcript having a first confidence score below a threshold confidence level; processing the embedding of the audio signal and the masked sequence of tokens by an output decoder comprising a non-autoregressive bidirectional decoder neural network to generate a speech processing output; and the audio processing output comprises: a predicted token for each of a plurality of masked tokens in the masked token sequence generating an output transcription of the utterance; and A label indicating a classification of the utterance; one or more labels indicating a plurality of parameter types associated with a plurality of respective first words in the output transcript, where the plurality of parameter types are associated with the classification of the utterance. Equipped with The method, wherein the encoder neural network is based on a conformer architecture.
3. Generating the output transcription of the utterance comprises iteratively refining a plurality of token predictions for the masked tokens in the masked token sequence, the iterative refinement process comprising: generating a predicted token for each of the plurality of masked tokens in the masked token sequence, wherein each predicted token is associated with the first confidence score of being the correct token; for each predicted token, replacing a corresponding masked token in the masked token sequence with the predicted token when the first confidence score is above the threshold confidence level and retaining the corresponding masked token in the masked token sequence when the first confidence score is equal to or less than the threshold confidence level; The method of claim 1 or claim 2, comprising:
4. The processing by the encoder includes, in one or more hidden layers of the encoder neural network, generating an intermediate layer embedding of the audio signal and an intermediate layer transcription of the utterance, wherein the intermediate layer transcription comprises one or more tokens, each of the one or more tokens being associated with a second confidence score of being a correct token for its position within the intermediate layer transcript; modifying the middle layer transcript to generate a middle layer masked token sequence, where modifying the middle layer transcript comprises masking the one or more tokens in the middle layer transcript having a second confidence score below a middle layer threshold confidence level; processing the intermediate layer embedding of the audio signal and the intermediate layer masked token sequence with an intermediate layer decoder comprising a non-autoregressive bidirectional decoder neural network to generate an intermediate layer decoded output; combining the intermediate layer embedding of the audio signal and the intermediate layer decoded output to provide input for one or more subsequent neural network layers of the encoder neural network following a respective intermediate layer of the encoder neural network; processing the input by the one or more subsequent neural network layers to generate the embedding of the audio signal and the initial transcription of the speech; The method of claim 1 or claim 2, comprising:
5. The intermediate layer decoded output is a masked token probability distribution comprising a probability distribution over a plurality of candidate tokens for predicting a token for each of a plurality of masked tokens in the masked token sequence of the middle layer; an utterance classification probability distribution comprising a probability distribution over a plurality of labels indicative of a classification of the utterance; a parameter type probability distribution comprising a probability distribution over a plurality of labels indicating a parameter type for a plurality of respective second words in the intermediate layer transcript; The method of claim 4 comprising:
6. Combining the intermediate layer embedding of the audio signal and the intermediate layer decoded output includes: combining the masked token probability distribution, the utterance classification probability distribution, and the parameter type probability distribution; performing a linear projection of a plurality of combined probability distributions such that the dimensionality matches the intermediate-layer embedding of the audio signal; summing the linearly projected combined probability distribution and the intermediate layer embedding of the audio signal; The method of claim 5 comprising:
7. the audio signal comprises a plurality of audio segments, and generating the mid-layer transcription of the utterance comprises generating a probability distribution over a plurality of candidate tokens for transcribing each audio segment, and generating the mid-layer transcription based on the plurality of combined probability distributions for each audio segment; Combining the intermediate layer embedding of the audio signal and the intermediate layer decoded output includes: determining an audio segment position corresponding to each token in the intermediate layer transcript; mapping the masked token probability distributions and the parameter type probability distributions to a plurality of audio segment locations based on a plurality of corresponding tokens in the intermediate-layer transcript; associating the utterance classification probability distribution with each of the plurality of audio segment positions in the mapping; Equipped with Combining the masked token probability distribution, the utterance classification probability distribution, and the parameter type probability distribution comprises: combining the probability distributions over a plurality of candidate tokens having a plurality of corresponding audio segment positions according to the mapping, the probability distributions of the masked tokens, the speech classification probability distribution, and the parameter type probability distribution. The method according to claim 6.
8. The method of claim 4 , wherein the initial transcription and / or the intermediate-layer transcription of the utterance are generated based on a CTC algorithm.
9. prepend a prefix token to the masked token sequence prior to processing by the output decoder to generate the label indicative of the classification of the utterance; and / or prepend a prefix token to the sequence of masked tokens of the intermediate layer prior to processing by the intermediate layer decoder to generate the utterance classification probability distribution; The method of claim 5 further comprising:
10. The method of claim 4 , wherein the decoder neural network of the output decoder and / or the intermediate layer decoder is based on a Transformer architecture.
11. The method of claim 1 or claim 2, wherein the classification of the utterance comprises an indication of an action to be taken in response to the utterance.
12. a memory having processor-readable instructions stored therein; a processor configured to read and execute instructions in the memory; 3. A system comprising: said processor-readable instructions comprising instructions arranged to control said system to perform a method according to claim 1 or claim 2.
13. A non-transitory computer readable medium carrying computer readable instructions configured to cause a computer to perform the method of claim 1 or claim 2.
Citation Information
Patent Citations
Speech recognition device
JP1990052400A
Voice input support device
JP2012078650A
Text generation device, text generation program and text generation method
JP2020034704A
Decoding method in artificial neural network, speech recognition device, and speech recognition system
JP2020086436A
Speech recognition device, speech recognition method and speech recognition program
JP2021039216A