End-to-end speech recognition suitable for multi-speaker applications
By using multi-head encoder and decoder in end-to-end speech recognition system, combined with GTC objective function and directed graph supervision information, the problem of speech recognition output delay in multi-speaker applications is solved, and more efficient real-time speech recognition is achieved.
Patent Information
- Application Number
- CN202380072998.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-26
- Filing Date
- 2023-07-12
- Publication Date
- 2025-05-30
AI Technical Summary
The existing end-to-end speech recognition system has output delay problems in multispeaker applications and it is difficult to handle speech separation and recognition tasks simultaneously, resulting in increased delays.
By performing speaker identification at the encoder level and allowing the decoder to decode both speech and speaker, using a combination of multi-head encoder and decoder, combining GTC objective functions and supervised information of directed graphs, the neural network is trained to reduce output latency.
It realizes reducing the delay in speech recognition output in multispeaker applications, improves the real-time performance of the speech recognition system, and is suitable for streaming and online speech recognition tasks.
Smart Images

Figure CN120077431A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to artificial intelligence (AI) systems for speech recognition, and more particularly to methods and systems for end-to-end speech recognition suitable for multi-speaker applications. Background Art
[0002] Neural networks can reproduce and model non-linear processes. Therefore, in the past few decades, neural networks have been used in many applications in various disciplines. A neural network can be learned (or trained) by processing examples, each of which contains a known "input" and "result", thereby forming a probability-weighted association between the two, which is stored in the data structure of the network itself. Training of a neural network from a given example is typically performed by determining the difference between the processed output (usually a prediction) of the network and the target output, which is also referred to herein as the training label. This difference represents the error that training aims to reduce. The network then adjusts its weighted associations according to a learning rule and using this error value. Successive adjustments will cause the neural network to produce an output that increasingly resembles the target output. After a sufficient number of these adjustments, training can be terminated based on certain criteria.
[0003] This type of training is generally referred to as supervised learning. During supervised learning, a neural network "learns" to perform a task by considering examples, usually without being programmed with task-specific rules. For example, in image recognition, they can learn to recognize images containing cats by analyzing example images that have been labeled as "cat" or "no cat" and using the results to identify cats in other images. They do this without any prior knowledge of what a cat is (e.g., a cat has fur, a tail, whiskers, and a cat-like face). Instead, they automatically generate identifying features from the examples they process.
[0004] However, in order to perform this supervised learning, the images need to be labeled as cats or dogs. This labeling is a tedious and laborious process. In addition, in this image recognition example, the labeling is unambiguous. The image contains a cat, a dog, or does not contain a cat or a dog. Such unambiguous labeling is not always possible. For example, some training applications address sequence problems where timing is a variable. The time variable can create one-to-many or many-to-one ambiguities in such training, where the input sequence has a different length from the output sequence.
[0005] Specifically, some methods for training neural networks use the connectionist temporal classification (CTC) objective function algorithm. CTC is a loss function that is used to train a neural network when there is no available temporal alignment information between a labeled training sequence and a longer sequence of label probabilities output by the neural network, where the longer sequence of label probabilities output by the neural network is computed from an observed sequence input to the neural network. This missing temporal alignment information creates a temporal ambiguity between the sequence of label probabilities output by the neural network and the supervisory information used for training, which is the labeled training sequence that can be resolved using the CTC objective function.
[0006] However, the CTC objective function is only suitable for resolving temporal ambiguity during neural network training. If other types of ambiguity need to be considered, the CTC objective function cannot do so.
[0007] A generalization of the CTC objective function is graph based temporal classification (GTC), which is a loss function for training deep neural networks using a graph representation in the loss function. The GTC loss function is used to handle sequence-to-sequence temporal alignment ambiguity resolution using a deep neural network. GTC can take graph-based supervisory information as input to describe all possible alignments between an input sequence and an output sequence in order to learn the best possible alignment from the training data.
[0008] An example of a sequence-based input to a neural network that requires a solution for temporal and label ambiguity is an audio input. The audio input can be in the form of speech from one or more speakers, which may need to be recognized and separated for audio applications.
[0009] One example of such an audio application is in an automatic speech recognition (ASR) system that is widely deployed for various interface applications such as voice search. However, it is challenging to manufacture a speech recognition system that achieves high recognition accuracy. This is because such manufacturing requires in-depth language knowledge of the target language that the ASR system accepts. For example, a set of phonemes, vocabulary, and pronunciation dictionaries are essential for manufacturing such an ASR system. The phoneme set needs to be carefully defined by linguists of the language. The pronunciation dictionary needs to be manually created by assigning one or more phoneme sequences to each word in a vocabulary that includes over 100,000 words. In addition, some languages do not have explicit word boundaries, so we may need to tokenize to create a vocabulary from a text corpus. Therefore, it is quite difficult to develop a speech recognition system, especially for minority languages. Another problem is that the speech recognition system is decomposed into several modules including an acoustic, a dictionary, and a language model, and these modules are optimized separately. This architecture may lead to local optima, but each model needs to be trained to match the other models.
[0010] Recently, end-to-end neural network models and sequence-to-sequence neural network models have gained increasing attention and popularity in the ASR community. The output of an end-to-end ASR system is typically a sequence of graphemes, which can be a single letter or a larger unit such as a word fragment and an entire word. The attraction of end-to-end ASR is that, compared with traditional ASR systems, it achieves a simplified system architecture by consisting of neural network components and avoiding the need for language expert knowledge to build the ASR system.
[0011] End-to-end ASR systems can directly learn all elements of a speech recognizer, including pronunciation, acoustic, and language models, which avoids the need for language-specific linguistic information and text normalization. These ASR systems perform a sequence-to-sequence transformation, where the input is a sequence of acoustic features extracted from audio frames at a specific rate, and the output is a sequence of characters. The sequence-to-sequence transformation allows various language features to be considered to improve recognition quality.
[0012] However, the improvement in the quality of end-to-end ASR systems comes at the cost of output latency due to the need to accumulate acoustic feature sequences and / or acoustic frame sequences for joint recognition. Therefore, end-to-end ASR systems are less suitable for online / streaming ASR that requires low latency.
[0013] A variety of techniques (such as triggered attention or restricted self-attention) have been developed for reducing output latency in end-to-end ASR systems. See, for example, U.S. Patent #11,100,920. However, these techniques are not applicable or at least not directly applicable to multi-speaker recognition and / or multi-speaker streaming applications. This is because multi-speaker applications involve two separate tasks: speaker separation and speech recognition. Currently, speaker separation in multi-speaker ASR systems is a preprocessing or postprocessing technique that introduces additional latency that current methods for streaming end-to-end speech recognition cannot handle.
[0014] Accordingly, there is a need to reduce output latency in multi-speaker applications applicable to end-to-end and / or sequence-to-sequence speech recognition applications. SUMMARY OF THE INVENTION
[0015] An object of some embodiments is to reduce output latency in multi-speaker applications configured for end-to-end and / or sequence-to-sequence speech recognition applications. An example of such an application is a streaming speech recognition application. Some embodiments are based on the understanding that, in order to reduce latency in multi-speaker speech recognition applications, the speech separation task and the speech recognition task should be considered jointly such that speech recognition is performed concurrently with speech separation. Doing so prevents additional latency in speech recognition caused by preprocessing or postprocessing techniques for speech separation.
[0016] Additionally or alternatively, some embodiments are based on the recognition that if speech separation is considered jointly with speech recognition, then speech separation can be replaced with speaker identification. In contrast to speech separation, which is considered an independent task, speaker identification can be regarded as a task subordinate to speech recognition. Thus, speaker identification can be implemented as an internal process of speech recognition.
[0017] With this in mind, some embodiments are based on the understanding that speech recognition in end-to-end ASR systems is typically performed using an encoder and a decoder. To make speaker identification an internal process, some embodiments perform speaker identification at the encoder level while allowing the decoder to decode both speech and speakers. In this way, speech separation is transformed into a part of decoding that does not cause additional latency. However, to achieve this effect, the encoder needs to be a multi-head or multi-output encoder that produces symbols for each audio frame and the identity of the speaker.
[0018] Some embodiments are based on the recognition that supervised information in a directed graph having nodes connected by edges representing labels and transitions between the labels allows for the imposition of flexible rules for training a neural network. For example, some embodiments disclose training a neural network using a GTC objective function without inserting blank labels between all training labels or using multiple different blank labels. Additionally or alternatively, some embodiments disclose training a neural network with a GTC objective using a topology similar to a hidden Markov model (HMM) for each label, which can include multiple states. Additionally or alternatively, some embodiments disclose using a directed graph having transitions between nodes associated with costs or weighting factors to train a neural network with a GTC objective.
[0019] In addition to using the supervised information residing on the directed graph, some embodiments modify the GTC objective function to accommodate label alignment. For example, the GTC objective function is defined by maximizing the sum of the conditional probabilities of all node sequence paths having a specific start node and end node, which can be generated from a given directed graph by unfolding the graph to the length of the label probability sequence output by the neural network. The GTC training loss and gradients can be efficiently computed by a dynamic programming algorithm that is based on computing forward variables and backward variables and stitching the two together.
[0020] The GTC-based training of the neural network aims to update the trainable parameters of the neural network by optimizing the label predictions of the neural network such that the best overall predicted label sequence can be generated from the directed graph encoding the labeled information and the error for all possible label sequence predictions for a set of training samples and graph-based labeled information pairs is minimized. Examples of trainable parameters include the weights of the neurons of the neural network, hyperparameters, etc.
[0021] Additionally or alternatively, some embodiments are based on the recognition that the GTC objective function and the directed graph allow for the consideration of not only multiple label sequences but also different probabilities of multiple label sequences. This consideration is advantageous for the GTC objective function because it can adapt the supervised information to specific situations. To this end, in some embodiments, the directed graph is weighted with different weights for at least some of the edges or transitions. The weights of these transitions are used to compute the conditional probabilities of the label sequences.
[0022] Some embodiments are based on the recognition that GTC can be used to encode the N-best list of a pseudo-label sequence as a graph for semi-supervised learning. To this end, some embodiments disclose an extension of GTC to model the posterior of both labels and label transitions via a neural network, which can be applied to a wider range of tasks. The extended GTC (GTC-e) is used for multi-speaker speech recognition tasks. The transcription of multi-speaker speech and speaker information are represented by a graph, where the speaker information is associated with the ASR output that transforms and utilizes nodes. Using GTC-e, multi-speaker ASR modeling becomes very similar to single-speaker ASR modeling because the tokens of multiple speakers are recognized as a single merged sequence in chronological order.
[0023] Additionally, the method of training a neural network model using a loss function to learn the mapping of an input sequence to an output sequence that is typically shorter (such as CTC and Recurrent Neural Network Transducer (RNN-T)) is a commonly used loss function in automatic speech recognition (ASR) technology. CTC and RNN-T losses are designed for unaligned training of neural network models to learn the mapping of an input sequence (e.g., acoustic features) to an output label sequence that is typically shorter (e.g., words or sub-word units). While the CTC loss requires the neural network output to be conditionally independent, the RNN-T loss provides an extension to train a neural network whose output frames are conditionally dependent on previous output labels. To perform training without knowing the alignment between the input sequence and the output sequence, both loss types are marginalized over a set of all possible alignments. Such alignments are derived from the supervisory information (label sequence) by applying specific instructions that define how to extend the label sequence to match the length of the input sequence. In both cases, such instructions include using transformation rules specific to the loss type and additional blank labels.
[0024] However, changing the training grid of the transducer model to achieve a strictly monotonic alignment between the input sequence and the output sequence can leave other aspects of the RNN-T (such as emitting ASR labels within a single time frame) unchanged.
[0025] Some embodiments are based on the recognition of the GTC-Transducer (GTC-T) objective, which extends GTC to output a conditional-dependent neural network similar to RNN-T. In one embodiment, GTC-T allows the user to define label transitions in graph format and thus easily explore new grid structures for transducer-based ASR. In the embodiment, a CTC-like grid is used to train a GTC-T-based ASR system. Additionally, the GTC-T objective allows the use of different graph topologies to construct the training grid, e.g., a graph type corresponding to a CTC-like topology or a graph type corresponding to a single RNN-T (or RNA) loss type.
[0026] Accordingly, one embodiment discloses an end-to-end automatic speech recognition (ASR) system, comprising: a processor; and a memory storing instructions thereon. The processor is configured to execute the stored instructions to cause the ASR system to collect a sequence of acoustic frames providing a digital representation of an acoustic signal, the acoustic signal including a mixture of speech performed by multiple speakers. The processor is further configured to encode each frame of the sequence of acoustic frames using a multi-head encoder that encodes the likelihood of each frame into a transcription output and the likelihood of the identity of the speaker, to produce a sequence of likelihoods of the identity of the speaker corresponding to the sequence of acoustic frames and a sequence of likelihoods of the transcription output. The processor is further configured to decode the sequence of likelihoods of the transcription output and the sequence of likelihoods of the identity of the speaker using a decoder that performs alignment, to produce a sequence of transcription outputs annotated with the identity of the speaker. Additionally, the processor is configured to submit the sequence of transcription outputs annotated with the identity of the speaker to a downstream application.
[0027] Accordingly, one embodiment discloses a computer-implemented method for performing end-to-end ASR. The method includes the steps of: collecting a sequence of acoustic frames providing a digital representation of an acoustic signal, the acoustic signal including a mixture of speech performed by multiple speakers. The method further includes encoding each frame of the sequence of acoustic frames using a multi-head encoder that encodes the likelihood of each frame into a transcription output and the likelihood of the identity of the speaker, to produce a sequence of likelihoods of the transcription output corresponding to the sequence of acoustic frames and a sequence of likelihoods of the identity of the speaker. The method further includes: decoding the sequence of likelihoods of the transcription output and the sequence of likelihoods of the identity of the speaker using a decoder that performs alignment, to produce a sequence of transcription outputs annotated with the identity of the speaker. Additionally, the method includes submitting the sequence of transcription outputs annotated with the identity of the speaker to a downstream application. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1
[0029] Figure 1 is a block diagram showing an end-to-end automatic speech recognition (ASR) system according to an example embodiment.
[0030] Figure 2
[0031] Figure 2 shows a block diagram of internal components of an end-to-end ASR system according to an example embodiment. Figure 1 of
[0032] Figure 3
[0033] Figure 3 shows a Figure 1 Example architecture of an end-to-end ASR system.
[0034] Figure 4
[0035] Figure 4 Shows an extension of the GTC for an end-to-end ASR system for performing multi-speaker separation according to an example embodiment.
[0036] Figure 5
[0037] Figure 5 Shows the architecture of an end-to-end ASR system using a neural network trained on the GTC-e objective function according to an example embodiment. Figure 1 of the end-to-end ASR system.
[0038] Figure 6
[0039] Figure 6 Shows a working example of a neural network according to an example embodiment. Figure 5 of the neural network.
[0040] Figure 7A
[0041] Figure 7A Is a schematic diagram showing the workflow of training a neural network using a graph-based temporal classification (GTC) objective function according to an example embodiment.
[0042] Figure 7B
[0043] Figure 7B Shows a sequence of probability distributions output by a neural network according to an example embodiment.
[0044] Figure 7C
[0045] Figure 7C Shows an exemplary directed graph according to an example embodiment.
[0046] Figure 7D
[0047] Figure 7D Shows an example of a possible unconstrained repetition of labels during the unfolding of a directed graph according to an example embodiment.
[0048] Figure 7E
[0049] Figure 7E Shows an exemplary monotonic directed graph according to an example embodiment.
[0050] Figure 7F
[0051] Figure 7F Shows a monotonic directed graph modified based on constraints on label repetition according to an example embodiment.
[0052] Figure 8
[0053] Figure 8 Shows the steps of a method for training a neural network using a GTC objective function according to an example embodiment.
[0054] Figure 9
[0055] Figure 9 Shows a beam search algorithm used during the decoding operation of a neural network according to an example embodiment.
[0056] Figure 10
[0057] Figure 10 Shows Table 1 according to an example embodiment, which shows the greedy search results of the word error rate (WER) using the GTC-e objective function compared with other methods.
[0058] Figure 11
[0059] Figure 11 Shows Table 2 according to an example embodiment, which shows the greedy search results of the ASR performance of an ASR system based on the oracle token error rate based on the GTC-e objective function.
[0060] Figure 12
[0061] Figure 12 Shows Table 3 according to an example embodiment, which shows the beam search results of the ASR performance of an ASR system based on the WEB based on the GTC-e objective function.
[0062] Figure 13
[0063] Figure 13 Shows Table 4 according to an example embodiment, which shows the beam search results of the ASR performance of an ASR system based on the WER for multiple speakers based on the GTC-e objective function.
[0064] Figure 14A
[0065] Figure 14A Shows the neural network architecture of an ASR system implemented using the GTC-T objective function according to an example embodiment.
[0066] Figure 14B
[0067] Figure 14B Shows the pseudocode of a beam search algorithm for GTC-T with a similar CTC graph according to an example embodiment.
[0068] Figure 14C
[0069] Figure 14C Shows a comparison of ASR results of CTC, RNN-T, and GTC-T losses on the HKUST benchmark according to an example embodiment.
[0070] Figure 14D
[0071] Figure 14D Shows a comparison of ASR results of CTC, RNN-T, and GTC-T losses on the LibriSpeech dataset benchmark according to an example embodiment.
[0072] Figure 15
[0073] Figure 15 Shows a block diagram of a computer-based system trained using a GTC-e objective function according to an example embodiment. Detailed Description
[0074] In the following description, for the purpose of illustration, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, devices and methods are shown in block diagram form only in order to avoid obscuring the present disclosure.
[0075] As used in this specification and the claims, the terms "for example," "exemplify," and "such as" and the verbs "include," "have," "contain," and their other verb forms, when used in conjunction with a list of one or more components or other items, are each construed as open-ended, meaning that the list is not to be considered as excluding other additional components or items. The term "based on" means at least partially based on. Additionally, it should be understood that the language and terminology used herein are for the purpose of description and should not be regarded as limiting. Any headings used in this specification are for convenience only and have no legal or limiting effect.
[0076] In recent years, automatic speech recognition (ASR) has made significant progress, especially due to the exploration of neural network architectures that improve the robustness and generalization ability of ASR models. The growth of end-to-end ASR models features a simplified ASR architecture that has a single neural network with frameworks such as connectionist temporal classification (CTC), attention-based encoder-decoder models, and recurrent neural network transducer (RNN-T). Additionally, graph modeling has traditionally been used in ASR (such as using systems based on hidden Markov models (HMMs)), and weighted finite-state transducers (WFSTs) are used to combine several modules, including pronunciation dictionaries, context dependencies, and language models. Recently, through a new loss function called graph-based temporal classification (GTC), the use of graph representations in loss functions for training deep neural networks has also been proposed, which is a generalization of CTC for handling sequence-to-sequence problems. GTC can take graph-based supervision information as input to describe all possible alignments between the input sequence and the output sequence for learning the best possible alignment from the training data.
[0077] GTC is used to boost ASR performance via semi-supervised training by using the N-best list of ASR hypotheses converted to graph representations to train the ASR model using unlabeled data. However, in the original GTC, only the posterior probabilities of ASR labels are trained, and trainable label conversions are not considered.
[0078] Some embodiments are based on the recognition that extending GTC to handle label conversions will allow for modeling information about the labels. For example, in a multi-speaker speech recognition scenario where there is some overlap between the speech signals of multiple speakers, transition weights can be used to model speaker predictions aligned with frame-level ASR label predictions such that when an ASR label is predicted, it is also detected whether it belongs to a specific speaker.
[0079] Figure 1 FIG. 100 is a block diagram showing an end-to-end ASR system 104 according to an example embodiment. The end-to-end ASR system 104 includes a memory 105 storing instructions thereon. The instructions are executed by a processor 106 to cause the end-to-end ASR system 104 to perform some operations. The operations of the end-to-end ASR system 104 are described below in the form of various embodiments.
[0080] In one implementation, the end-to-end ASR system 104 is configured to collect a sequence of acoustic frames providing a digital representation of an acoustic signal that includes a mixture of speech performed by multiple speakers. For example, a first speaker 101 outputs a first speech signal, a second speaker 102 outputs a second speech signal, the first speech signal and the second speech signal overlap, and the overlapping speech 103 corresponding to the mixture of the speech of the first speaker 101 and the second speaker 102 is collected by the end-to-end ASR system 104. The end-to-end ASR system 104 includes an input interface that converts the overlapping speech into a digital representation of the acoustic signal corresponding to the sequence of frames in the overlapping speech 103.
[0081] The overlapping speech 103 thus corresponds to an input acoustic sequence that is processed by the end-to-end ASR system 104 to generate a sequence of transcription outputs 107 annotated with the identities of the speakers, and the sequence of transcription outputs is submitted to a downstream application. Each sequence of transcription outputs is a transcription of a discourse, or a portion of the discourse represented by the corresponding input acoustic signal. For example, the end-to-end ASR system 104 can obtain the overlapping speech 103 (also interchangeably referred to hereinafter as the acoustic signal) and generate a corresponding transcription output 107 that is a transcription of the discourse represented by the input acoustic signal 103 and is annotated with the speaker ID of at least one of the multiple speakers, such as the first speaker 101 or the second speaker 102.
[0082] The input acoustic signal 103 can include a sequence of multiple audio data frames that are a digital representation of a discourse, such as a continuous data stream. The sequence of multiple audio data frames can correspond to a sequence of time steps, for example, where each audio data frame is associated with a 25-millisecond audio stream data that is further shifted 10 milliseconds in time from a previous audio data frame. Each audio data frame in the sequence of multiple audio data frames can include feature values of a frame characterizing a portion of the discourse at the corresponding time step. For example, the sequence of multiple audio data frames can include filter bank spectral feature vectors.
[0083] The end-to-end ASR system 104 obtains the input acoustic sequence and processes the input acoustic sequence to generate a sequence of transcription outputs. Each sequence of transcription outputs is a transcription of a discourse, or a portion of the discourse represented by the corresponding input acoustic signal. For example, the end-to-end ASR system 104 can obtain the input acoustic signal 103 corresponding to a mixture of acoustic signals of multiple speakers, such as the first speaker 101 and the second speaker 102, and generate a corresponding transcription output 107 in chronological order, where the transcription output 107 is a transcription of the discourse represented by the multiple speakers through the input acoustic signal 103.
[0084] The transcription output 107 may include a sequence of transcribed segments of the utterance represented by the input acoustic signal 103. The transcription output may include one or more characters. For example, the transcription output may be a character or a sequence of characters from the Unicode character set. For example, the character set may include the alphabets of English, Asian languages, Cyrillic, and Arabic. The character set may also include Arabic numerals, space characters, and punctuation marks. Additionally or alternatively, the transcription output may include bits, words, and other language constructs.
[0085] To that end, the end-to-end ASR system 104 is configured to perform a series of operations, including an encoding operation, a decoding operation, and an output operation, which are shown by way of example in Figure 2 by way of example.
[0086] Figure 2 FIG. 200 is a block diagram showing internal components of an end-to-end ASR system 104 according to some embodiments of the present disclosure. The end-to-end ASR system 104 includes an encoder 201, a decoder 202, and an output generation module 203. The encoder 201, the decoder 202, and the output generation module 203 are examples of operations performed by the end-to-end ASR system 104 by a processor 106 executing stored computer instructions corresponding to each of these operations.
[0087] The encoder 201 is a multi-head encoder having one head corresponding to each of a plurality of speakers, such as a first speaker 101 and a second speaker 102. The encoder 201 is configured such that the end-to-end ASR system 104 encodes each frame in the input sequence of acoustic frames of the input acoustic signal 103 into a likelihood of a transcription output and a likelihood of the identity of the speaker using the multi-head encoder 201 that encodes each frame, to produce a sequence of likelihoods of the identity of the speaker and a sequence of likelihoods of the transcription output corresponding to the sequence of acoustic frames of the input acoustic signal 103.
[0088] In addition, the decoder 202 is configured to decode the sequence of likelihoods of the transcription output and the sequence of likelihoods of the identity of the speaker provided by the encoder 201. The decoder 202 is an alignment-based decoder for producing an aligned sequence of transcription outputs annotated with the identity of the speaker.
[0089] The sequence of transcription outputs annotated with the identity of the speaker is submitted by the output generation module 203 as the transcription output 107 to a downstream application. The downstream application may be an online streaming-based application, such as an online music providing application, an online video rendering application, a live sports event application, a live teleconference application, etc.
[0090] In an example, for the end-to-end ASR system 104, the encoder 201 is an acoustic encoder, and the decoder 202 is an attention-based decoder. The acoustic encoder processes the input acoustic signal 103 and generates a sequence of encoder states that provides an alternative (e.g., higher) representation for the input acoustic signal 103. The sequence of encoder states can include an alternative sequence of multiple audio data frames corresponding to a second set of time steps. In some embodiments, the alternative representation of the input acoustic sequence is downsampled to a lower frame rate, i.e., the second set of time steps in the alternative representation is less than the first set of time steps in the input acoustic sequence. The attention-based decoder is trained to process the encoder states representing the alternative representation of the input acoustic signal 103 and generate a transcription output from the sequence of encoder states provided to the attention-based decoder.
[0091] Some embodiments are based on the recognition that an attention-based ASR system may need to observe an entire speech utterance segmented by speech pauses in order to assign weights to each input frame for identifying each transcription output 203. Since there is no prior knowledge about which part of the input acoustic signal is relevant to identifying the next transcription output and the need to assign weights to each encoder state, an attention-based decoder typically needs to process large input sequences. Such processing allows for leveraging attention on different parts of the utterance, but also increases the output latency and is thus not practical for streaming / online speech recognition.
[0092] Some embodiments are based on the recognition that an example of prior knowledge about the relevance of different parts of the input sequence to identifying the next transcription output is an indication of the position of the frames corresponding to the transcription output to be identified in the input sequence. In practice, if the transcription output position is known, the attention-based decoder can be forced to place more attention on these positions and less or no attention on other positions by restricting the input sequence. In this way, for each transcription output, the attention-based network can focus its attention on its position in the input sequence. This guided attention reduces the need to process large input sequences, which in turn reduces the output latency, making the attention-based decoder practical for streaming / online recognition.
[0093] To this end, decoder 202 is an alignment decoder that is trained to determine the position of the encoder state in the encoded state sequence of the encoded transcription output (such as characters, bits, words, etc.). For example, Connectionist Temporal Classification (CTC) is an objective function and associated neural network output for training a Recurrent Neural Network (RNN) (such as a Long Short-Term Memory (LSTM) network) to solve the problem of time-varying sequences. The CTC-based ASR system is an alternative to the attention-based ASR system. The CTC-based neural network generates an output for each frame of the input sequence, that is, the input and output are synchronized, and before collapsing the neural network output to the output transcription, a beam search algorithm is used to find the best output sequence. The performance of the attention-based ASR system can be better than that of the CTC-based ASR system. However, some embodiments are based on the recognition that the input and output frame alignment used by the intermediate operations of the CTC-based ASR system can be used by the attention-based ASR system to address its above-mentioned output latency drawback.
[0094] Figure 3 An example architecture of such a CTC-based ASR system 300 is shown, where encoder 201 is a self-attention encoder 301. The CTC-based ASR system 300 also includes an attention-based decoder 303.
[0095] Encoder 301 processes the input acoustic signal 103 and generates an encoder state sequence 302, thereby providing an alternative (e.g., higher) representation for the input acoustic signal 103. The encoder state sequence 302 may include an alternative sequence of multiple audio data frames corresponding to a second set of time steps. The attention-based decoder 303 is trained to process the encoder state sequence 302 representing the alternative representation of the input acoustic signal 103 and generate a transcription output 304 (corresponding to output 203) from the encoder state sequence provided to the attention-based decoder 303.
[0096] The CTC-based ASR system 300 further includes a decoder 202, which is an alignment decoder 305 that utilizes alignment information 306. The alignment information 306 includes an alignment of the transcription output sequence annotated with the identities of the speakers in the multi-speaker input acoustic signal 103. Such a CTC-based ASR system 300 includes a partitioning module 307 configured to partition the encoder state sequence 302 into a set of partitions 308. For example, the partitioning module 307 can partition the encoder state sequence at each position 306 of the identified encoder states such that the number of partitions 308 is defined by (e.g., equal to) the number of identified encoder states 302 encoding the transcription output. In this way, the attention-based decoder 303 receives as input not the entire sequence 302 but a portion 308 of the sequence, and each portion may include a new transcription output to form the transcription output sequence 304. In some implementations, the combination of the alignment decoder 305, the attention-based decoder 303, and the partitioning module 307 is referred to as a triggered attention decoder. In effect, the triggered attention decoder can process a portion of the utterance when it is received, making the CTC-based ASR system 300 practical for streaming / online mode recognition.
[0097] In some existing end-to-end ASR systems, it is assumed that the label sequences of different speakers are output at different output heads, or the prediction of the sequence of a speaker can only start when the sequence of the previous speaker is completed.
[0098] However, in the end-to-end ASR system 104 disclosed in various embodiments provided herein, the multi-speaker ASR problem is not implicitly regarded as a source separation problem using separate output layers for each speaker or a cascaded process of identifying each speaker one by one. Instead, the prediction of the ASR labels of multiple speakers is regarded as a sequence of acoustic events, independent of the source.
[0099] For this purpose, some embodiments use a generalized form of CTC previously disclosed at GTC and use an extended GTC (GTC-e) loss to achieve multi-speaker separation.
[0100] Figure 4 FIG. 400 is a schematic diagram showing an extension of the GTC 401 objective function of an end-to-end ASR system 104 for performing multi-speaker separation according to some embodiments.
[0101] The GTC 401 objective function is extended to provide the GTC-e 402 loss, which allows training of two separate predictions for an end-to-end ASR system 104 aligned at the frame level, one prediction for speakers (such as speakers 101 and 102) and one prediction for the ASR output (such as output 203). To efficiently utilize the speaker prediction during decoding, the existing frame-synchronous beam search algorithm of GTC 401 is adapted to GTC-e 402.
[0102] The GTC 401 objective function provides an output in the form of a directed graph 403, in which nodes represent labels and edges represent label transitions. On the other hand, the GTC-e 402 objective function provides a directed graph 404 as output, in which nodes represent tokens and edges represent speaker identifiers (IDs). Thus, the GTC-e 402 objective function is configured to perform multi-speaker ASR by treating the ASR outputs of multiple speakers as a sequence of mixed events with a chronologically meaningful ordering.
[0103] To this end, the GTC-e 402 objective function is used as the loss function of a neural network, which is trained to receive an input sequence of labels corresponding to multiple speakers and provide, as output, the labels and speaker identifiers separated chronologically for each label.
[0104] Figure 5 The architecture of an end-to-end ASR system 104 using a neural network 501 trained on the GTC-e 402 objective function is shown. The neural network 501 includes an encoder 201 and a decoder 202 described in Figure 2 . The neural network 501 is trained to achieve multiple objectives of speech recognition and speaker identification.
[0105] In various embodiments, the encoder 201 is a multi-head encoder and the decoder 202 is a time-aligned decoder (as Figure 3 shown). The multi-head encoder and the time-aligned decoder are parts of the neural network 501 that are end-to-end trained to recognize and / or transcribe the speech of each speaker. To this end, the neural network 501 is trained to achieve multiple objectives, namely, speech recognition and speech identification. To achieve this training, in some implementations, multiple loss functions (one loss function for speech recognition and another loss function for speaker identification) are used to train the neural network 501. Doing so allows simplifying the construction of loss functions designed for different applications and / or reusing traditional loss functions.
[0106] To this end, the neural network 501 is trained to minimize a loss function that includes a first component associated with the error in speech recognition and a second component associated with the error in speaker identification.
[0107] However, using multiple loss functions or multiple components of a loss function can create synchronization issues in the outputs of different heads of a multi-head encoder. This is because there is no alignment information between the acoustic frames of the input acoustic signal 103 and the labels, and thus aligning each transcription and each speaker ID information separately will result in inconsistent alignments. To this end, some embodiments use a single loss function configured to simultaneously minimize errors in speech recognition and speaker identification to train the multi-head encoder 201 of the neural network 501.
[0108] Some embodiments are based on the recognition that the CTC objective, which allows the decoder 202 to enforce alignment, can be used to train the end-to-end ASR system 104. For example, in speech audio, there can be multiple time slices corresponding to a single phone. Since the alignment of the observed sequence with the target labels is unknown, training using the CTC objective predicts the probability distribution at each time step.
[0109] When there is no available temporal alignment information between the longer label probability sequence output by the neural network 501 and the training label sequence, the CTC objective uses a graph-based loss function to train the neural network 501, and the longer label probability sequence output by the neural network 501 is computed from the observed sequence input to the neural network 501. This missing temporal alignment information creates a temporal ambiguity between the label probability sequence output by the neural network 501 and the supervisory information used for training, which is the label training sequence that can be resolved using the CTC objective function.
[0110] However, the CTC objective function is only suitable for resolving temporal ambiguity during the training of the neural network. If other types of ambiguity need to be considered, the CTC objective function will fail. Therefore, the aim of some embodiments is to enhance the CTC objective function to account for other ambiguities, such as speaker identification.
[0111] Some embodiments are based on the recognition that although the definition of the CTC objective and / or CTC rules is not graph-based, the problems or limitations of the CTC objective can be illustrated by a directed graph and solved using a graph-based definition. Specifically, if the CTC rules cause the supervisory information of the training label sequence to reside on a graph that enforces the alignment between the label probability sequence generated by the neural network and the training label sequence, it would be advantageous to extend the principle of this graph to solve speaker alignment.
[0112] In an example, an extended CTC objective function is used to train the neural network 501. As is well known, GTC is a generalized form of the CTC objective function. Thus, in one implementation, the neural network 501 is trained to use the GTC-e402 objective function, also referred to as the GTC-e 402 loss function. The GTC-e 402 objective function (or extended CTC objective function) is used to enforce an alignment between the input and output on the graph using nodes that indicate speech identity outputs, which are also referred to as transcription outputs. The edges of the graph indicate transitions between multiple speakers. Such a graph is shown in Figure 6 as follows.
[0113] Figure 6 FIG. 600 shows a working example of the neural network 501 according to an example implementation. The working example 600 shows a graph 602 having a plurality of nodes and edges. A node such as node 603 is depicted with the text "Hello", and an edge 604 is depicted with the text "s1". In FIG. 602, each node represents a label, and an edge connecting two nodes represents the possibility of a transition between those two nodes. Some implementations are based on the understanding that one way to resolve speaker ambiguity is to annotate the nodes and edges not only with labels but also with the identity of the speaker. Thus, in FIG. 602, nodes (such as node 603, node 605, etc.) are associated with labels indicating ASR outputs. For example, node 603 is associated with the label "Hello", node 605 indicates a start node, and edge 604 indicates a speaker with identity s1, and edge 606 indicates a speaker with identity s2. Similarly, other nodes and edges in FIG. 602 are annotated, although not all annotations are shown for brevity and do not limit the scope of the present disclosure.
[0114] Additionally or alternatively, some implementations are based on the understanding that in FIG. 602, for each ASR output, a speaker label is predicted at the frame level in the form of a label on the node and in the form of an annotation on the edge. The speaker information can be considered as a transition probability in FIG. 602, and such annotation allows synchronization of the speaker and the ASR label prediction at the frame level.
[0115] As Figure 6 shown, the neural network 501 receives a multi-speaker overlapping speech input acoustic signal 103. For brevity, the overlapping speech input acoustic signal 103 is from two speakers s1 and s2 (in Figure 1Overlapping speech formation by, respectively, a first speaker 101 and a second speaker 102 (shown in the figure). Speaker s1 has the utterance "Hello, cat", and speaker s2 has the utterance "Hi, dog". The neural network 501 uses the encoder 201 and the decoder 202 and processes the overlapping speech input acoustic signal 103 based on the extended CTC objective function and the GTC-e402 objective function. As a result of the processing, Figure 602 is obtained, where the nodes of Figure 602 indicate the transcription outputs corresponding to the utterances "Hello", "Hi", "cat", and "dog" in chronological order, and the edges give the corresponding speaker IDs s1, s2, s1, and s2 in chronological order. The transcription output 107 from the neural network 501 thus includes both the synchronized label output 107a and the speaker identification output 107b. This synchronization is performed for each frame of the input acoustic signal 103.
[0116] In an embodiment, the GTC-e 402 objective function uses the supervision information from a directed graph of nodes connected by edges representing labels and transitions between labels, where the directed graph represents the sequence of probability distributions output by the neural network 501 and the possible alignment paths of the labels. The explanation of the GTC-e 402 objective function is covered in the following description.
[0117] To understand the principle of the GTC-e 402 objective function, it is first necessary to understand the principle of the GTC objective function.
[0118] Figure 7A is a schematic diagram showing the workflow of training a neural network 701 using a graph-based temporal classification (GTC) objective function 702 according to an exemplary embodiment. The neural network 701 is trained to output a sequence of probability distributions 703 of the observation sequence 705, where the sequence of probability distributions 703 represents the label probabilities at each moment. The type of the observation sequence 705 input to the neural network 701 and the multiple label sequences 706a depend on the type of application using the neural network 701.
[0119] For example, for a neural network 701 associated with an ASR system, an observation sequence 705 provided at an input interface of the neural network 701 is associated with a speech utterance, and multiple tag sequences 706a can correspond to words, sub-words, and / or characters from an alphabet of a particular language. Additionally, in an acoustic event detection application where the neural network 701 can be trained to detect different acoustic events occurring in a particular time span in an acoustic scene, the observation sequence 705 can include different audio features of sounds included in the particular time span in the acoustic scene. In such a case, the multiple tag sequences 706a can include tags corresponding to different entities that produce the sounds or cause the acoustic events. For example, for a meowing sound in an acoustic scene, the tag "cat sound" can be used, and similarly, for a barking sound of a dog, the tag "dog sound" can be used. Thus, the observation sequence 705 and the multiple tagging sequences 706a vary according to the application.
[0120] The neural network 701 is trained using a GTC objective function 702, where the GTC objective function 702 uses supervision information from a directed graph 704. The directed graph 704 includes multiple nodes connected by edges, where the edges represent tags and transitions between the tags. Some embodiments are based on the recognition that presenting the supervision information on the directed graph 704 allows different rules to be applied to train the neural network in a manner consistent with the principles of such training. This is because the structure of the directed graph 704 is consistent with the differentiable methods used by the forward-backward algorithm for training. Thus, if the rules to be imposed on the training are represented as part of the structure of the directed graph 704, such rules can be imposed on the training in a differentiable manner consistent with the forward-backward algorithm.
[0121] For example, in one embodiment, the directed graph 704 represents multiple possible alignment paths between a probability distribution sequence 703 and multiple tag sequences 706a. Such a directed graph allows the neural network 701 to be trained using the GTC objective to perform alignment between its input and output in both the time domain and the tag domain. To achieve such multi-alignment, the structure of the directed graph 704 is non-monotonic, i.e., it specifies a non-monotonic alignment between a tag sequence in the multiple tag sequences 706a and the probability distribution sequence 703.
[0122] Additionally or alternatively, in one embodiment, the directed graph 704 represents a constraint 706b on tag repetition. The constraint 706b on tag repetition specifies a minimum number of repetitions of a tag, a maximum number of repetitions of a tag, or both. The constraint 706b on tag repetition can reduce the number of possible sequences of tags that can be generated during the unfolding of the directed graph 704 for time alignment and accelerate the calculation of the GTC loss.
[0123] The observation sequence 705 can correspond to features extracted by a feature extraction method. For example, the observations can be obtained by partitioning the input signal into overlapping blocks and extracting features from each block. The type of features extracted can vary depending on the type of input. For example, for a speech utterance, the features extracted from a sequence of chunks of an audio sample can include a spectral decomposition of the input signal and additional signal processing steps to mimic the frequency resolution of the human ear. For example, each feature frame extracted from the input speech utterance can correspond to a moment in the observation sequence 705, e.g., where each frame of the speech utterance is associated with a 25 millisecond audio sample that is further shifted in time by 10 milliseconds from a previous frame of the speech utterance. Each feature frame of the speech utterance in the sequence of feature frames of the speech utterance can include acoustic information characterizing the portion of the utterance at the corresponding time step. For example, the sequence of feature frames of the audio data can include filter bank spectral energy vectors.
[0124] Input and Output of Neural Network
[0125] In various embodiments, the input to the neural network 701 is the observation sequence 705, and the output of the neural network 701 is a sequence of probability distributions 703 (also referred to as likelihoods) over a set of labels. For clarity of explanation, the probability distributions 703 generated by the neural network 701 are explained below using an exemplary embodiment, where the neural network 701 is trained for automatic speech recognition (ASR). However, this example is not intended to limit the scope, applicability, or configuration of the embodiments of the present disclosure.
[0126] Figure 7B A sequence of probability distributions 703 computed by a neural network 701 trained for ASR from multiple observation sequences 705 is shown in accordance with an example embodiment. In conjunction with Figure 7A Explanation Figure 7B .. The input to the neural network 701 includes an observation sequence 705 having features extracted from a speech utterance. The neural network 701 is trained based on supervision information including a directed graph 704 that encodes possible speech recognitions with some ambiguity.
[0127] The directed graph 704 and the sequence of probability distributions 703 are processed by the GTC objective function 702 to optimize the timing and label alignment of the labels for the input observation sequence in the directed graph 704 and to determine the gradients for updating the parameters of the neural network 701. The neural network 701 trained using the GTC objective function 702 produces a matrix of the probability sequence 703, where the columns correspond to time steps and each row corresponds to a label (here, the letters of the English alphabet).
[0128] In Figure 7BIn the example, the neural network 701 outputs a D×T-dimensional matrix (where D represents the label dimension and T represents the time dimension, and in the given example, D = 29 and T = 30) or a sequence of probability distributions 703, where the letters of the English alphabet and some special characters correspond to D = 29 labels. Each column (D-dimensional) in the D×T matrix corresponds to probabilities that sum to 1, i.e., the matrix represents the probability distribution over all labels at each time step. In this example, the labels correspond to the characters from the English alphabet A-Z plus the additional symbols "_", ">", and "-", where "-" represents a blank token or blank symbol. The sequence of probability distributions 703 defines the probabilities of different labels at each time step, which are computed by the neural network 701 from the observation sequence 705. For example, as Figure 7B shown, the probability of observing the label "B" at the fourth time step is 96%, the probability of the label "O" is 3%, and the probabilities of the remaining labels are close to zero. Thus, the most likely label sequence in the output of this example will have the letter "B" or "O" at the fourth time position. At inference time, various techniques (such as prefix beam search) can be used to extract the final label sequence from the sequence of probability distributions 703 over the labels.
[0129] Furthermore, by using the GTC objective, the neural network 701 is trained to maximize the probability of the label sequence corresponding to the sequence of nodes and edges included in the directed graph 704 in the sequence of probability distributions 703. For example, assume that the true transcription of the input speech utterance corresponds to "BUGS_BUNNY", however, the true transcription is unknown. In this case, the directed graph 704 can be generated from the list of ASR hypotheses of the speech utterance corresponding to "BUGS_BUNNY". For example, the list of ASR hypotheses represented by the directed graph 704 can be "BOX_BUNNY", "BUGS_BUNNI", "BOG_BUNNY", etc. (where each letter of the English alphabet corresponds to a label). Since it is unknown whether any hypothesis is correct or what part of the hypothesis is correct, such a list of multiple hypotheses of the speech utterance corresponding to "BUGS_BUNNY" contains ambiguous label information, different from the true information of only "BUGS_BUNNY".
[0130] During GTC training, the directed graph 704 will be unfolded to the length of the sequence of probability distributions 703, where each path in the unfolded graph from a particular start node to a particular end node represents an alignment path and a label sequence. Such a graph may include a non-monotonic alignment between the sequence of probability distributions 703 output by the neural network 701 and the label sequence 706a encoded in the graph. One of the alignment paths included by the directed graph 704 may correspond to label sequences such as: "-BOOXXX_BBUN-NI", "B-OOX--BUNN-NY-", "BU-GS-_-BUN-Y-" etc. (where "-" represents a blank symbol). Each label sequence in the directed graph 704 includes a time alignment and a label alignment. By processing the directed graph 704 and training the neural network 701, the GTC objective function 702 optimizes the time and label alignments of the labels in the directed graph 704 and the sequence of probability distributions 703. The GTC objective function 702 is used to train the neural network 701 to maximize the probability of the label sequences included by the directed graph 704. Transition weights residing on the edges of the directed graph 704 can be used during training to emphasize more likely alignment paths. To this end, in an example embodiment, each hypothesis may be provided with a score by the neural network 701. Additionally, each hypothesis can be ranked based on the score. Further, based on the ranking, weights can be assigned to the transitions corresponding to each hypothesis such that the weight of the transition corresponding to the top-ranked hypothesis is greater than the weights of the transitions corresponding to the subsequent hypotheses in the N-best hypotheses. For example, based on context information, the hypothesis "BOG" may have a higher ranking compared to another hypothesis "BOX". Thus, the weight connecting the labels "O" and "G" can be greater than the weight of the connection between "O" and "X". Thus, the label sequence with a higher transition weight will be assigned a higher probability score and will thus be selected to correctly transcribe the input speech utterance.
[0131] Directed graph with non-monotonic alignment
[0132] In some embodiments, the supervision information is included by the structure of the directed graph 704, where the supervision information is used by the GTC objective function 702 to resolve one or more ambiguities (such as time and label ambiguities) to train the neural network 701. Thus, the supervision information specifies one or a combination of the non-monotonic alignments between the multiple label sequences 706a and the sequence of probability distributions 703. Based on the non-monotonic alignment, the directed graph 704 can output multiple unique label sequences.
[0133] Figure 7CIllustrates an exemplary directed graph 700c according to an example embodiment. The directed graph 700c includes a plurality of nodes 707a, 707b, 707c, and 707d, where each node represents a label. For example, node 707a represents the label "A", node 707b represents the label "B", node 707c represents the label "C", and node 707d represents the label "D". The directed graph 700c starts with a start node 711a and ends with an end node 711b. In Figure 7C it, the start node and the end node are connected to labels with dashed lines to illustrate that there may be other nodes in the directed graph 700c that are not shown for simplicity and clarity of illustration.
[0134] The directed graph 700c is a non-monotonic directed graph, thus providing a non-monotonic alignment between the label sequence of the directed graph 700c and the sequence of probability distributions 703 output by the neural network 705 during training. In different embodiments, the non-monotonic alignment can be implemented differently to capture label and temporal ambiguity through multiple paths of the nodes of the directed graph 700c.
[0135] For example, as Figure 7C shown, the non-monotonic alignment in the directed graph 700c can be constructed by connecting at least one node to different nodes representing different labels. For example, the node 707a representing the label A is connected to the node 707b representing the label B by an edge 709ab, and is also connected to the node 707c representing the label C by an edge 709ac. This separate connection allows the creation of multiple different label sequences defined by multiple different paths through the graph, such as the ACD sequence and the ABD sequence between the start node and the end node.
[0136] Another example of the non-monotonic alignment encoded in the structure of the directed graph 700c is a loop formed by edges connecting multiple non-blank nodes. In the directed graph 700c, the loop is formed by the edges 709ab and 709ba that allow multiple paths through the graph, such as ABACD or ABABD.
[0137] Some embodiments are based on the recognition that since the non-monotonic directed graph 700c encodes different label sequences, not all sequences are equally likely. Therefore, it is necessary to impose unequal probabilities on the structure of the directed graph 700c.
[0138] Another advantage of the directed graph 700c is its ability to encode the probability of a transition as the weight of an edge (which in turn encodes the probability of different paths). To this end, at least some of the edges in the non-monotonic directed graph 700c are associated with different weights (w), thus making the directed graph 700c a weighted directed graph 700c. For example, the edge 709ab can be weighted with a weight w 2 and the edge 709ba can be weighted with a weight w1 Weighted, edge 709bd can have a weight w 3 Weighted, edge 709ac can have a weight w 4 Weighted, edge 709cd can have a weight w 5 Weighted. Additionally, based on the weights, the conditional probability of the node sequence can be changed. For example, if the weight w 2 is greater than the weight w 1 , then in a specific sequence of nodes, the conditional probability of transitioning from node 707a to node 707b is greater than the conditional probability of transitioning from node 707b to node 707a.
[0139] A directed graph with a constraint on label repetition
[0140] Figure 7D Shows the repetition of labels during the expansion of the directed graph 700d according to an example embodiment. Figure 7D Includes the directed graph 700d on the left and the expanded directed graph 710d on the right. The directed graph 700d includes a label sequence corresponding to the transcription "HELLO WORLD". Assume that there are more observations in the observation sequence 705 provided to the neural network 701 than there are labels in the labeled sequence (i.e., the transcription). For example, the number of letters in the transcription "HELLO WORLD" is 10, and the number of observations (and corresponding conditional probabilities) can be 30. Therefore, in order to match or align the number of labels with the number of observations, some labels in the transcription are repeated during the expansion of the graph. For example, the letter "E" in the transcription "HELLO WORLD" can be repeated multiple times.
[0141] However, due to the lack of a constraint on the number of times a label can be repeated, it results in a waste of unnecessary computational power because the GTC objective function is required to analyze the possible transitions from each repeated label. To this end, the directed graph 700d includes a constraint 706b on label repetition. The constraint 706b in the directed graph 700d can include the minimum number of times a label is allowed to repeat in the label sequence or the maximum number of times a label is allowed to repeat in the label sequence, or both. This is because it is unlikely to observe the letter "E" in so many consecutive time frames, as in the exemplary expansion 712.
[0142] Therefore, as an addition or alternative to the non-monotonic alignment of the directed graph 700d, some embodiments use the structure of the directed graph 700d to impose a constraint on label repetition during training, which specifies the minimum number of repetitions of a label, the maximum number of repetitions of a label, or both. Such a constraint on label repetition of the nodes representing the labels can be achieved by removing the self-transitions of the nodes and adding transitions from the nodes to other nodes representing the same label.
[0143] Figure 7EAn exemplary directed graph 700e with a constraint 706b on label repetition according to an example embodiment is shown. The directed graph 700e starts with a start node 713a and ends with an end node 713b. The monotonic directed graph 700e includes a plurality of nodes 714x, 715y, 714y, and 714z, where each node represents a label. For example, node 714x represents the label "X", node 714y represents the label "Y", node 714z represents the label "Z", and node 715y represents another label "Y". In this example, a sequence of connected nodes representing the same label is formed by nodes 714y and 715y.
[0144] The directed graph 700e is monotonic because although there are multiple paths through the graph connecting the start node and the end node, after the folding process, only a single label sequence XYZ can be formed.
[0145] For example, the monotonic directed graph 700e can specify different label sequences during the unfolding of the monotonic directed graph 700e, such as X→X→X→Y→Z→Z→ or X→Y→Y→Z or X→Y→Z. However, after folding these label sequences, only one label sequence, i.e., X→Y→Z, is generated. In some embodiments, multiple monotonic directed graphs can be combined to form a non-monotonic directed graph (such as non-monotonic directed graph 700c), which is used to train the neural network 701.
[0146] In addition, in the monotonic directed graph 700e, it can be defined that a specific label (e.g., the label "Y") should not repeat more than twice, and the labels "X" and "Z" can repeat multiple times. This information is encoded in the structure of the graph and used in an automated manner during unfolding. For example, nodes 714x and 714z have self-transitions and can therefore repeat any number of times allowed by the unfolding. In contrast, nodes 714y and 715y corresponding to the label "Y" do not have self-transitions. Therefore, in order to travel through the graph between the start node and the end node, the path can be 714x - 714y - 714z (where the label "Y" corresponding to node 714y is repeated once) or 714x - 714y - 715y - 714z (where the label "Y" corresponding to nodes 714y and 715y is repeated twice). In addition, the directed graph 700e allows the repetition of other labels (such as the labels "X" and "Z" that are currently repeating multiple times without any constraints) to be modified or constrained. The directed graph 700e can be modified to a directed graph 700f to impose constraints on the other labels "X" and "Z".
[0147] Figure 7F Another exemplary directed graph 700f with a constraint 706b on label repetition according to an example embodiment is shown. In Figure 7F, the structure of the monotone directed graph 700f imposes the following constraint: the label "X" can only be repeated three times in sequence, for which reason, the node 716x representing the label "X" and the node 718x also representing the label "X" can be connected to the original node 714x. In this example, the sequence of connected nodes representing the same label is formed by nodes 714x, 716x, and 718x.
[0148] In a similar manner, label "Z" can be constrained to always repeat twice, and so on. To this end, node 717z can be connected to the original node 714z. In this way, directed graph 700f provides great flexibility to optimize the training of neural network 701.
[0149] Constraints 706b on repetitions are advantageous for speech-related applications. For example, for a directed graph 700f to be used by a neural network 701 corresponding to an ASR system configured to transcribe in English, it may be known in advance that an output corresponding to a label "U" is unlikely to be observed over multiple consecutive frames. Therefore, the label "U" may be constrained to be repeated only a limited number of times in order to reduce computational complexity and speed up computation of the GTC objective.
[0150] The advantages of the constraint 706b on repetition are not limited to speech-related applications. For example, the directed graph 700f and the neural network 701 may correspond to an acoustic event detection system implemented to detect acoustic events in a home environment. A short event like "door bang" is unlikely to occur over many consecutive observation frames. Therefore, the structure of the directed graph 700f may define a constraint 706b on the repetition of the label "door bang".
[0151] Training with the GTC objective using a directed graph
[0152] In various embodiments, the neural network 701 is trained based on the GTC objective function 702 to transform the observation sequence 705 into the probability distribution sequence 703. In addition, the neural network 701 is configured to expand the directed graph 704 to generate all possible label sequences from multiple label sequences 706a, so that the length of the label sequence matches the length of the probability distribution sequence 703. Expanding the directed graph 704 includes generating a label sequence and an alignment path according to the structure of the directed graph 704 by finding a path from the start node to the end node on the nodes and edges of the directed graph 704 whose length is the length of the probability distribution sequence 703. Each path in the expanded graph corresponds to a sequence of nodes and edges of a fixed length starting from a specific start node and ending at a specific end node. Each possible path corresponding to the sequence of nodes and edges in the expanded graph can be mapped to a label sequence.
[0153] In addition, the neural network 701 updates one or more parameters of the neural network 701 based on the GTC objective function 702, which is configured to maximize the sum of the conditional probabilities of all possible label sequences 706a generated by unfolding the directed graph 704. One or more parameters of the neural network 701 updated by the neural network 701 may include neural network weights and biases, as well as other trainable parameters (such as, embedding vectors, etc.).
[0154] In some embodiments, the directed graph 704 is a weighted graph having at least some edges associated with different weights. In addition, the GTC objective function 702 is configured to learn time alignment and label alignment to obtain an optimal pseudo-label sequence from the weighted directed graph 704, such that the training of the neural network 701 using the GTC function 702 updates the neural network 701 to reduce the loss with respect to the optimal pseudo-label sequence. The neural network 701 trained using the GTC objective function 702 transforms the observation sequence 705 into a sequence of probability distributions 703 over all possible labels at each time step. In addition, the trained neural network 701 maximizes the probability of the label sequence at the output of the neural network 701, where the label sequence corresponds to a sequence of nodes and edges present in the directed graph 704.
[0155] Thus, the GTC objective function 702 enables the neural network 701 to utilize label information in graph format to learn and update the parameters of the neural network 701.
[0156] The directed graph 704 provides the supervision information used by the GTC objective function 702 during the training of the neural network 701. In the directed graph 704, the label sequence is represented by multiple nodes and edges. In addition, the directed graph 704 may include a non-monotonic alignment between the probability distribution sequence 703 and the multiple label sequences 706a represented by the directed graph 704. The non-monotonic alignment or monotonic alignment is defined as the number of label sequences that can be generated from the directed graph 704 by transitioning from a specific start node to a specific end node after removing label repetitions and blank labels. The non-monotonic alignment allows the directed graph 704 to output multiple unique label sequences, while a monotonic graph would only allow the output of a single label sequence.
[0157] Due to the non-monotonic alignment feature, the directed graph 704 includes information associated not only with the changes in the label sequence in the time domain but also with the changes in the label sequence within the label domain itself. Due to the changes in the label sequence within the label domain, the directed graph 704 includes multiple paths passing through the multiple nodes and edges of the directed graph 704, where each path corresponds to at least one of the multiple label sequences 706a. Therefore, each edge in the directed graph 704 has a direction from one node towards another node.
[0158] Thus, the unaligned features allow the directed graph 704 to consider different label sequences during training, which allows the use of ambiguous label information to train the neural network 701 to account for the uncertainty about the correct transcription of the training samples.
[0159] In addition, the directed graph 704 allows at least one label in the label sequence to be repeated a specific minimum number of times and a specific maximum number of times during the unfolding of the directed graph 704, in order to reduce the number of possible label paths that can be generated from the unfolded graph and to accelerate the calculation of the GTC loss.
[0160] In some embodiments, the non-monotonic directed graph 704 is a weighted graph having at least some edges associated with different weights. Further, based on the weights of the corresponding edges in the directed graph 704, the conditional probability of a sequence of nodes can be calculated during training.
[0161] For ease of explanation, the GTC objective function is explained here with respect to a neural network corresponding to an ASR system. Consider a feature sequence X of length T′ derived from a speech utterance, which is processed by the neural network 701 to output a sequence of posterior distributions Y = (y 1 , …, y T ) of length T, which may differ from T′ due to downsampling, where y t represents the vector of posterior probabilities at time t, and represents the posterior probability of output symbol k at time t. For GTC, the label information used for training is represented by the graph , where the graph corresponds to the directed graph 704. The GTC objective function 702 marginalizes over all possible node sequences that can be obtained from the graph , which includes all valid node patterns as well as all valid time alignment paths. Thus, the conditional probability given the graph is defined by the sum over all node sequences in , which can be written as:
[0162]
[0163] where, represents the search function that unfolds into all possible node sequences of length T (not counting non-emitting start and end nodes), π represents a single node sequence and alignment path, and p(π|X) is the posterior probability of the path π given the feature sequence X. The posterior probability is used to calculate the conditional probability of the path π. The calculation of the conditional probability is explained in detail later.
[0164] We introduce several more notations that will be useful for deriving . Using g = 0, …, G+1 for the graph The nodes are indexed, and they are sorted in breadth - first search order from 0 (non - emitting start node) to G + 1 (non - emitting end node). In addition, the output symbol observed at node g is denoted by l(g), and the transition weight on the edge (g, g′) (which connects node g to node g′) is denoted by W (g,g′) denoted. Finally, by π t:t′ =(π t ,..., π t′ ) represents the subsequence of nodes of π from time index t to time index t′. In addition, π 0 and π T+1 correspond to the non - emitting start node 0 and the non - emitting end node G + 1.
[0165] To efficiently compute the conditional probability of a given graph compute the forward variable α and the backward variable β, and compute the conditional probability based on α and β. To this end, GTC uses the following equations to compute the forward probabilities (or forward variables) for g = 1,..., G.
[0166]
[0167] where denotes the sub - graph that starts at node 0 and ends at node g. The sum is taken over all possible π of its subsequence up to time index t that can be generated in t steps according to the sub - graph . In addition, the backward variable β is computed similarly using the following equation:
[0168]
[0169] where denotes the sub - graph that starts at node G and ends at node G + 1. By using the forward variable and the backward variable, the probability function
[0170]
[0171] For gradient - descent training, the loss function
[0172] (5)
[0173] must be differentiated with respect to the network output. For any symbol (where denotes the set of all possible output symbols or labels), it can be written as:
[0174]
[0175] Because is proportional to so
[0176]
[0177] And according to (4), the following equation can be derived:
[0178]
[0179] wherein represents the set of nodes in where the symbol k is observed.
[0180] In order to backpropagate the gradient through the softmax function, before applying softmax, the derivative with respect to the unnormalized network output is required, and this derivative is
[0181]
[0182] By substituting (8) and the derivative of the softmax function into (9), equation (10) is obtained
[0183]
[0184] wherein, the following fact is used
[0185]
[0186] The GTC objective function 702 learns the time and label alignment from the supervision information of the directed graph , and the GTC objective function 702 is used to train the neural network 701. The following explains Figure 8 the training.
[0187] The neural network 701 is trained using the GTC objective function 702, and the GTC objective function 702 enables the neural network 801 to solve the time alignment or time ambiguity and the label alignment or label ambiguity, so as to learn the best alignment between the probability distribution sequence 703 and the label sequence represented by the directed graph 704.
[0188] Figure 8 shows the steps of the method 800 for training the neural network 701 using the GTC objective function 702 according to an exemplary embodiment. In combination with Figure 7A explain Figure 8 . In Figure 8In it, at step 801, the output of neural network 701 for the given observation sequence X is calculated to obtain the posterior probability of any output symbol k at time t represented by At step 803, the directed graph
[0189] can be expanded to the length of the probability distribution sequence Y. During the process of expanding the directed graph the labels represented by the nodes and edges of the graph may appear repeatedly so that the length of the label sequence matches the corresponding length of the probability distribution sequence Y. At step 805, the GTC loss function as shown in equation (5) is calculated by summing the conditional probabilities of all node sequences π in the expanded graph
[0190] The summation is calculated efficiently using dynamic programming.
[0191] At step 807, the gradients of the neural network parameters are calculated using the derivatives of the GTC objective function 702 with respect to all possible output symbols, as shown in equations (10) and (4) above, which are calculated efficiently using the forward-backward algorithm and backpropagation. For this purpose, the forward-backward algorithm determines the forward variable α and the backward variable β, where α and β are used to determine the mathematically expressed in equation (12)
[0192] At step 809, the parameters of neural network 701 can be updated according to the gradients calculated in step 807. To update the parameters, a neural network optimization function can be implemented, which defines the rules for updating the parameters of neural network 701. The neural network optimization function can include at least one of the following: Stochastic Gradient Descent (SGD), SGD with momentum, Adam, AdaGrad, AdaDelta, etc.
[0193] At step 811, it can be determined whether to repeat steps 801 to 809 by iterating over the training samples (i.e., pairs of observation sequences and graphs or by iterating over batches of training samples based on at least one of the following: the GTC loss converges to the optimum or satisfies the stopping criterion.
[0194] Some embodiments are based on the recognition that the above GTC objective function 702 needs to be extended to the GTC-e 402 objective function to enable its application to the trained neural network 501 operating under multi-speaker conditions. In the GTC objective function 702, the neural network 701 only predicts the posterior on the nodes. However, in the GTC-e 402 objective function, the neural network 501 even predicts the weights on the edges of a directed graph (such as FIG. 602). For this purpose, it has been discussed that in FIG. 602, the nodes indicate tokens or labels, and the edges indicate speaker transitions. For this purpose, in the extended GTC formula, there are two transition weights on the edge (g, g′) (which connects node g to node G′). First is the deterministic transition weight represented by W (g,g′) which has been described above when discussing the GTC objective function 702, and additionally, the neural network 501 has a predicted transition weight, which is represented as (g,g′) The predicted transition weight in the GTC-e 402 objective function is an additional posterior probability distribution representing the transition weight on the edge (g, g′) at time t, where I(g, g′) ∈ I and I is the set of indices of all possible transitions. The posterior probability is obtained as the output of the softmax.
[0195] Furthermore, in the GTC-e 402 objective function, the forward probability α t (g) defined in Equation (2) is modified to:
[0196]
[0197] where αt(g) represents the total probability of the subgraph of containing all paths starting from node 0 and terminating at node g 0 at time t. It can be calculated for g = 1,..., G. Additionally, if g corresponds to the start node, then α 0 (g) is equal to 1, otherwise α
[0198] t (g) is equal to 0.Furthermore, in the GTC-e formula, the backward probability β
[0199]
[0200] defined in Equation (3) is modified to: where represents the subgraph of
[0201] containing all paths starting from node g and terminating at node G + 1. Similar to GTC, the calculations of α and β can be efficiently performed using the forward-backward algorithm.The neural network 501 is optimized by gradient descent. For any symbol k ∈ u, the gradient of the loss with respect to the label posterior can be obtained in the same way as in CTC and GTC before applying softmax and the corresponding unnormalized network output The key idea is to use the forward and backward variables given in Equation (4) to express the probability function at t
[0202]
[0203] For the transition probability of transition i ∈ I The derivative of the loss with respect to the network output is similar but has some important differences. Here, the key is to express at t as:
[0204]
[0205] Then, the derivative of with respect to the transition probability can be written as:
[0206]
[0207] where, denotes the set of edges in corresponding to transition i.
[0208] To backpropagate the gradient through the softmax function of the derivative with respect to the unnormalized network output before applying softmax is needed, which is
[0209]
[0210] The gradient of the transition weights is derived by substituting (14) and the derivative of the softmax function into (15):
[0211]
[0212] Using the fact that:
[0213]
[0214] and
[0215]
[0216] Thus, using the above GTC-e 402 formula, the neural network 501 is used to perform speech recognition tasks and speaker separation tasks. Specifically, the neural network 501 can use different decoders that can perform time alignment of the sequence of likelihoods (or probabilities) of the transcription output that can perform label or speech recognition and the sequence of likelihoods of the speaker's identity. For example, one embodiment extends the principle of suffix beam search to the multi-speaker scenario. It should be noted that beam search cannot be used in multi-speaker applications that employ speech separation as a preprocessing task or a postprocessing task. However, introducing a multi-head encoder allows adapting the suffix beam search to produce a sequence of transcription outputs annotated with the speaker's identity.
[0217] Figure 9 FIG. shows a beam search algorithm used during the decoding operation of the neural network 501 according to an example embodiment.
[0218] Since the output of the GTC-e 402 objective function contains tokens from multiple speakers, the existing time-synchronous prefix beam search algorithm is modified as shown in Figure 9 FIG. The main modifications are three folds. First, the speaker transition probability 901 is used in score calculation. Second, when expanding the prefix, all possible speaker IDs 902 are considered to account for all possible speakers. Third, when calculating the language model (LM) score of the prefix, subsequences of different speakers are considered separately 903.
[0219] These modifications are used by the decoder 202 of the neural network 501 to perform beam search to produce a sequence of chronologically ordered language tokens, where each token is associated with a speaker identity.
[0220] In some embodiments, at inference, the LM is adopted via shallow fusion. The LM includes 2 long short-term memory (LSTM) neural network layers, and 1024 units each are trained using stochastic gradient descent and the official LM training text data of LibriSpeech, where sentences that appear in the 860h training data subset are excluded. ASR decoding is based on the time-synchronous prefix beam search algorithm. A decoding beam size of 30, a score-based pruning threshold of 14.0, an LM weight of 0.8, and an insertion reward factor of 2.0 are used.
[0221] Figure 10 FIG. shows Table 1, which shows the greedy search results of the ASR performance of the ASR system 104 based on the GTC-e 402 objective function.
[0222] The word error rate (WER) is shown in Table 1. It is observed from the table that the ASR system 104 based on the GTC-e402 objective function is superior to the normal ASR model. Table 1 shows the WER of three models: the single-speaker CTC model 1001, the PIT-CTC model 1002, and the GTC-e model 1003. The GTC-e model 1003 is the ASR system 104 based on GTC-e 402 disclosed in various embodiments described herein. The GTC-e model achieves performance close to that of the PIT-CTC model 1002, especially in the low overlap rate cases (0%, 20%, 40%) 1004.
[0223] Figure 11 Table 2 is shown, and Table 2 shows the greedy search results of the ASR performance based on the oracle token error rate of the ASR system 104 based on the GTC-e 402 objective function.
[0224] Table 2 shows the oracle TER of the PIT-CTC 1002 and the GTC-e model 1003 by comparing only the tokens from all output sequences with all reference sequences regardless of speaker assignment.
[0225] The average test TER of PIT-CTC 1101 and GTC-e 1102 are 22.8% and 25.0% respectively, from which it is determined that the token recognition performance is comparable.
[0226] The GTC-e 1003 is able to accurately predict the activation of most tokens, which is a very good performance metric.
[0227] Figure 12 Table 3 is shown, and Table 3 shows the beam search results of the ASR performance of the ASR system 104 based on the WER based on the GTC-e 402 objective function.
[0228] For the beam search decoding results in Table 3, for the language model, a 16-layer transformer-based LM trained on the external full LibriSpeech data is used. Using the text. The beam size of GTC-e1003 is set to 40, while the beam size of PIT-CTC1002 is cut in half to keep the average beam size per speaker the same. Through beam search, the word error rate is significantly improved.
[0229] Figure 13Table 4 is shown, which shows the beam search results of the ASR performance of the ASR system 104 based on the WEB-based GTC-e 402 objective function for multiple speakers. Table 4 shows rows of WER for different overlap situations for the GTC-e 2-speaker model 1301 (such as the GTC-e 402 objective function of the ASR system 104), speaker 1 1302, and speaker 2 1303.
[0230] As can be seen from the table, the GTC-e model 1301 does not favor any speaker and provides equivalent WER for each speaker.
[0231] Based on the performance results, it can be determined that the GTC-e 402 objective function is beneficial for multi-speaker separation and speech recognition tasks and has good performance. Therefore, the GTC-e 402 objective function can be used in various neural network architectures to perform end-to-end ASR.
[0232] Figure 14A The neural network architecture 1400a of the ASR system implemented using the GTC-e 402 objective function according to an example embodiment is shown.
[0233] In some embodiments, the neural network architecture 1400a corresponds to a transformer-based neural network architecture that employs the proposed GTC-T loss function to train a neural network (e.g., neural network 501).
[0234] In an embodiment, the GTC-T function is explained here with respect to the neural network corresponding to the ASR system. Consider a feature sequence X of length T′ derived from a speech utterance, which is processed by the neural network 501 to produce an output sequence of length T, which may potentially be different from T′ due to downsampling. This output sequence contains a set of posterior probability distributions at each point, since the neural network 501 conditionally depends on the previous label outputs generated by the ASR system and thus has different states that produce multiple posterior probability distributions of the labels. For example, v t,i represents the posterior probability of the neural network state i at time step t, and represents the posterior probability of the output label k of state i at time t. The GTC-T objective function is marginalized over all possible label alignment sequences represented by the graph For GTC, the label information used for training is represented by the graph where the graph corresponds to the directed graph 704. Thus, the conditional probability given the graph is defined by the sum over all node sequences in of length T, which can be written as:
[0235]
[0236] Among them, represents the search function that extends to a grid of length T (not counting non-emitting start and end nodes), π represents a single-node sequence and alignment path, and p(π|X) is the posterior probability of path π given the feature sequence X. The posterior probability is used to calculate the conditional probability of path π given the feature sequence X.
[0237] The nodes are sorted in breadth-first search order and indexed using g = 0,..., G+1, where 0 corresponds to the non-emitting start node and G+1 corresponds to the non-emitting end node. l(g) represents the output symbol observed at node g, and through W g,g′ and I g,g′ represent the transition weight and the decoder state index on the edge connecting nodes g and g'. Finally, π t:t′ =(π t ,..., π t′ ) is the subsequence of nodes of π from time index t to t'. Note that π 0 and π T+1 correspond to the non-emitting start node 0 and the non-emitting end node G+1.
[0238] In RNN-T, the conditional probability p(y|X) given the label sequence y is efficiently computed by a dynamic programming algorithm that is based on computing forward and backward variables and combining them to compute p(y|X) at any given time t[2]. In a similar manner, the GTC-T forward probability can be computed for g = 1,..., G using the following formula.
[0239]
[0240] Among them, represents the subgraph that contains all paths from node 0 to node g. Summing over all possible π that can be generated in t steps for the subsequence up to time index t in the subgraph . Note that if g corresponds to the start node, then α 0 (g) is equal to 1, otherwise it is equal to 0. The backward variable β is computed similarly for g = 1,..., G using the following formula.
[0241]
[0242] Among them, represents the subgraph of G that contains all paths from node g to node G+1. Based on the forward and backward variables at any t, the probability function
[0243]
[0244] For gradient descent training, the derivative of the loss function with respect to the network output must be calculated.
[0245] (20), the network output can be written as
[0246]
[0247] For any symbol k ∈ U and any decoder state i ∈ I, where U represents the set of all possible output symbols and I represents the set of all possible decoder state indices. With respect to the derivative can be written as
[0248]
[0249] where denotes the set of edges corresponding to decoder state i in and where the label k is observed at node g'. To backpropagate the gradient through the softmax function of
[0250]
[0251] Finally, the gradient of the neural network output is
[0252]
[0253] where Equation (24) is derived by substituting (22) and the derivative of the softmax function of
[0254]
[0255] and
[0256]
[0257] Figure 14A Fig. 1400a shows the neural network architecture of an ASR system implemented using the GTC-T objective function.
[0258] In some embodiments, the neural network architecture 1400a corresponds to a transformer-based neural network architecture that employs the proposed GTC-T loss function 1401 to train a neural network (e.g., neural network 501), where the GTC-T loss function 1401 corresponds to the GTC-T objective function. In the neural network architecture 1400, 80-dimensional log spectral energy plus 3 additional features for pitch information are used as acoustic features, which are derived from the audio input 1402 using the feature extraction module 1403.
[0259] In some embodiments, the neural network architecture 1400a includes a two-layer convolutional neural network (CNN) model 1405, followed by a stack of E = 12 transformer-based encoder layers 1406, a linear layer 1407, a prediction network 1408, a connection network 1409, and a final softmax function 1410 to map the neural network output to a posterior probability distribution. In some example embodiments, each layer of the two-layer CNN model 1405 can use a stride of 2, a kernel size of 3×3 with 320 channels, and a rectified linear unit (ReLU) activation function. Additionally, a linear neural network layer 1407 is applied to the output of the last CNN layer. Sine position encoding 1411 is added to the output of the two-layer CNN model 1405 before feeding it to the transformer-based encoder 1406. Each transformer layer employs a feed-forward neural network module with an internal dimension of 1540, a 320-dimensional multi-head self-attention layer with 4 attention heads, and layer normalization. Residual connections are applied to the outputs of the multi-head self-attention and feed-forward modules.
[0260] In an embodiment, the HKUST and LibriSpeech ASR benchmarks are used for evaluation. HKUST is a corpus of Mandarin telephone speech recordings with over 180 hours of transcribed speech data, and LibriSpeech includes nearly 1k hours of read English audiobooks. In an example, the ASR system is configured to first extract 80-dimensional log mel spectrogram energy plus 3 additional features for pitch information. The derived feature sequence is processed by a VGG neural network that downsamples the feature sequence to a frame rate of 40 ms before feeding it to the encoder 1406. The encoder 1406 includes 12 Conformer blocks, where each block includes a self-attention layer, a convolutional module, and two macaron-style feed-forward neural network modules. Additionally, the input to each component of the Conformer block is layer-normalized, and dropout is applied to the outputs of several neural network layers.
[0261] The hyperparameters of the encoder 1406 are d model = 256, d = 2048, dh = 4 and E = 12, while for LibriSpeech, d model and d h are increased to 512 and 8 respectively. For the CTC model, a linear layer and a softmax function are used to project the output of the encoder neural network onto multiple output labels (including the blank label) to derive the probability distribution over the labels. For the GTC-T and RNN-T loss types, two additional neural network components are used, a prediction network 1408 and a joining network 1409. The prediction network 1408 includes a single long short-term memory (LSTM) neural network and a dropout layer. The prediction network 1408 acts as a language model and receives the previously emitted ASR labels (ignoring the blank label) as input. The prediction network 1408 transforms the input of the previously emitted ASR labels received into the embedding space. The joining network 1409 uses a linear layer 1407 and a tanh activation function to combine the encoder frame sequence and the neural network output. Additionally, softmax 1410 is used to map the neural network output to the posterior probability distribution. Dropout with a probability of 0.1 is used after the multi-head self-attention module, after the feed-forward module, and for the internal dimensions of the feed-forward module.
[0262] In some embodiments, data augmentation based on SpecAugement is used for training. In a specific example, the ASR output symbols include the blank symbol plus 5000 sub-words obtained by the SentenePiece method, which are generated only from the transcriptions of the "clean" 100h LibriSpeech training data subset. The ASR model is trained using the Adam optimizer, where β 1 = 0.9, β 2 = 0.98, ∈ = 10 -9 , and the learning rate schedule has 25000 warm-up steps. The learning rate factor and the maximum number of training epochs are set to 1.0 and 50 for HKUST, and 5.0 and 100 for LibriSpeech.
[0263] In some embodiments, a task-specific LSTM-based language model (LM) is trained and employed via shallow fusion during decoding. For HKUST, the LM includes 2 LSTM layers, each with 650 units. For LibriSpeech, alternatively, 4 LSTM layers each with 2048 units are used. For LibriSpeech, the effect of a strong transformer-based LM (Tr-LM) with 16 layers is also tested. The ASR output labels include the blank token plus 5000 sub-word units obtained for LibriSpeech, or the blank token plus 3653 character-based symbols for the HKUST task.
[0264] Figure 14B Shows the pseudo-code 1400b of a beam search algorithm for GTC-T with a similar CTC graph according to an example embodiment. In Figure 14B , l corresponds to the prefix sequence, and at time step t, the prefix probability is divided into those ending with a blank (b) or not ending with a blank (nb) and and θ1 and θ2 are used as thresholds for locally pruning the set of posterior probabilities and for pruning the set of prefix / hypotheses based on scores. More specifically, the function PRUNE(Ω next , p asr , P, θ 2 ) performs two pruning steps. First, the set of hypotheses residing in Ω next is restricted to the P best hypotheses using the ASR score p asr , and then any ASR hypothesis whose ASR score is less than log p best -θ 2 is also removed from the set, where p best represents the best prefix ASR score in the set. The posterior probability v t,i is generated by the neural network using the function NN ET (X, l, t), where X represents the input feature sequence, and i represents the neural network state dependent on the prefix l. The posterior probability of the ASR label k at time frame t and state i is represented by . In addition, α and β are the LM and label insertion reward weights, and |l| represents the sequence length of the prefix l. The symbol represents the blank label, and <sos>Indicates the start of a sentence symbol.
[0265] Figure 14C Shows a comparison 1400c of the ASR results of CTC loss, RNN-T loss, and GTC-T loss on the HKUST benchmark according to an example embodiment.
[0266] In Figure 14C it shows the ASR results of CTC loss, RNN-T loss, and GTC-T loss on the HKUST benchmark. Joint CTC / RNN-T training and parameter initialization for GTC-T training via CTC pre-training greatly improve the ASR results of both the RNN-T-based model and the GTC-T-based model. For example, CTC-based initialization only affects the parameters of the encoder 1406, while the parameters of the prediction network 1408 and the connection network 1409 remain randomly initialized. The ASR results show that for GTC-T training, the use of a CTC-like graph performs better compared to the MonoRNN-T graph. In addition, the GTC-T model outperforms the RNN-T model by 0.5% on the HKUST development test set. Although the use of an LM via shallow fusion does not help significantly improve the word error rate (WER) of the RNN-T- and GTC-T-based ASR models, the CTC-based ASR results are improved between 0.7% and 1.0%. For HKUST, the CTC system also outperforms both the RNN-T system and the GTC-T system.
[0267] Figure 14D Shows a comparison 1400d of the ASR results of CTC loss, RNN-T loss, and GTC-T loss on the LibriSpeech dataset benchmark according to an example embodiment.
[0268] In Figure 14D it shows the ASR results on the larger LibriSpeech dataset. RNN-T and GTC-T outperform the CTC results. For example, using a CTC-like graph, CTC-based initialization, a transformer-based LM, and a beam size of 30 for decoding in GTC-T achieved a WER of 5.9% for other test conditions of LibriSpeech. Despite using a strong LM and a large beam size, this is 0.9% better than the best CTC result. The GTC-T result is also 0.3% better than the best RNN-T result. In addition, similar to the HKUST experiment, GTC-T using a CTC-like graph obtains better results than using the MonoRNN-T graph. However, Figure 14D The results also show that the parameter initialization of the encoder 1406 is particularly important for GTC-T training, and without initialization, the training converges more slowly. For LibriSpeech, when not using an external LM, the RNN-T model performs better than GTC-T.
[0269] Example implementation
[0270] Figure 15 FIG. shows a block diagram of a computer-based system 1500 trained using the GTC-e 402 objective function according to an example embodiment. The computer-based system 1500 may correspond to an end-to-end ASR system 104, an acoustic event detection system, etc.
[0271] The computer-based system 1500 includes a plurality of interfaces that connect the system 1500 to other systems and devices. The system 1500 includes an input interface 1501 configured to receive a plurality of observation sequences 1509, such as a stream of acoustic frames representing features of a speech utterance. Additionally or alternatively, the computer-based system 1500 may receive a plurality of observation sequences from various other types of input interfaces. In some embodiments, the system 1500 includes an audio interface configured to obtain a plurality of observation sequences 1509 (i.e., a stream of acoustic frames) from an acoustic input device 1503. For example, the system 1500 may use a plurality of observation sequences 1509 including acoustic frames in an ASR application or an acoustic event detection application.
[0272] The input interface 1501 is further configured to obtain, for each of the plurality of observation sequences 1509, a plurality of label training sequences 1525, where there is no time alignment between the plurality of label training sequences 1525 and a sequence of probability distributions corresponding to the observation sequences input to the neural network and output by the neural network.
[0273] In some embodiments, the input interface 1501 includes a network interface controller (NIC) 1505 configured to obtain a plurality of observation sequences 1509 and a plurality of label training sequences 1525 via a network 1507, which may be one or a combination of a wired network and a wireless network.
[0274] The network interface controller (NIC) 1505 is adapted to connect the system 1500 to a network 1507 via a bus 1523, which connects the system 1500 to a sensing device (e.g., input device 1503). Additionally or alternatively, the system 1500 may include a human-machine interface (HMI) 1511. The human-machine interface 1511 within the system 1500 connects the system 1500 to a keyboard 1513 and an indicating device 1515, where the indicating device 1515 may include a mouse, trackball, touchpad, joystick, pointing stick, stylus, or touch screen, etc.
[0275] The system 1500 includes a processor 1521 configured to execute stored instructions 1517, and a memory 1519 that stores instructions executable by the processor 1521. The processor 1521 may be a single-core processor, multi-core processor, computing cluster, or any number of other configurations. The memory 1519 may include random access memory (RAM), read-only memory (ROM), flash memory, or any other suitable memory system. The processor 1521 may be connected via the bus 1523 to one or more input devices and output devices.
[0276] The instructions 1517 may implement a method for training a neural network associated with the system 1500 using the GTC-e 402 objective function. According to some embodiments, the system 1500 may be used to implement various applications of neural networks, such as end-to-end speech recognition, acoustic event detection, image recognition, etc. To this end, the computer memory 1519 stores a directed graph 1528, a language model 1527, and the GTC-e 402 objective function. To train the system 1500 using the GTC-e 402 objective function, the directed graph 1528 includes a plurality of nodes connected by edges, where each node represents a label and each edge represents a speaker ID.
[0277] In addition, a path is generated that represents a label training sequence through the nodes and edges of the directed graph 1528, where there are multiple paths.
[0278] In some embodiments, the directed graph 1528 is a weighted graph of nodes weighted by an associated score corresponding to the probability that the transcriptional output of the node is the true transcriptional output at a moment. In some embodiments, the transitions from one node to another are weighted, where the weights can be estimated from the scores of the strong language model (LM) 1527. The directed graph 107 is used by the GTC-e 402 objective function, where the GTC-e 402 objective function is used to train the system 1500 to transform each observation sequence 1509 among a plurality of observation sequences into a sequence of probability distributions over all possible labels at each moment by maximizing the probability of a label sequence corresponding to the sequence of nodes and edges included in the directed graph 1528 at the output of the system 1500, wherein the system 1500 includes an output interface 1535 configured to output a sequence of labels and edges and their likelihoods in terms of probability distributions.
[0279] In some embodiments, the output interface 1535 can output, on a display device 1533, each probability corresponding to each label at each timestamp of the sequence of probability distributions. The sequence of probability distributions can be displayed as a matrix. Examples of the display device 1533 include a computer monitor, a television, a projector, or a mobile device, etc. The system 1500 can also be connected to an application interface 1529, and the application interface 1529 is adapted to connect the system 1500 to an external device 1531 to perform various tasks, such as sound event detection.
[0280] Embodiments
[0281] This specification only provides exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the following description of the exemplary embodiments will provide those skilled in the art with a description of the implementation for implementing one or more exemplary embodiments. Various changes can be made to the functions and arrangements of the elements without departing from the spirit and scope of the disclosed subject matter set forth in the appended claims. Specific details are given in the following description to provide a thorough understanding of the embodiments. However, those of ordinary skill in the art can understand that the embodiments can be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter can be shown in block diagram form as components so as not to obscure the embodiments with unnecessary details. In other cases, well-known processes, structures, and technologies can be shown without unnecessary details so as not to obscure the embodiments. Additionally, the same reference numerals and markings in the various figures indicate the same elements.
[0282] In addition, each embodiment can be described as a process depicted as a flowchart, process view, data flow diagram, structure diagram, or block diagram. Although a flowchart may describe operations as a sequential process, many operations can be performed in parallel or concurrently. In addition, the order of the operations can be rearranged. A process can terminate when its operations are completed, but can have additional steps not discussed or included in the figure. Moreover, not all operations in any particularly described process may occur in all embodiments. A process can correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the function can correspond to the return of the function to the calling function or the main function.
[0283] In addition, embodiments of the disclosed subject matter can be implemented, at least in part, manually or automatically. The manual or automatic implementation can be performed or at least assisted by using a machine, hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the necessary tasks can be stored in a machine-readable medium. A processor can perform the necessary tasks.
[0284] In addition, the embodiments of the present disclosure and the functional operations described in this specification can be implemented in digital electronic circuits, in tangible, specifically implemented computer software or firmware, in computer hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. In addition, some embodiments of the present disclosure can be implemented as one or more computer programs, that is, one or more modules of computer program instructions encoded on a tangible non-transitory program carrier for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. In addition, the program instructions can be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical signal, optical signal, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0285] A computer program (which may also be referred to or described as a program, software, software application, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may or may not correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0286] For example, a computer suitable for executing a computer program can be based on a general-purpose or special-purpose microprocessor or both, or any other type of central processing unit. Generally, the central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a central processing unit for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include or be operatively coupled to receive data from and / or transfer data to one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data. However, a computer need not have such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few.
[0287] To provide interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and an indicating device, such as a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from the devices used by the user; for example, by sending a web page in response to a request received from a web browser on the user's client device.
[0288] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component (such as a data server), or a middleware component (such as an application server), or a front-end component (such as a client computer having a graphical user interface or a web browser), through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), such as the Internet.
[0289] The computing system can include clients and servers. The clients and servers are typically remote from each other and typically interact through a communication network. The relationship between the client and the server arises from computer programs that run on the respective computers and have a client-server relationship with each other.
[0290] Although the present disclosure has been described with reference to certain preferred embodiments, it should be understood that various other changes and modifications can be made within the spirit and scope of the present disclosure. Accordingly, aspects of the claims cover all such changes and modifications that fall within the true spirit and scope of the present disclosure.< / sos>
Claims
1. An end-to-end automatic speech recognition (ASR) system, the ASR system comprising: a processor; and a memory storing instructions thereon, wherein the processor is configured to execute the stored instructions to cause the ASR system to: collect an acoustic frame sequence providing a digital representation of an acoustic signal, the acoustic signal including a mixture of speech performed by multiple speakers; encode each frame using a multi-head encoder that encodes the likelihood of transcribing the output and the likelihood of the identity of the speaker for each frame in the acoustic frame sequence to produce a sequence of likelihoods of the identity of the speaker corresponding to the acoustic frame sequence and a sequence of likelihoods of the transcribing output; decode the sequence of likelihoods of the transcribing output and the sequence of likelihoods of the identity of the speaker using a decoder that performs alignment to produce a sequence of transcribing outputs annotated with the identity of the speaker; and submit the sequence of transcribing outputs annotated with the identity of the speaker to a downstream application.
2. The ASR according to claim 1, wherein, the decoder uses beam search to generate a chronologically ordered sequence of language tokens, where each token is associated with the identity of the speaker.
3. The ASR according to claim 2, wherein, the beam search is configured to perform operations including one or a combination of the following: (1) generate speaker transition probabilities and language token probabilities, (2) calculate scores of language tokens, (3) expand the prefix lists of all speakers in the set of possible speakers, and (4) calculate scores of prefixes by separately considering subsequences of different speakers.
4. The ASR according to claim 1, wherein, the encoder includes an acoustic encoder configured to process an input acoustic signal and generate an encoder state sequence, and the decoder includes an attention-based decoder.
5. The ASR according to claim 1, wherein, the encoder and the decoder form at least a part of a neural network, and at least a part of the neural network is trained to achieve multiple objectives by minimizing a loss function, the loss function including a first component associated with errors in speech recognition and a second component associated with errors in speaker identification.
6. The ASR system according to claim 5, wherein, a connectionist temporal classification (CTC) objective function is used to train the neural network.
7. The ASR system according to claim 5, wherein, the encoder and the decoder form at least a part of the neural network trained using an extended CTC objective function to enforce alignment between the input and the output on a graph having nodes indicating transcribing outputs and edges indicating speaker transitions.
8. The ASR system according to claim 7, wherein, the extended CTC objective function is a graph-based temporal classification (GTC-e) objective function, where the GTC-e objective function uses supervision information from a directed graph of nodes connected by edges representing labels and transitions between labels, where the directed graph represents possible alignment paths of sequences of probability distributions output by the neural network and the labels.
9. The ASR system according to claim 8, wherein, the directed graph represents multiple possible alignment paths of the probability distribution sequence and the label sequence, such that the structure of the directed graph that may be traversed allows multiple unique label sequences, the multiple unique label sequences being obtained after collapsing label repetitions and removing blank labels from the multiple unique label sequences, thereby resulting in a non-monotonic alignment between the label sequence and the probability distribution sequence.
10. The ASR system according to claim 9, wherein, the non-monotonic alignment is encoded into the structure of the directed graph by allowing transitions from one label to multiple other non-blank labels, by allowing transitions from one label to multiple other blank labels, or both.
11. The ASR system according to claim 7, wherein, the extended CTC objective function is a graph-based temporal classification transducer GTC-T objective function.
12. The ASR system according to claim 7, wherein, the nodes of the directed graph indicate tokens from all speakers in chronological order.
13. The ASR system according to claim 7, wherein, the edges of the directed graph indicate speaker identification information.
14. A computer-implemented method for end-to-end automatic speech recognition ASR, the computer-implemented method comprising the steps of: collecting a sequence of acoustic frames providing a digital representation of an acoustic signal, the acoustic signal including a mixture of speech performed by multiple speakers; encoding each frame in the sequence of acoustic frames using a multi-head encoder that encodes the likelihood of the frame as a transcription output and the likelihood of the identity of the speaker, to produce a sequence of likelihoods of the identity of the speaker corresponding to the sequence of acoustic frames and a sequence of likelihoods of the transcription output; decoding the sequence of likelihoods of the transcription output and the sequence of likelihoods of the identity of the speaker using a decoder that performs alignment, to produce a sequence of transcription outputs annotated with the identity of the speaker; and submitting the sequence of transcription outputs annotated with the identity of the speaker to a downstream application.
15. The method according to claim 14, wherein, the decoder uses beam search to produce a chronologically ordered sequence of language tokens, wherein each token is associated with the identity of a speaker.
16. The method according to claim 15, wherein, the beam search is configured to perform operations including one or a combination of the following: (1) generating speaker transition probabilities and language token probabilities, (2) calculating scores of language tokens, (3) expanding the prefix lists of all speakers in the set of possible speakers, and (4) calculating scores of prefixes by separately considering subsequences of different speakers.
17. The method according to claim 14, wherein, the encoder includes a self-attention encoder, and the decoder includes an attention-based decoder.
18. The method according to claim 14, wherein, The encoder and the decoder form at least a part of a neural network, and at least a part of the neural network is trained to achieve multiple objectives by minimizing a loss function, where the loss function includes a first component associated with an error in speech recognition and a second component associated with an error in speaker identification.
19. The method according to claim 18, wherein, a connectionist temporal classification (CTC) objective function is used to train the neural network.
20. The method according to claim 19, wherein, the encoder and the decoder form at least a part of the neural network trained with an extended CTC objective function to enforce an alignment between the input and the output on a graph having nodes indicating transcription outputs and edges indicating speaker transitions.
Citation Information
Patent Citations
System and method for end-to-end speech recognition with triggered attention
US11100920B2