Training of a model for processing sequence data

The novel training method for CTC models by shifting predicted sequences relative to input sequences addresses time delays in ASR systems, enabling timely output predictions and reducing latency for real-time applications.

CN115244616BActive Publication Date: 2025-07-15INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180019824.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-03
Filing Date
2021-03-19
Publication Date
2025-07-15
Estimated Expiration
2041-03-19

AI Technical Summary

Technical Problem

The existing CTC loss function results in a time delay between the acoustic features and the output symbol during training, affecting the efficiency of the streaming automatic speech recognition system.

Method used

By performing forward shift training on the CTC model, adjust the length of the predicted sequence to match the input sequence, and update the model parameters using the shifted predicted sequence to reduce time delay.

Benefits of technology

Effectively reduces the time delay between model output and input, so that the trained model can output prediction results earlier, suitable for streaming applications, and improves recognition efficiency without affecting accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115244616B_ABST
    Figure CN115244616B_ABST
Patent Text Reader

Abstract

A technique for training a model is disclosed. Training samples are obtained, where the training samples include an observed input sequence and a target sequence having symbols of a length different from that of the observed input sequence. The observed input sequence is fed into the model to obtain a predicted sequence. The predicted sequence is shifted by an amount relative to the observed input sequence. The shifted predicted sequence and the target sequence of symbols are used to update the model based on a loss.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND ART

[0001] The present disclosure generally relates to machine learning, and more particularly to techniques for training models for processing sequence data.

[0002] End-to-end automatic speech recognition (ASR) systems using the CTC (Connectionist Temporal Classification) loss function have received much attention due to their ease of training and decoding efficiency. End-to-end ASR systems use a CTC model to predict a sequence of sub-words or words with or without a subsequent language model. CTC-based ASR can operate faster than related NN (neural network) / HMM (hidden Markov model) hybrid systems. Thereby, a significant reduction in power consumption and computational resource costs is expected.

[0003] The combination of a unidirectional LSTM model and the CTC loss function is one of the promising ways to build a streaming ASR. However, generally, such a combination suffers from a time delay between acoustic features and output symbols during decoding, which increases the latency of the streaming ASR. The related NN / HMM hybrid system trained from frame-level forced alignment between acoustic features and output symbols does not suffer from time delay. Compared with the hybrid model, the CTC model is usually trained with training samples of acoustic features and output symbols having different lengths. This means there is no time alignment supervision. The CTC model trained without frame-level alignment produces output symbols after the model consumes sufficient information of the output symbols, which results in a time delay between acoustic features and output symbols.

[0004] To reduce the time delay between acoustic features and output symbols, a method of imposing constraints on CTC alignment has been proposed (Andrew Senior et al., "Acoustic modeling with CD-CTC-sMBR LSTM RNN", Proc. ASRU, 2015, pp. 604–609). It has been studied that the delay can be limited by restricting a set of search paths used in the forward-backward algorithm to those paths where the delay between the CTC label and the "ground truth" alignment does not exceed a certain threshold. However, the method disclosed in this document requires iterative steps to prepare frame-level forced alignment before CTC model training.

[0005] U.S. Patent Application 20170148431A1 discloses end-to-end deep learning systems and methods for identifying speech in distinct languages such as English or Chinese. The entire pipeline of hand-designed components is replaced by neural networks, and end-to-end learning allows for the processing of a wide variety of speech, including noisy environments, accents, and different languages. However, the techniques disclosed in this patent document attempt to modify the neural network topology.

[0006] U.S. Patent Application 20180130474A1 discloses methods, systems, and devices, including computer programs encoded on a computer storage medium for learning pronunciations from acoustic sequences. The method includes: stacking one or more frames of acoustic data to produce a sequence of modified frames of acoustic data; processing the sequence of modified frames of acoustic data through an acoustic modeling neural network including one or more recurrent neural network (RNN) layers and a final CTC output layer to produce a neural network output. The techniques disclosed in this patent document only adjust the input to the encoder to reduce the frame rate.

[0007] Accordingly, there is a need for a novel training technique that can effectively reduce the time delay between the output and the input of a model, where the model is trained using training samples with input observations and output symbols of different lengths. Summary of the Invention

[0008] According to an embodiment of the present invention, there is provided a computer-implemented method for training a model. The method includes obtaining training samples, the training samples including an input sequence of observations and a target sequence of symbols having a length different from the input sequence of observations. The method further includes feeding the input sequence of observations into the model to obtain a predicted sequence. The method further includes shifting the predicted sequence relative to the input sequence of observations by an amount. The method further includes updating the model based on a loss using the shifted predicted sequence and the target sequence of symbols.

[0009] The method according to an embodiment of the present invention enables the trained model to output predictions at an appropriate timing to reduce the delay of the prediction process relative to the input.

[0010] In a preferred embodiment, the predicted sequence may be shifted forward relative to the input sequence of observations to produce a shifted predicted sequence and the model is unidirectional. The method enables the trained model to output predictions earlier to reduce the delay of the prediction process relative to the input. The model trained by this method is suitable for streaming applications.

[0011] In a specific embodiment, the model may be a model based on a recurrent neural network. In a particular embodiment, the loss may be a CTC (Connectionist Temporal Classification) loss.

[0012] In a specific embodiment, the shifted prediction sequence includes an adjustment such that the lengths of the shifted prediction sequence and the observed input sequence are the same.

[0013] In a preferred embodiment, the shifted prediction sequence and using the shifted prediction sequence to update the model can be performed at a predetermined rate. Thereby, the method enables the trained model to balance the accuracy and latency of the prediction process.

[0014] In a specific embodiment, the model can be a neural network-based model with multiple parameters. Feeding the input sequence includes forward propagation through the neural network-based model. Updating the model includes performing backpropagation through the neural network-based model to update the multiple parameters.

[0015] In a further preferred embodiment, the model can be an end-to-end speech recognition model. Each observation in the input sequence of the training samples can represent acoustic features, and each symbol in the target sequence of the training samples can represent a phone, context-dependent phone, character, word segment, or word. Thus, the method can enable the speech recognition model to output the recognition result at an appropriate timing to reduce the overall latency of the speech recognition process, or provide more time for subsequent processes to improve the recognition accuracy.

[0016] Also described and claimed herein are computer systems and computer program products related to one or more aspects of the present invention.

[0017] According to other embodiments of the present invention, there is provided a computer program product for decoding using a model. The computer program product includes a computer-readable storage medium having program instructions embodied thereon. The program instructions are executable by a computer to cause the computer to perform a method that includes feeding an input into the model to obtain an output. The model is trained by: obtaining training samples that include an observed input sequence and a target sequence of symbols having a length different from the observed input sequence. It can be further trained by feeding the observed input sequence into the model to obtain a prediction sequence, shifting the prediction sequence by an amount relative to the observed input sequence, and updating the model based on a loss using the shifted prediction sequence and the target sequence of symbols.

[0018] The computer program product according to the embodiments of the present invention is capable of outputting predictions at an appropriate timing to reduce the latency of the prediction process relative to the input.

[0019] Additional features and advantages are realized by the techniques of the present invention. Other embodiments and aspects of the present invention are described in detail herein and are considered to be part of the claimed invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The subject matter regarded as the invention is particularly pointed out and distinctly claimed in the claims. The above and other features and advantages of the present invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:

[0021] Figure 1 A block diagram of a speech recognition system according to an exemplary embodiment of the present invention is shown, the speech recognition system including a forward-shifted CTC (Connectionist Temporal Classification) training system for training a CTC model for speech recognition;

[0022] Figure 2 A schematic diagram of a unidirectional LSTM CTC model as an example of a CTC model to be trained according to an embodiment of the present invention is shown;

[0023] Figure 3 A speech signal for an example sentence "this is true" and resultant phoneme probabilities calculated for the example speech signal by bidirectional and unidirectional LSTM phoneme CTC models trained through a standard training process are shown;

[0024] Figure 4 A manner of forward-shifted CTC training with one frame shift according to an exemplary embodiment of the present invention is depicted;

[0025] Figure 5 Is a flowchart depicting a novel forward-shifted CTC training process for training a CTC model for speech recognition according to an exemplary embodiment of the present invention;

[0026] Figure 6 Posterior probabilities of a phoneme CTC model trained by forward-shifted CTC training are shown, where the maximum number of frames to be shifted is set to 1 and the rate of samples to be shifted varies from 0.1 to 0.3;

[0027] Figure 7 Posterior probabilities of a phoneme CTC model trained by forward-shifted CTC training are shown, where the rate of samples to be shifted is set to 0.1 and the maximum number of frames to be shifted varies from 1 to 3;

[0028] Figure 8 Posterior probabilities of a word CTC model trained by forward-shifted CTC training are shown, where the maximum number of frames to be shifted is set to 1 and the rate of samples to be shifted is set to 0.1;

[0029] Figure 9 Time delays by phoneme and word CTC models relative to a hybrid model are shown; and

[0030] Figure 10Shows a schematic diagram of a computer system according to one or more embodiments of the present invention. Detailed implementation

[0031] Hereinafter, the present invention will be described with reference to specific embodiments, but those skilled in the art will understand that the embodiments described below are only mentioned by way of example and are not intended to limit the scope of the present invention.

[0032] One or more embodiments according to the present invention relate to a computer-implemented method, a computer system, and a computer program product for training a model for processing sequence data, wherein a predicted sequence obtained from the model being trained is shifted by an amount relative to an observed input sequence, and the shifted predicted sequence is used to update the model based on a computed loss.

[0033] Hereinafter, first with reference to Figures 1 - 4 , a computer system for training a model according to an exemplary embodiment of the present invention will be described, wherein the model to be trained is a CTC (Connectionist Temporal Classification) model for speech recognition and the sequence data to be processed is a sequence of acoustic features. Then, with reference to Figure 5 , a computer-implemented method for training a model according to an exemplary embodiment of the present invention will be described, wherein the model to be trained by this method is a CTC model for speech recognition, and the sequence data to be processed is a sequence of acoustic features. Then, reference will be made to Figures 6 - 9 An experimental study of a novel CTC training for speech recognition according to an exemplary embodiment of the present invention will be described. Finally, with reference to Figure 10 , the hardware configuration of a computer system according to one or more embodiments of the present invention will be described.

[0034] Hereinafter, with reference to Figure 1 , a block diagram of a speech recognition system 100 including a CTC training system 110 with forward shift according to an exemplary embodiment of the present invention is described.

[0035] As Figure 1 shown, the speech recognition system 100 may include a feature extraction module 104 for extracting acoustic features from an input; and a speech recognition module 106 for performing speech recognition on the input.

[0036] The speech recognition system 100 according to an exemplary embodiment of the present invention further includes: a CTC training system 110 with forward shift for performing a novel CTC training to obtain a trained CTC model 170 that constitutes the speech recognition module 106; and a training data memory 120 for storing a set of training data used in the novel CTC training performed by the CTC training system 110 with forward shift.

[0037] The feature extraction module 104 can receive the audio signal data 102 digitized by sampling an audio signal at a predetermined sampling frequency and a predetermined depth as input. For example, an audio signal can be input from a microphone. The feature extraction module 104 can also receive the audio signal data 102 from a remote client device via a network such as the Internet. The feature extraction module 104 is configured to extract acoustic features from the received audio signal data 102 by any known acoustic feature analysis to generate a sequence of the extracted acoustic features.

[0038] The acoustic features can include, but are not limited to, MFCC (Mel Frequency Cepstral Coefficients), LPC (Linear Predictive Coding) coefficients, PLP (Perceptual Linear Prediction) cepstral coefficients, log Mel spectrograms, or any combination thereof. The acoustic features can further include 'dynamic' acoustic features, for example, static delta features and double delta features of the aforementioned acoustic features.

[0039] Note that the elements of the acoustic feature sequence are called "frames", and the audio signal data 102 includes a series of sampled values of the audio signal at a predetermined frequency. Generally, the audio signal data 102 is sampled at 8,000 Hz for narrowband audio and at 16,000 Hz for wideband audio. The duration of each frame in the acoustic feature sequence can be (but is not limited to) about 10 to 40 milliseconds.

[0040] The speech recognition module 106 is configured to convert an input sequence of the extracted acoustic features into an output sequence of words. The speech recognition module 106 uses the CTC model 170 to predict the most reasonable speech content of the input sequence of the extracted acoustic features and outputs the result 108.

[0041] The speech recognition module 106 according to an exemplary embodiment of the present invention uses the CTC model 170 and can be an end-to-end model. In a specific embodiment, the speech recognition module 106 can include a sub-word (e.g., phoneme, character) unit end-to-end model. In other embodiments, the speech recognition module 106 can include a word unit end-to-end model. Examples of the units of the end-to-end model can include phonemes, characters, context-dependent phonemes (e.g., triphones and pentaphones), word fragments, words, etc. The speech recognition module 106 at least includes the CTC model 170. The CTC model 170 is the target of the novel CTC training performed by the forward-shifted CTC training system 110. The CTC model 170 is defined as a model trained by using a CTC loss function, and its architecture is not limited.

[0042] When the speech recognition module 106 is configured with a subword (e.g., phoneme) unit end-to-end model, in addition to the CTC model 170 that outputs a subword sequence, the speech recognition module 106 further includes appropriate language models such as n-gram models and neural network-based models (e.g., RNN (recurrent neural network)) and a lexicon. When the speech recognition module 106 is configured with a word unit end-to-end model, the speech recognition module 106 may only include the CTC model 170 that directly outputs a word sequence and does not require a language model and a lexicon.

[0043] Moreover, the speech recognition module 106 can perform speech recognition only using a neural network and does not require a complex speech recognition decoder. However, in other embodiments, the language model can be further applied to the results of the word unit end-to-end model to improve the accuracy of speech recognition. Also, in the described embodiments, the speech recognition module 106 receives an input sequence of acoustic features. However, in another embodiment, the original waveform of the audio signal data 102 can also be received by the speech recognition module 106. Thus, the original audio signal data 102 can be regarded as a kind of acoustic feature.

[0044] The speech recognition module 106 finds the word sequence with the highest probability based on the input sequence of acoustic features and outputs the word sequence as the result 108.

[0045] Figure 1 The forward-shifted CTC training system 110 shown in is configured to perform novel CTC training to obtain the CTC model 170 that at least partially constitutes the speech recognition module 106.

[0046] In the described embodiments, the training data memory 120 stores a set of training data, and each training data includes speech data and a corresponding transcription.

[0047] Note that the speech data stored in the training data memory 120 can be given in the form of an acoustic feature sequence after feature extraction, and this feature extraction can be the same as the speech data used for inference performed by the feature extraction module 104 in the front-end process. If the speech data is given in the form of the same audio signal data as the audio signal data 102 used for inference, the speech data can be subjected to feature extraction before training to obtain an acoustic feature sequence. In addition, the transcription can be given in the form of a sequence of phonemes, context-dependent phonemes, characters, word fragments, or words in a manner depending on the units targeted by the CTC model 170.

[0048] In the described embodiments, each training sample is given as a pair of an observed input sequence and a target sequence of symbols, where the observations are acoustic features and the symbols are sub-words (e.g., phonemes) or words. The training data can be stored in an internal or external memory operatively coupled to the processing circuitry.

[0049] The forward-shifted CTC training system 110 performs a novel CTC training process to obtain the CTC model 170. During the novel CTC training process, the forward-shifted CTC training system 110 performs a predetermined process on the predicted sequence obtained from the CTC model, which is trained prior to CTC calculation and parameter update of the CTC model.

[0050] Before describing the novel CTC training, first, an exemplary architecture of the CTC model is described.

[0051] Reference Figure 2 , a schematic diagram of the LSTM CTC model is shown as an example of the CTC model. To train the CTC model, pairs of unaligned sub-word (e.g., phoneme) / word sequences and audio signal data are fed. The LSTM CTC model 200 may include: an input component 202 for receiving an input sequence of acoustic features obtained from a given audio signal data via feature extraction; an LSTM encoder 204; a softmax function 206; and a CTC loss function 208. As an input to the CTC model, frame stacking is also envisioned, where consecutive frames are stacked together as a super-frame.

[0052] The LSTM encoder 204 converts the input sequence of acoustic features into high-level features. Figure 2 The LSTM encoder 204 shown in [[ ]] is unidirectional. Note that the term "unidirectional" means that the network obtains information only from past states and does not obtain future states, as compared to a bidirectional model that obtains information from both past and future states simultaneously. The use of a unidirectional model is preferred for streaming (and possibly real-time) ASR because the unidirectional model does not require the entire frame sequence before decoding. When the acoustic features arrive, the unidirectional model can output predictions sequentially.

[0053] The softmax function 206 calculates a probability distribution by normalization based on the output high-level features obtained from the LSTM encoder 204. The CTC loss function 208 is a specific type of loss function designed for sequence labeling tasks.

[0054] Note that the related NN / HMM hybrid system training requires frame-level alignment and requires the length of the target sequence of phonemes to be equal to the length of the input sequence of acoustic features. This frame-level alignment can typically be achieved by forced alignment techniques. However, this frame-level alignment makes the training process complex and time-consuming.

[0055] Compared with related NN / HMM systems, the target sequence of sub-words or words required to train a CTC model can have a different length from the input sequence of acoustic features. Generally, the length of the input sequence of acoustic features is much longer than the target sequence of sub-words or words. That is, frame-level alignment is not required, and there is no time alignment supervision for training the CTC model.

[0056] Due to the above properties, a CTC model with a unidirectional LSTM encoder trained without frame-level alignment produces output symbols after the CTC model consumes sufficient information of the output symbols, which results in a time delay between the acoustic features and the output symbols (sub-words or words). This time delay is not the type that can be reduced by investing a large amount of resources.

[0057] Figure 3 The waveform of the speech signal of the example sentence "this is true" at the top is shown. Figure 3 Also shown in the middle and at the bottom are the resultant phoneme probabilities calculated by bidirectional and unidirectional LSTM phoneme CTC models trained through a standard CTC training process for the example speech signal.

[0058] The CTC model emits a sharp and sparse posterior distribution over the target output symbols (sub-words or words), where most frames emit the blank symbol with a high probability and a few frames emit the target output symbols of interest. It should be noted that, at each time index, in addition to at least the blank, the symbol with the highest posterior probability is referred to as a'spike' in this paper. In Figure 3 , for the purpose of convenience, the probability of the blank symbol is omitted. The spike timing emitted by the trained CTC model is generally not controlled.

[0059] As Figure 3 shown at the bottom, the spike timing from the unidirectional LSTM model is delayed from the Figure 3 actual acoustic features and speech signal shown at the top. The output corresponding to the spikes detecting the phonemes 'DH', 'IH', 'S', 'IH', 'Z', 'T', 'R', and 'UW' has a time delay with respect to the input acoustic features.

[0060] Note that the bidirectional LSTM model outputs posterior probabilities aligned with the acoustic features, as Figure 3 shown in the middle. This means that the bidirectional LSTM model gives spike signals in a more timely manner than the unidirectional model. This is because the bidirectional LSTM model summarizes the entire input sequence of acoustic features before decoding and utilizes information from both past and future states. Thus, using the unidirectional model is preferred for streaming ASR.

[0061] To reduce the time delay between the peak timing and the acoustic features, the forward-shifted CTC training system 110 according to an exemplary embodiment of the present invention performs a forward shift on the predicted sequence (posterior probability distribution) obtained from the CTC model, which is trained by backpropagation before CTC calculation and parameter update.

[0062] Return reference Figure 1 and Figure 4 , and a detailed block diagram of the forward-shifted CTC training system 110 is further described. Figure 4 Describes the manner of forward-shifted CTC training with one frame shift according to an exemplary embodiment of the present invention.

[0063] As Figure 1 shown, the forward-shifted CTC training system 110 may include: an input feeding module 112 for feeding an input sequence of acoustic features into the CTC model being trained to obtain a predicted sequence, wherein a forward shift module 114 is used to forward-shift the obtained predicted sequence; and an update module 116 for updating the parameters of the CTC model being trained in a manner based on the shifted predicted sequence.

[0064] The input feeding module 112 is configured to first obtain a training sample, which includes an input sequence of acoustic features as the correct label and a target sequence of sub-words or words. The input feeding module 112 is further configured to feed the input sequence of acoustic features included in each training sample into the CTC model being trained to obtain a predicted sequence.

[0065] Let X denote a sequence of acoustic feature vectors over T time steps, and x t be the acoustic feature vector at time index t (t = 1,..., T) in the sequence X. As Figure 4 shown at the top, by performing a conventional forward propagation through the CTC model from the acoustic feature vector sequence X = {x1,..., x T}, a predicted sequence O = {o1,..., o T} is obtained, where o t (t = 1,..., T) represents the prediction at each time index t, and each prediction o t is the posterior probability distribution over the target output symbols (sub-words or words).

[0066] The forward shift module 114 is configured to shift the obtained predicted sequence O by a predetermined amount relative to the input sequence of the acoustic feature vectors X to obtain a shifted predicted sequence O'. In a preferred embodiment, the obtained predicted sequence O is forward-shifted relative to the input sequence of the acoustic feature X.

[0067] The forward shift module 114 is further configured to adjust so that the length of the shifted prediction sequence O' is the same as the input sequence of acoustic feature vectors X. In a particular embodiment, the last element o of the prediction corresponding to the predetermined amount to be shifted may be used. T One or more copies of are used to fill the end of the predicted sequence O' to make adjustments. Therefore, the length of the predicted sequence O' remains T. The last element o of the prediction is used T A copy of is an example. In the described embodiment, padding can be used as a way of adjustment. However, the way of adjustment is not limited to padding. In another specific embodiment, the adjustment can be performed by trimming the sequence of acoustic feature vectors X from one end according to a predetermined shift amount.

[0068] When the predetermined shift amount (the number of frames to be shifted) is one frame, the shifted prediction sequence O' maintains the remaining predictions except for the beginning, and the shifted prediction sequence O' is a set {o2, ..., o T , o T},like Figure 4 Note that the last element of the prediction is o T is doubled in the prediction shift sequence O'. It should also be noted that only the prediction (posterior probability distribution) is shifted, and the other predictions including the input acoustic feature vector are not shifted.

[0069] In a preferred embodiment, the shifting of the prediction sequence O is not performed for each training sample, but only for a portion of the training samples, which can be determined by a predetermined rate. The predetermined rate at which the training samples are to be shifted can range from about 5% to 40%, more preferably, about 8% to 35%. In addition, the unit amount of training samples to be shifted can be one training sample or a group of training samples (e.g., a small batch).

[0070] In addition, in certain embodiments, the amount to be shifted (or the number of frames to be shifted) can be fixed to an appropriate value. The fixed value can depend on the goal of reducing the delay time and the duration of each frame. A delay reduction commensurate with the duration of each frame and the number of frames to be shifted will be obtained.

[0071] In another embodiment, the amount to be shifted may be determined probabilistically within a predetermined range. The term "probabilistically" means relying on a predetermined distribution, such as a uniform distribution. The predetermined range or the upper limit of the predetermined range may be determined in a manner depending on the goal of reducing the delay time and the duration of each frame. As will be shown later by experiments, a reduction in delay commensurate with the duration of each frame and the average number of frames to be shifted will be obtained.

[0072] The update module 116 is configured to update the model based on a loss function using a shifted predicted sequence O' and a target sequence of symbols (sub-words or words) included in the training samples. As Figure 4 shown at the bottom of

[0073] CTC calculation includes a process of CTC alignment estimation. Let y denote a sequence of target output symbols of length L and y i (i = 1,..., L) be the i-th sub-word or word in the target sequence y. Compared with the alignment-based NN / HMM hybrid system training where L is required to be equal to T, CTC introduces an additional blank symbol which extends the length-L sequence y into a set of length-T sequences Φ(y), thus allowing unaligned training. Each sequence y^ (y^ is an element of Φ(y) and is a set of {y1^, y2^, y3^,... T-1 ^, y T ^}) in this set of length-T sequences is one of the CTC alignments between the sequence of acoustic feature vectors X and the sequence of target output symbols y.

[0074] For example, assume the given output phoneme sequence is 'ABC' and the length of the input sequence is 4. In this case, the possible phoneme sequences would be {AABC, ABBC, ABCC, ABC_, AB_C, A_BC, _ABC}, where "_" represents the blank symbol

[0075] The CTC loss is defined as the sum of the symbol posterior probabilities of all possible CTC alignments, as follows:

[0076]

[0077] CTC training maximizes the sum of the possible output sequences or minimizes the negative of that sum, while allowing blank outputs for any frame probability. The update module 116 updates the parameters of the CTC model to minimize the CTC loss L CTC . Note that minimizing the loss (CTC loss) includes maximizing the negative of the loss, which can be referred to as a reward, utility, or fitness. The update module 116 can calculate the CTC loss L based on the shifted predicted sequence O' CTC , and based on the CTC loss L CTC perform backpropagation through the entire network to update the parameters of the CTC model each time a training sample (online) or a set of training samples (e.g., mini-batch) is processed.

[0078] The background idea for CTC training for forward shifting is as follows: If before forward shifting, the CTC alignment (ŷ1, ŷ2, ŷ3, ..., ŷ T-1 ^, ŷ T ^) has a high probability P(ŷ|X) in the above equation (1), then the forward shifting of the prediction sequence is converted into a high probability of the CTC alignment for forward shifting (ŷ2, ŷ3, ..., ŷ T-1 ^, ŷ T ^, ŷ T ^). Due to the high probability of the CTC alignment for forward shifting, the entire network of the CTC model is trained to promote the CTC alignment for forward shifting through backpropagation, which results in a reduced time delay after the entire training.

[0079] In the described embodiments, the CTC model is described as a unidirectional LSTM model. However, the architecture of CTC, which is the target of the novel CTC training for forward shifting, is not restricted and can be any RNN type model, which includes basic RNN, LSTM (Long Short-Term Memory), GRU (Gated Recurrent Unit), Elman network, Jordan network, Hopfield network, etc. Moreover, the RNN type model can include more complex architectures, such as any one of the foregoing RNN type models used in combination with other architectures (such as CNN (Convolutional Neural Network), VGG, ResNet, and Transformer).

[0080] Note that the novel CTC training for forward shifting is only for training. The same trained CTC model as that trained by conventional CTC training can be used. The topology (e.g., the way neurons are connected) and configuration (e.g., the number of hidden layers and units) of the CTC model, as well as the way of decoding with the CTC model, remain unchanged.

[0081] In a specific embodiment, each of the modules 104, 106 described in Figure 1 and each of the modules 112, 114, and 116 of the CTC training system 110 for forward shifting can but are not limited to be implemented as a software module, which includes program instructions and / or data structures combined with hardware components (such as processors, memories, etc.); as a hardware module including electronic circuits; or a combination thereof.

[0082] They can be implemented on a single computer device (such as a personal computer and a server machine) or in a distributed manner (such as a computer cluster of computer devices, a client-server system, a cloud computing system, an edge computing system, etc.) on multiple devices.

[0083] The processing circuitry of a computer system implementing the forward-shifted CTC training system 110 can be operably coupled to an internal or external storage device or medium that provides the training data memory 120 and the memory for the parameters of the CTC model 170, which can be provided using any internal or external storage device or medium.

[0084] Also in a particular embodiment, the feature extraction module 104 and the speech recognition module 106 (including the CTC model 170 trained by the forward-shifted CTC training system 110) are implemented on a computer system on the user side, while the forward-shifted CTC training system 110 is implemented on a computer system on the provider side of the speech recognition system.

[0085] In a further variant embodiment, the feature extraction module 104 is implemented only on the user side and the speech recognition module 106 is implemented on the provider side. In this embodiment, the computer system on the client side transmits only a sequence of acoustic features to the computer system on the provider side and receives the decoding result 108 from the provider side. In another variant embodiment, both the feature extraction module 104 and the speech recognition module 106 are implemented on the provider side, and the computer system on the client side transmits only the audio signal data 102 to the computer system on the provider side and receives the decoding result 108 from the provider side.

[0086] Hereinafter, referring to Figure 5 , a novel forward-shifted CTC training process for training a CTC model for speech recognition according to an exemplary embodiment of the present invention is described. Figure 5 is a flowchart depicting the novel forward-shifted CTC training process. Note that Figure 5 the process shown in Figure 1 can be executed by a processing circuitry, such as the processing unit of a computer system implementing the forward-shifted CTC training system 110 and its modules 112, 114, and 116 shown in

[0087] For example, Figure 5 the process shown in

[0088] can start at step S100 in response to receiving a request from an operator for the novel forward-shifted CTC training.

[0089] In step S101, the processing unit can set training parameters, which include the maximum number of frames to be shifted (shift amount) and the rate of samples to be shifted in the novel forward-shifted CTC training.

[0089] In step S102, the processing unit can prepare a set of training samples from the training data memory 120. Each training sample can include an input sequence of acoustic feature vectors X having length T and a target sequence of symbols (sub-words (e.g., phonemes) or words) y having length L.

[0090] In step S103, the processing unit may initialize the CTC model. Appropriately set the initial values of the parameters of the CTC model.

[0091] In step S104, the processing unit may pick one or more training samples from the prepared set. A small batch of training samples may be picked.

[0092] In step S105, for each picked training sample, the processing unit may perform forward propagation through the CTC model by feeding an input sequence of the acoustic feature vector X to obtain a predicted sequence O of length T.

[0093] In step S106, the processing unit may determine whether to perform a forward shift based on the rate of samples to be shifted given in step S101. A small batch may be randomly selected as the target of forward shift at a predetermined rate.

[0094] In step S107, the processing unit may branch the process depending on the way the determination made in step S106. In step S107, when the processing unit determines that the picked training sample is the target of forward shift (yes), the process may proceed to S108.

[0095] In step S108, the processing unit may determine the number of frames to be shifted based on the maximum number of frames to be shifted given in step S101. The number of frames to be shifted may be probabilistically determined based on a specific distribution. In a specific embodiment, for the selected small batch, the number of frames to be shifted may be determined from an integer uniform distribution up to an upper bound (the maximum number of frames to be shifted).

[0096] In step S109, the processing unit may perform a forward shift on the predicted sequence O obtained in step S105 to generate a predicted shifted sequence O'.

[0097] On the other hand, in response to determining in step S107 that the picked training sample is not the target of forward shift (no), the process may proceed directly to S110.

[0098] In step S110, the processing unit may calculate the CTC loss using the shifted predicted sequence O' or the predicted original sequence O, and perform backpropagation through the CTC model to update the parameters of the CTC model. A forward shift may be performed on the selected small batch. For the remaining small batches, CTC training may proceed as normal.

[0099] In step S111, the processing unit may determine whether the process ends. When a predetermined convergence condition or termination condition is met, the processing unit may determine that the process will be terminated.

[0100] In response to determining that the process is not ended (No) in step S111, the process can loop back to S104 for subsequent training samples. On the other hand, in response to determining that the process is ended (Yes) in step S111, the process can proceed to S112. In step S112, the processing unit can store the currently obtained parameters of the CTC model into a suitable storage device, and the process can end in step S113.

[0101] According to the above embodiments, a novel training technique is provided, which is capable of reducing the time delay between the output and the input of a model in an effective manner, and the model is trained using training samples with input observations and output symbols of different lengths.

[0102] The novel CTC training enables the trained model to output predictions at appropriate timings to reduce the delay of the prediction process relative to the input. Preferably, the trained model outputs predictions earlier, and the delay of the prediction process relative to the input can be reduced. The trained model is suitable for streaming ASR applications. As shown in the experimental results described later, the time delay and speech recognition accuracy can be balanced by tuning the training parameters of the novel CTC training.

[0103] Although the actual total delay of streaming end-to-end ASR is affected by other factors, reducing the time delay between acoustic features and symbols will result in lower latency of streaming ASR or allow more time for subsequent (post) processes to improve accuracy. Compared with conventional CTC training, the novel forward-shifted CTC training does not require any additional information, such as frame-level forced alignment. In addition, at decoding, the same CTC model trained with this forward-shifted training can be used as the model trained with conventional CTC training.

[0104] In addition, as shown in the experimental results described later, there is almost no adverse effect on the accuracy of the CTC model for speech recognition.

[0105] In the above embodiments, the CTC model trained by the forward-shifted CTC training system 110 is described as being directly used as the CTC model 170 that constitutes the speech recognition module 106. However, in other embodiments, the CTC model trained by the forward-shifted CTC training system 110 may not be directly used as the CTC model 170. In a specific embodiment, the CTC model trained by the forward-shifted CTC training system 110 may be used in a knowledge distillation framework. For example, a unidirectional LSTM trained by the forward-shifted CTC training system 110 can be used as a guiding CTC model, and a bidirectional LSTM model trained under the guidance of the guiding CTC model can be used as a teacher model for a student unidirectional LSTM model (which serves as the CTC model 170) for knowledge distillation. For example, a bidirectional LSTM model trained under the guidance of a guiding CTC model can be used as a teacher model for knowledge distillation of a student unidirectional LSTM model, where the student unidirectional LSTM model is trained based on forward shifting.

[0106] It is also noted that in still other embodiments, the CTC model trained by the forward-shifted CTC training system 110 can be used not only as the CTC model 170. In other specific embodiments, post-fusion involving the CTC model trained by the forward-shifted CTC training system 110 is also envisioned.

[0107] It should be noted that the languages to which the novel training for speech recognition according to the exemplary embodiments of the present invention is applicable are not limited, and such languages may include, but are not limited to, for example, Arabic, Chinese, English, French, German, Japanese, Korean, Portuguese, Russian, Swedish, Spanish. Since the novel training has a non-aligned nature, the GMM / HMM system for forced alignment can be omitted. In addition, when a word unit end-to-end model is adopted, no dictionary and no language model are required. Therefore, the novel training is applicable to some languages for which it is difficult to prepare a GMM / HMM system and / or a dictionary.

[0108] Furthermore, in the foregoing embodiments, the novel forward-shifted CTC training has been described as being applied to speech recognition. However, the applications to which the CTC model is applicable are not limited to speech recognition. The CTC model can be used in different sequence recognition tasks other than speech recognition. Moreover, the problem of latency appears not only in speech recognition but also in other sequence recognition tasks. Such sequence recognition tasks may include handwritten text recognition from images or stroke sequences, optical character recognition, gesture recognition, machine translation, etc. Therefore, it is expected to apply the novel forward-shifted CTC training to such other sequence recognition tasks.

[0109] Although advantages have been described and will be described hereinafter with respect to one or more specific embodiments obtained in accordance with the present invention, it should be understood that some embodiments may not have these potential advantages and that these potential advantages are not necessarily required for all embodiments.

[0110] Experimental study

[0111] Encode and execute the program for the forward-shifted CTC training system 110 shown in Figure 1 and the forward-shifted CTC training process described in Figure 5 using a given speech data set pair. Conduct ASR experiments using a standard English conversational phoneme speech data set to verify the operation of the novel forward-shifted CTC training. Apply the novel forward-shifted CTC training to train a unidirectional LSTM phoneme CTC model and a unidirectional LSTM word CTC model. Measure the time delay from a sufficiently strong offline hybrid model trained by frame-level forced alignment.

[0112] Experimental setup

[0113] Use 262 hours of segmented speech from the standard 300-hour Switchboard-1 audio with transcripts. For acoustic features, extract 40-dimensional logMel filter bank energies on 25 msec frames every 10 msec. This static feature and its delta and double-delta coefficients are used for frame stacking with a downsampling rate of 2. For evaluation, use the Switchboard (SWB) and CallHome (CH) subsets of the NIST Hub5 2000 evaluation data set. Considering that the training data includes data similar to SWB, testing on the CH test set is a mismatched scenario for the model.

[0114] For the unidirectional LSTM phoneme CTC model, use 44 phonemes and the blank symbol from the Switchboard pronunciation dictionary. For decoding, train a 4-gram language model with 24M words from the Switchboard and Fisher transcripts with a vocabulary size of 30K. Construct the CTC decoding graph. For the neural network architecture, stack 6 unidirectional LSTM layers (unidirectional LSTM encoder) with 640 units and a fully connected linear layer of 640 x 45, followed by a softmax activation function. Initialize all neural network parameters as samples with a uniform distribution on (-∈, ∈), where ∈ is the reciprocal square root of the input vector size.

[0115] For the word CTC model, words that occur at least five times in the training data are selected. This results in an output layer with 10,175 words and the blank symbol. The same 6 unidirectional LSTM layers are selected, and 1 fully connected linear layer with 256 units is added to reduce the computation. A fully connected linear layer of 256x10,176 is placed, followed by a softmax activation function. To better converge, the unidirectional LSTM encoder part is initialized with the trained phoneme CTC model. Other parameters are initialized in a similar way to the phoneme CTC model. For decoding, simple peak picking is performed on the output word posterior distribution, and duplicates and blank symbols are removed.

[0116] All models are trained for 20 epochs and use a learning rate that starts from 0.01 and anneals at (0.5) after epoch 10 1 / 2 per epoch using Nesterov-accelerated stochastic gradient descent. The batch size is 128.

[0117] For the novel forward-shifted CTC training, two parameters include "shift max" and "shift rate", where "shift max" indicates the maximum number of frames to be shifted, and "shift rate" is the rate of the mini-batch whose output posterior probability is shifted. Training mini-batches are randomly selected at the "shift rate" and forward-shifted for the selected mini-batches. For the shift size, a shift size is selected from a uniform distribution of integers greater than 0 to an upper limit provided by the training parameter "shift max" for each selected mini-batch.

[0118] Post-spike signal

[0119] As mentioned above, the CTC model emits a very sharp posterior distribution. Examine the posterior probabilities of the utterance "this (DHIHS) is (IHZ) true (TRUW)" from the SWB test set.

[0120] Figure 6 The posterior probabilities of the phoneme CTC model trained by the novel forward-shifted CTC training are shown, where the maximum number of frames to be shifted (shift max) is set to 1 and the sample rate to be shifted (shift rate) varies from 0.1 to 0.3 (Examples 1 - 3). Compared with the post-spike at the top from the conventional CTC training (Comparative Example 1), the post-spike from the forward-shifted CTC training (Examples 1 - 3) appears earlier, which is the expected behavior of the novel forward-shifted CTC training.

[0121] Figure 7Shows the posterior probabilities of the phoneme CTC model with forward shifting in CTC training, where the rate of samples to be shifted (shift rate) is set to 0.1 and the maximum number of frames to be shifted (shift maximum) varies from 1 to 3 (Examples 1, 4, and 7). Although some earlier spikes also appear here, some spikes are not forward shifted, especially when the maximum number of frames to be shifted (shift maximum) is large.

[0122] Finally, study the posterior probabilities of the word CTC model with forward shifting in CTC training. Figure 8 Shows the posterior probabilities of the word CTC model with forward shifting in CTC training, where the maximum number of frames to be shifted (shift max) is set to 1, and the rate of samples to be shifted (shift rate) is set to 0.1 (Example 10). As Figure 8 shown, it is confirmed that earlier spike timings are obtained from the CTC model trained with forward shifting. Note that for all word CTC models with conventional and forward shifting in CTC training, a unidirectional LSTM phoneme CTC model trained with conventional CTC training is used for initialization.

[0123] Latency from the hybrid model

[0124] Next, study the time delay of each word after decoding on the SWB and CH test sets. Figure 9 Shows the definition of the time delay through the phoneme and word CTC models relative to the hybrid model. In Figure 9 , each box represents the unit for output symbol prediction for each model. Note that due to the lower frame rate achieved through input frame stacking and subsampling in the CTC model, the unit sizes between the hybrid model and the CTC model are different. For the hybrid model, "-b", "-m", and "-e" represent three states of the HMM, and for simplicity, the identifiers of the context-dependent variants of each state are omitted from this figure.

[0125] To set the ground truth for the timing of words, a sufficiently strong offline hybrid model trained on the 2000-hour Switchboard + Fisher dataset using iterative and careful forced alignment steps is used. More specifically, a combination of two bidirectional LSTMs and a residual network (ResNet) acoustic model and an n-gram language model is used for decoding, and timestamps are obtained for each word. The WERs of the SWB and CH test sets with this hybrid model are 6.7% and 12.1% respectively, which are much better than those with the following CTC models. This hybrid model is not for streaming ASR and is used to obtain proper alignment for reference. Additionally, it is noted that this hybrid model is trained with more training data.

[0126] For the output from the phoneme CTC model, the timestamps after decoding are also obtained. For hybrid and phoneme CTC decoding, the same graph-based static decoder is used, while appropriately handling the blank symbol. For the word CTC model, the first spike of each word occurrence is assumed to be its start time. To measure the latency, the recognition results from the hybrid and CTC models are first aligned with a dynamic programming-based string matching, and the average latency at the start of the aligned words is calculated, as Figure 9 shown. For the unidirectional LSTM phoneme CTC model, the shift maximum is changed from 1 to 3, and the shift rates from 0.1 to 0.3 are investigated (Examples 1-9). The conditions and evaluation results of the examples and comparative examples of the unidirectional LSTM phoneme CTC model are summarized in Table 1.

[0127] Table 1

[0128]

[0129]

[0130] Although there are some fluctuations in the WER and time delay, it is demonstrated that a constant reduction in latency is obtained by using the novel forward-shifted CTC training. This trend is common in the matched SWB and unmatched CH test sets. For example, as written in bold in Table 1, the time delay is reduced by 25 milliseconds without observing a negative impact on the WER. By setting the shift rate larger, a further reduction in time delay is obtained at the expense of the WER, which can provide an option for developers of streaming applications to tune the time delay. Similar to the previous studies on spike timing, no additional reduction in time delay is observed by setting the shift maximum larger.

[0131] For the unidirectional LSTM word CTC model, the same settings as in the previous studies on spike timing are used (Example 10). The conditions and evaluation results of Example 10 and Comparative Example 2 for the unidirectional LSTM word CTC model are summarized in Table 2.

[0132] Table 2

[0133]

[0134] It is confirmed that the training with the novel forward shift reduces the time delay by approximately 25 milliseconds while observing a marginal WER degradation in the SWB test set.

[0135] It is demonstrated that the novel forward-shifted CTC training can reduce the time delay between the acoustic features and the output symbols in unidirectional LSTM phoneme and word CTC models. The novel forward-shifted CTC training is also studied to enable the trained model to generate spikes earlier. It is also confirmed that in most cases, the time delay can be reduced by about 25 msec without negatively affecting the WER. Notably, the CTC model trained with the novel forward-shifted training simply generates output symbols earlier and can be used without changing the existing decoder of the model trained with conventional CTC training. Note that reducing the delay to less than 200 milliseconds (which is called the acceptable limit in human-computer interaction) is desirable, and different efforts have been made and combined to achieve this. The 25 milliseconds obtained only by a simple modification of the CTC training pipeline is preferred and can be combined with other efforts.

[0136] Computer hardware components

[0137] Now referring to Figure 10 , a schematic diagram of an example of a computer system 10 that can be used in a speech recognition system 100 is shown. Figure 10 The computer system 10 shown in [[ ]] is implemented as a computer system. The computer system 10 is only one example of a suitable processing device and is not intended to impose any limitation on the scope of use or functionality of the embodiments of the present invention described herein. In any case, the computer system 10 is capable of implementing and / or performing any of the functions set forth above.

[0138] The computer system 10 can operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that can be suitable for use with the computer system 10 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, in-vehicle devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems or devices, etc.

[0139] The computer system 10 can be described in the general context of computer system-executable instructions, such as program modules, executed by the computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types.

[0140] As Figure 10As shown, computer system 10 is shown in the form of a general-purpose computing device. The components of computer system 10 may include, but are not limited to, a processor (or processing unit) 12 and a memory 16, which is coupled to processor 12 via a bus including a memory bus or memory controller, and a processor or local bus using any one of a variety of bus architectures.

[0141] Computer system 10 generally includes a variety of computer system readable media. Such media can be any available media accessible by computer system 10, and it includes volatile and non-volatile media, removable and non-removable media.

[0142] Memory 16 may include computer system readable media in the form of volatile memory, such as random access memory (RAM). Computer system 10 may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 18 may be provided for reading from and writing to a non-removable non-volatile magnetic medium. As will be further depicted and described below, storage system 18 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of embodiments of the present invention.

[0143] A program / utility having a set (at least one) of program modules may be stored in storage system 18 by way of example and not limitation, as well as an operating system, one or more application programs, other program modules, and program data. Each or some combination of the operating system, one or more application programs, other program modules, and program data may include an implementation of a network environment. The program modules generally perform the functions and / or methods of embodiments of the present invention as described herein.

[0144] The computer system 10 may also communicate with one or more peripheral devices 24, such as a keyboard, a pointing device, an automotive navigation system, an audio system, etc.; a display 26; one or more devices that enable a user to interact with the computer system 10; and / or any device that enables the computer system 10 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may occur via an input / output (I / O) interface 22. Additionally, the computer system 10 may communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet), via a network adapter 20. As shown, the network adapter 20 communicates with other components of the computer system 10 via a bus. It should be understood that, although not shown, other hardware and / or software components may be used in conjunction with the computer system 10. Examples include, but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0145] Computer program implementation

[0146] The present invention may be a computer system, a method, and / or a computer program product. The computer program product may include one or more computer-readable storage media having thereon computer-readable program instructions for causing a processor to execute aspects of the present invention.

[0147] A computer-readable storage medium may be a tangible device that is capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card, or a raised structure in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0148] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.

[0149] The computer-readable program instructions for carrying out operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (e.g., using an Internet service provider via the Internet). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuit, so as to carry out aspects of the present invention.

[0150] The present invention will be described below with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0151] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus produce means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, so that the computer-readable storage medium storing the instructions comprises a manufacture including instructions embodying aspects of implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0152] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0153] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or combinations of special purpose hardware and computer instructions.

[0154] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms as well. It will be further understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of the stated features, steps, layers, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, layers, elements, components, and / or combinations thereof.

[0155] All apparatus or steps in the claims, plus the corresponding structures, materials, acts, and equivalents of the functional elements, if any, are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of one or more aspects of the invention has been presented for purposes of illustration and description but is not intended to be exhaustive or limited to the invention in the form disclosed.

[0156] Numerous modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been chosen to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology found in the marketplace, or to enable those of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A computer-implemented method for training a model, comprising: Obtaining training samples, the training samples including an observed input sequence and a target sequence having symbols of a length different from that of the observed input sequence; Feeding the observed input sequence into the model to obtain a predicted sequence; Determining a time delay between spike timings in the predicted sequence and acoustic features of the input sequence, and reducing the time delay by shifting the predicted sequence forward relative to the input sequence for a portion of the training samples to produce a shifted predicted sequence; And Updating the model using the shifted predicted sequence and the target sequence of the symbols, based on a loss calculated by a loss function.

2. The method according to claim 1, wherein, The model is unidirectional.

3. The method according to claim 2, wherein The model is a model based on a recurrent neural network.

4. The method according to claim 2, wherein The loss is a connectionist temporal classification loss.

5. The method according to claim 2, wherein, Shifting the predicted sequence includes: Adjusting such that the lengths of the shifted predicted sequence and the observed input sequence are the same.

6. The method according to claim 2, wherein Performing the shifting of the predicted sequence at a predetermined rate and updating the model using the shifted predicted sequence.

7. The method according to claim 6, wherein The predetermined rate is in the range of 5% to 40%.

8. The method according to claim 2, wherein The amount of the shift is fixed.

9. The method according to claim 2, wherein Determining the amount of the shift probabilistically within a predetermined range.

10. The method according to claim 1, wherein The model is a neural network-based model having a plurality of parameters, wherein feeding back the input sequence includes forward propagation through the neural network-based model, and wherein updating the model includes performing backpropagation through the neural network-based model to update the plurality of parameters.

11. The method according to claim 1, wherein, The model includes an end-to-end speech recognition model, each observation in the input sequence of the training samples represents acoustic features, and each symbol in the target sequence of the training samples represents a phoneme, a context-dependent phoneme, a character, a word segment, or a word.

12. A computer system for training a model by executing program instructions, the computer system comprising: A memory storing the program instructions; A processing circuit in communication with the memory to execute the program instructions, wherein the processing circuit is configured to: Obtain training samples, the training samples including an observed input sequence and a target sequence having symbols of a length different from that of the observed input sequence; Feed the observed input sequence into the model to obtain a predicted sequence; Determine a time delay between spike timings in the predicted sequence and acoustic features of the input sequence, and reduce the time delay by shifting the predicted sequence forward relative to the input sequence for a portion of the training samples to produce a shifted predicted sequence; And Update the model using the shifted predicted sequence and the target sequence of the symbols, based on a loss calculated by a loss function.

13. The computer system according to claim 12, wherein, The model is unidirectional.

14. The computer system according to claim 13, wherein, The processing circuit is further configured to: Adjust to shift the predicted sequence such that the lengths of the shifted predicted sequence and the observed input sequence are the same.

15. The computer system according to claim 13, wherein, Perform the shifting of the predicted sequence and update the model using the shifted predicted sequence at a predetermined rate.

16. The computer system according to claim 12, wherein, The model is a model based on a recurrent neural network.

17. A computer program product for training a model, the computer program product comprising a computer-readable storage medium having program instructions embodied thereon, the program instructions being executable by a computer to cause the computer to perform a method comprising the following steps: Obtain a training sample, the training sample including an observed input sequence and a target sequence of symbols having a length different from that of the observed input sequence; Feed the observed input sequence into the model to obtain a predicted sequence; Determine a time delay between spike timings in the predicted sequence and acoustic features of the input sequence, and reduce the time delay by shifting the predicted sequence forward relative to the input sequence for a portion of the training sample to produce a shifted predicted sequence; and Update the model based on a loss calculated by a loss function using the shifted predicted sequence and the target sequence of the symbols.

18. The computer program product according to claim 17, wherein, The model is unidirectional.

19. The computer program product according to claim 18, wherein, Shifting the predicted sequence includes: Adjusting such that the lengths of the shifted predicted sequence and the observed input sequence are the same.

20. The computer program product according to claim 18, wherein, Performing the shifting of the predicted sequence at a predetermined rate and updating the model using the shifted predicted sequence.

21. The computer program product according to claim 17, wherein, The model is a model based on a recurrent neural network.

22. A computer program product for decoding using a model, the computer program product comprising a computer-readable storage medium having program instructions embodied thereon, the program instructions being executable by a computer to cause the computer to perform a method comprising the following steps: Feed an input into the model to obtain an output; Among them, The model is trained by: Obtaining a training sample, the training sample including an observed input sequence and a target sequence of symbols having a length different from that of the observed input sequence; Feeding the observed input sequence into the model to obtain a predicted sequence; Determining a time delay between spike timings in the predicted sequence and acoustic features of the input sequence, and reducing the time delay by shifting the predicted sequence forward relative to the input sequence for a portion of the training sample; And Updating the model based on a loss calculated by a loss function using the shifted predicted sequence and the target sequence of the symbols.

Citation Information

Patent Citations

  • End-to-end speech recognition

    US20170148431A1

  • Speech recognition with acoustic models

    US20180130474A1

  • Systems and methods for speech transcription

    CN107077842A