Attention-based end-to-end speech recognition method and device

By adopting attention-based end-to-end ASR training method in speech recognition, cross-entropy training, beam search and large-space training, the problem of inefficiency of existing MBR training is solved, and more efficient training and more matching performance with MAP decoding is achieved.

CN119993132APending Publication Date: 2025-05-13TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510152825.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-02-14
Filing Date
2020-02-07
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing Minimum Bayesian Risk (MBR) training methods are inefficient in speech recognition and do not match the performance of MAP decoding, resulting in insufficient training efficiency and performance.

Method used

Attention-based end-to-end (E2E) automatic speech recognition (ASR) training method is adopted to determine the best hypothesis and maximize the distance between the reference sequence and the best hypothesis through cross-entropy training, beam search and large-space training to improve the discriminantity of the model.

Benefits of technology

Improves the efficiency of speech recognition training, making it more suitable for MAP decoding, while providing similar performance to the NWER method on the same benchmark dataset and is easier to apply.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993132A_ABST
    Figure CN119993132A_ABST
Patent Text Reader

Abstract

A method of attention-based end-to-end (E2E) automatic speech recognition (ASR) training includes: performing cross entropy training on a model based on one or more input features of a speech signal; performing a beam search using the model to generate a list of n good hypotheses of output hypotheses; and determining an optimal hypothesis in the generated n-optimal hypothesis list. The method further includes determining a character-based gradient and a word-based gradient based on the model and a loss function that maximizes a distance between the reference sequence and the determined optimal hypothesis; and performing back propagation on the model based on the determined character-based gradient and the determined word-based gradient to update the model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Application No. 16 / 276,081, filed on February 14, 2019, the disclosure of which is incorporated herein by reference in its entirety.

[0003] This application files a divisional application for the Chinese patent application with application number 202080014798.9, application date February 7, 2020, and invention name “Large Interval Tracking for Attention-Based End-to-End Speech Recognition”. Technical Field

[0004] The present disclosure relates to speech processing technology, and more particularly to a method and device for attention-based end-to-end (E2E) automatic speech recognition (ASR) training. Background Art

[0005] Minimum Bayesian risk (MBR) training, such as minimum word error rate (MWER) training, aims to minimize the expected risk on the output hypothesis. The risk or loss L can be expressed as follows:

[0006]

[0007] where (χ, s) represents a sample in the training set D, χ represents the input features, s represents its corresponding sequence label, s′ represents the output hypothesis generated during training, and θ represents the model parameters. The difference between s and s′ is denoted as l(s′, s), which is the word or character level edit distance.

[0008] Accordingly, the output sequence during evaluation can be generated as follows:

[0009]

[0010] in, represents a candidate output sequence, and Indicates the output sequence selected by MBR decoding.

[0011] The search space of grows exponentially with its length, making MBR decoding mechanisms such as Recognition Output Voting Error Reduction (ROVER) inefficient. Although methods based on n-best lists or confusion networks can improve efficiency, beam search decoding based on maximum a posteriori (MAP) is still one of the most commonly used evaluation methods in practice. In MAP decoding, the output hypothesis with the highest (logarithmic) posterior is directly used for evaluation, as shown below:

[0012]

[0013] The mismatch between MBR training and MAP decoding indicates that there may be a training scheme with better efficiency and comparable performance than MBR training. Therefore, providing a technical solution with better efficiency and comparable performance than MBR training is a technical problem to be solved in the art. Summary of the invention

[0014] According to an embodiment, a method for attention-based end-to-end (E2E) automatic speech recognition (ASR) training includes: performing cross entropy training on a model based on one or more input features of a speech signal; performing beam search using the model to generate an n-best hypothesis list of output hypotheses; and determining the best hypothesis in the generated n-best hypothesis list. The method also includes: determining a character-based gradient and a word-based gradient based on the model and a loss function that maximizes the distance between a reference sequence and the determined best hypothesis; and performing back propagation on the model based on the determined character-based gradient and the determined word-based gradient to update the model.

[0015] According to an embodiment, a device for attention-based E2E ASR training includes: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code. The program code includes: a first execution code configured to cause at least one processor to perform cross entropy training on a model based on one or more input features of a speech signal; a second execution code configured to cause at least one processor to perform beam search using the model to generate an n-best hypothesis list of output hypotheses; and a first determination code configured to cause at least one processor to determine the best hypothesis in the generated n-best hypothesis list. The program code also includes: a second determination code and a third determination code, wherein the second determination code and the third determination code are configured to cause at least one processor to determine a character-based gradient and a word-based gradient based on the model and a loss function that maximizes the distance between a reference sequence and the determined best hypothesis; and a third execution code configured to cause at least one processor to perform back propagation on the model based on the determined character-based gradient and the determined word-based gradient to update the model.

[0016] According to an embodiment, a non-transitory computer-readable medium stores instructions that, when executed by at least one processor of a device, cause the at least one processor to: perform cross entropy training on a model based on one or more input features of a speech signal; perform beam search using the model to generate an n-best hypothesis list of output hypotheses; and determine the best hypothesis in the generated n-best hypothesis list. The instructions also cause the at least one processor to: determine a character-based gradient and a word-based gradient based on the model and a loss function that maximizes the distance between a reference sequence and the determined best hypothesis; and perform back propagation on the model based on the determined character-based gradient and the determined word-based gradient to update the model.

[0017] It can be seen that according to the method and device for end-to-end automatic speech recognition training based on attention of the present application, cross entropy training is performed on the model based on one or more input features of the speech signal; beam search is performed using the model to generate an n-best hypothesis list of output hypotheses; the best hypothesis is determined in the generated n-best hypothesis list; based on the model and the loss function that maximizes the distance between the reference sequence and the determined best hypothesis, character-based gradients and word-based gradients are determined; and back propagation is performed on the model based on the determined character-based gradients and the determined word-based gradients to update the model. Therefore, compared with MBR, the large interval-based training standard described in the present application is better matched with MAP decoding. In addition, the training scheme of the present application can provide performance similar to that of the NWER method on the same benchmark data set; in the present application, the large interval principle is combined with deep neural network learning, which is easier to apply than the current SVM deep learning combination method. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a diagram of an environment in which the methods, apparatuses, and systems described herein may be implemented, according to an embodiment;

[0019] Figure 2 yes Figure 1 diagrams of example components of one or more devices;

[0020] Figure 3 is a flow chart of a method for attention-based E2E ASR training according to an embodiment; and

[0021] Figure 4 is a diagram of an apparatus for attention-based E2E ASR training, according to an embodiment. DETAILED DESCRIPTION

[0022] Attention-based E2E ASR systems have a simplified training pipeline compared to conventional recognition systems, but their performance still needs to be improved. Many studies have devoted themselves to various training strategies to improve recognition accuracy.

[0023] Embodiments described herein include a sequence-by-sequence large-margin training scheme for an attention-based E2E ASR system. Rather than minimizing expected loss, this training scheme maximizes the margin. In other words, when the best hypothesis is not the reference, the model is trained to maximize the score difference between the best hypothesis and the reference, making the model more discriminative. Unlike the MBR-based MWER standard, this new approach focuses on only one hypothesis per utterance during training, but still achieves comparable performance. When both approaches employ character-level edit distance, the large-margin approach consistently outperforms MWER training. The large-margin training scheme is also a more concise formulation of the concept of large margins. It maintains the original model structure and is therefore easier to apply than current support vector machine (SVM)-based systems. The new model was tested on the benchmark SWB300h dataset. It achieved the same performance as the results of MWER training.

[0024] In detail, attention-based E2E ASR systems map input audio features into text sequences through three steps: encoding, attention, and decoding. The most commonly used decoder is a recurrent network trained with a point-wise cross-entropy loss. Recently, sequence-discriminative optimization criteria such as MBR-based MWER have been applied to improve model performance. MWER requires multiple assumptions during training. Because decoding is based on the commonly used MAP evaluation, large-margin-based training criteria may match MAP decoding better than MBR.

[0025] The large margin concept is often fused with SVM. By enlarging the margin between the reference sequence and the incorrect sequence, the upper bound of the generalization error can be minimized. In recent years, structured SVM (SSVM) has been combined with various deep neural networks (DNN) for speech recognition tasks. In these models, the softmax layer is replaced by the SSVM layer, and the training process consists of two stages. In the first stage, the weights of the SSVM layer are calculated using the cutting-plane algorithm for all training samples. The parameters in the DNN are then updated using the back-propagation algorithm. On tasks such as phone text message dictation, the sequence-level implementation of the deep neural support vector machine (DNSVM) shows better performance than the corresponding sequence-discriminatively trained DNN. The method described in this paper can maintain the original deep neural network structure and is therefore easier to apply than DNSVM.

[0026] Definitions of abbreviations and terms appearing in the detailed description include the following:

[0027] Speech Recognition System: A computer program that can recognize and translate speech signals into written characters / words.

[0028] Encoder-Decoder: A model architecture where the encoder network maps raw input to feature representations and the decoder takes the feature representations as input and produces output.

[0029] Attention-based end-to-end (E2E) model: A model with an encoder-decoder architecture coupled with an attention scheme that enables the model to learn to focus on specific parts of the input sequence during decoding.

[0030] Large Margin: Large margin or maximum margin is a learning principle that trains a model to maximize the distance to boundary examples.

[0031] Support Vector Machine (SVM): A discriminative training method that exploits the large margin principle to learn the optimal hyperplane for classifying new examples.

[0032] Minimum Bayesian Risk (MBR): A training / decoding principle that aims to minimize the expected error in classification.

[0033] Figure 1 is a diagram of an environment in which the methods, devices, and systems described herein may be implemented, according to an embodiment. Figure 1 As shown, the environment may include user devices, platforms, and networks. The devices of the environment may be interconnected via wired connections, wireless connections, or a combination of wired and wireless connections.

[0034] A user device includes one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with a platform. For example, a user device may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smart phone, a wireless phone, etc.), a wearable device (e.g., a pair of smart glasses or a smart watch), or a similar device. In some implementations, a user device may receive information from a platform and / or send information to a platform.

[0035] The platform includes one or more devices as described elsewhere herein. In some implementations, the platform may include a cloud server or cloud server group. In some implementations, the platform may be designed to be modular so that software components can be swapped in or out depending on specific needs. In this way, the platform can be easily and / or quickly reconfigured for different uses.

[0036] In some implementations, as shown, the platform can be hosted in a cloud computing environment. It is worth noting that although the implementations described herein describe the platform as being hosted in a cloud computing environment, in some implementations, the platform may not be cloud-based (i.e., may be implemented outside of a cloud computing environment) or may be partially cloud-based.

[0037] A cloud computing environment includes an environment that hosts a platform. A cloud computing environment can provide computing, software, data access, storage, and other services that do not require end users (e.g., user devices) to be aware of the physical location and configuration of the systems and / or devices of the hosting platform. As shown, a cloud computing environment can include a set of computing resources (collectively referred to as "computing resources" and individually referred to as a "computing resource").

[0038] Computing resources include one or more personal computers, workstation computers, server devices, or other types of computing and / or communication devices. In some implementations, computing resources can host platforms. Cloud resources can include computing instances executed in computing resources, storage devices provided in computing resources, data transmission devices provided by computing resources, etc. In some implementations, computing resources can communicate with other computing resources via wired connections, wireless connections, or a combination of wired and wireless connections.

[0039] If further Figure 1 As shown in the figure, the computing resources include a set of cloud resources, such as one or more applications ("APP"), one or more virtual machines ("VM"), a virtualized storage device ("VS"), one or more hypervisors ("HYP"), etc.

[0040] An application includes one or more software applications that can be provided to or accessed by a user device and / or platform. An application can eliminate the need to install and execute software applications on a user device. For example, an application can include software associated with a platform and / or any other software that can be provided via a cloud computing environment. In some implementations, an application can send information to or receive information from one or more other applications via a virtual machine.

[0041] A virtual machine includes a software implementation of a machine (e.g., a computer) such as a physical machine that executes a program. A virtual machine can be a system virtual machine or a process virtual machine, depending on the use and degree of correspondence of the virtual machine to any real machine. A system virtual machine can provide a complete system platform that supports the execution of a complete operating system ("OS"). A process virtual machine can execute a single program and can support a single process. In some implementations, a virtual machine can execute on behalf of a user (e.g., a user device) and can manage the infrastructure of a cloud computing environment, such as data management, synchronization, or long-duration data transfer.

[0042] A virtualized storage device includes one or more storage systems and / or one or more devices that use virtualization technology within a storage system or device of a computing resource. In some implementations, within the context of a storage system, the types of virtualization may include block virtualization and file virtualization. Block virtualization may refer to the extraction (or separation) of logical storage from physical storage so that the storage system can be accessed without considering physical storage or heterogeneous structures. Separation may allow administrators of storage systems flexibility in terms of administrator management of storage for end users. File virtualization may eliminate the dependency between data accessed at the file level and the location where the file is physically stored. This may enable optimization of storage usage, server consolidation, and / or performance of non-disruptive file migration.

[0043] A hypervisor can provide hardware virtualization technology that allows multiple operating systems (e.g., "guest operating systems") to execute simultaneously on a host computer such as a computing resource. The hypervisor can present a virtual operating platform to the guest operating systems and can manage the execution of the guest operating systems. Multiple instances of various operating systems can share virtualized hardware resources.

[0044] The network includes one or more wired and / or wireless networks. For example, the network may include a cellular network (e.g., a fifth generation (5G) network, a long-term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, etc. and / or a combination of these or other types of networks.

[0045] Figure 1 The number and arrangement of devices and networks shown are provided as examples. Figure 1 The devices and / or networks shown may contain additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or differently arranged devices and / or networks. Figure 1 Two or more of the devices shown may be implemented in a single device, or Figure 1 The single device shown may be implemented as multiple distributed devices. Additionally or alternatively, one or more devices of one set of devices of an environment (eg, one or more devices) may perform one or more functions described as being performed by another set of devices of the environment.

[0046] Figure 2 yes Figure 1 A diagram of example components of one or more devices of the present invention. The device may correspond to a user device and / or a platform. Figure 2 As shown, the device may include a bus, a processor, a memory, a storage component, an input component, an output component, and a communication interface.

[0047] The bus includes components that allow communication between components of the device. The processor is implemented in hardware, firmware, or a combination of hardware and software. The processor is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some implementations, the processor includes one or more processors that can be programmed to perform a function. The memory includes a random access memory (RAM), a read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by the processor.

[0048] The storage component stores information and / or software related to the operation and use of the device. For example, the storage component may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid-state disk), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette, a magnetic tape, and / or another type of non-volatile computer-readable medium and a corresponding drive.

[0049] Input components include components that allow the device to receive information, such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, buttons, switches, and / or a microphone). Additionally or alternatively, the input components may include sensors for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, and / or an actuator). Output components include components that provide output information from the device (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs)).

[0050] The communication interface includes transceiver-like components (e.g., a transceiver and / or a separate receiver and transmitter) that enable a device to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. A communication interface may allow a device to receive information from another device and / or provide information to another device. For example, a communication interface may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.

[0051] The device may perform one or more of the processes described herein. The device may perform these processes in response to a processor executing software instructions stored by a non-volatile computer-readable medium (e.g., a memory and / or storage component). Computer-readable media is defined herein as a non-volatile memory device. A memory device includes memory space within a single physical storage device or memory space distributed on multiple physical storage devices.

[0052] The software instructions can be read into the memory and / or storage component from another computer-readable medium or from another device via a communication interface. When executed, the software instructions stored in the memory and / or storage component can cause the processor to perform one or more processes described herein. Additionally or alternatively, hard-wired circuits can be used to replace software instructions or to be combined with software instructions to perform one or more processes described herein. Therefore, the implementation described herein is not limited to any specific combination of hardware circuits and software.

[0053] Figure 2 The number and arrangement of components shown are provided as examples. Figure 2 Components shown, a device may include additional components, fewer components, different components or differently arranged components. Additionally or alternatively, a set of components (e.g., one or more components) of a device may perform one or more functions described as being performed by another set of components of the device.

[0054] The embodiments described herein include a training method based on a sequence-by-sequence loss function. Like other sequence-based training methods, this training method can be applied after training the model with a point-by-point cross entropy loss. This training method is not limited to speech recognition, but can also be applied to other sequence-to-sequence tasks.

[0055] Figure 3 is a flow chart of a method for attention-based E2E ASR training according to an embodiment. In some implementations, the platform may perform Figure 3 In some implementations, the processing may be performed by another device or group of devices (e.g., a user device) that is separate from or includes the platform. Figure 3 One or more processing blocks of .

[0056] like Figure 3 As shown, in operation 310, the method includes performing cross entropy training on the model based on one or more input features of the speech signal. Based on the cross entropy training being performed, the method includes performing large interval training on the cross entropy trained model in operations 320 to 360.

[0057] Large margin training refers to a sequence-level training criterion that enlarges the distance between the reference sequence and the most competitive incorrect sequence. The following equation (1) shows the distance between the reference sequence s and the most competitive incorrect sequence The interval and the loss function L() defined on the entire training dataset d:

[0058]

[0059] Referring to the following equation (2), the threshold Controls the desired distance between the reference sequence and the most competing incorrect sequence. To filter out samples that already meet the threshold constraint, a hinge function [·] is applied. + The loss function is squared so that it affects not only the sign but also the value of its gradient.

[0060] In operation 320 , the method includes performing a beam search using the cross-entropy trained model to generate an n-best hypothesis list having, for example, the highest a posteriori output hypotheses or sequences.

[0061] In operation 330 , the method includes determining a best hypothesis in the generated list of n-best hypotheses.

[0062] Unlike Hidden Markov Model (HMM) based systems, attention-based E2E ASR systems cannot generate a complete posterior graph due to the explicit dependencies between the output tokens. A common approximation of the posterior graph is a list of n-best hypotheses. In the list of n-best hypotheses, the best hypothesis is selected as the most competing incorrect hypothesis. Therefore, the loss function is as follows in Equation (2), where, represents the best hypothesis:

[0063]

[0064] In the attention-based E2E system, the log-posterior of the hypothesis can be interpreted as a score. Therefore, the above equation (2) can be written as the following equation (3), where score() is unnormalized by the sequence length:

[0065]

[0066] In operation 340 , the method includes determining a character-based gradient based on the cross-entropy trained model and a loss function that maximizes a distance between a reference sequence and the determined best hypothesis.

[0067] In operation 350 , the method includes determining a word-based gradient based on the cross-entropy trained model and a loss function that maximizes a distance between a reference sequence and the determined best hypothesis.

[0068] Threshold It can be chosen to be word and / or character level edit distance. Each of the following equations (4) gives the gradient at the top layer, where δ(·) is the Kronecker delta (indicator) function, and Represented by γ+:

[0069]

[0070] exist Repeated training on the correct beginning segment may make the training process unstable and prone to overfitting. To avoid training instability and overfitting, The first misclassified word in starts to apply the gradient in equation (4) above, which is shown in equation (5) as follows:

[0071]

[0072] By manually assigning gradients with respect to the log-posteriori, gradients of other parameters can be automatically derived by deep learning platforms such as PyTorch and Chainer.

[0073] In operation 360, the method includes performing back propagation on the model based on the determined character-based gradient and the determined word-based gradient to update the model. The method returns to operation 320 to continue large-margin training.

[0074] The method can be implemented with different loss functions. In the above embodiments, character-based and word-based loss functions are used. When applied to other sequence learning tasks such as translation, the loss function can be based on the bilingual evaluation understudy (BLEU) score, etc. Even if the goal of the method is to maximize the interval, multiple boundary examples can be applied. In addition to those described in the embodiments, there are various ways to select a set of hypotheses in which the hypothesis with the highest posteriori is beam searched. Finally, the gradient calculations on the hypotheses may also be different. Gradients in all positions or only gradients at special positions can be used.

[0075] although Figure 3 An example block diagram of the method is shown, but in some implementations, Figure 3 Compared to the blocks depicted in the method, the method may include additional blocks, fewer blocks, different blocks, or differently arranged blocks. Additionally or alternatively, two or more blocks in the blocks of the method may be executed in parallel.

[0076] The embodiments provide similar performance to that of the MWER method on the same benchmark dataset. The embodiments combine the large margin principle with deep neural network learning, which is easier to apply than the current SVM deep learning combination method.

[0077] Models according to embodiments are compared with the benchmark SWB300h dataset. The settings are the same as those in MWER, where the input is 40 dim log mel fbank features and the output is 49 characters. The E2E framework is fed with label attachment scores (LAS) as input to 6 bidirectional long short-term memories (BiLSTM) as encoders and 2 LSTMs as decoders. The baseline cross entropy model is the same as that of MWER, which is trained with the cross entropy criterion of the predetermined sampling authorization. The Adam optimizer is used for training, and the initial learning rate is 7.5*10 -7 The dropout rate is chosen to be 0.2 and the mini-batch size is 8. For large-margin training, results using multiple hypotheses are reported in addition to large-margin training based on the best hypothesis. The many hypotheses and multi-hypothesis large-margin criteria in MWER are the same. Results without using an external language model are denoted “w / o LM” and results using a language model are denoted “w / LM”.

[0078] Compared to the baseline WER of 13.3%, the results of large interval training based on the best hypothesis obtained a relative improvement of 6.8%, and the results of large interval training based on four best hypotheses obtained a relative improvement of 8.3%, as shown in Table 1:

[0079] Table 1. Results and comparisons of long-interval training

[0080]

[0081] Table 3 shows a comparison with other published results:

[0082] Table 3. Comparison with previously proposed attention-based end-to-end systems

[0083]

[0084] Figure 4 is a diagram of an apparatus for attention-based E2E ASR training according to an embodiment. Figure 4 As shown, the device includes a first execution code, a second execution code, a first determination code, a second determination code, a third determination code and a third execution code.

[0085] The first execution code is configured to perform cross entropy training on the model based on one or more input features of the speech signal.

[0086] The second execution code is configured to perform a beam search using the model to generate an n-best hypothesis list of output hypotheses.

[0087] The first determination code is configured to determine a best hypothesis in the generated n-best hypothesis list.

[0088] The second determining code and the third determining code are respectively configured to determine a character-based gradient and a word-based gradient based on the model and a loss function that maximizes a distance between a reference sequence and the determined best hypothesis.

[0089] The third execution code is configured to perform back propagation on the model based on the determined character-based gradient and the determined word-based gradient to update the model.

[0090] The output hypothesis in the list of n-best hypotheses may have the highest posterior in the output sequence of the model.

[0091] The second execution code may also be configured to perform the beam search again using the back-propagated model to regenerate the n-best hypothesis list.

[0092] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementations to the precise forms disclosed.Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations.

[0093] As used herein, the term "component" is intended to be broadly interpreted as hardware, firmware, or a combination of hardware and software.

[0094] It will be apparent that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software codes, and it should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0095] Even though particular combinations of features are defined in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features may be combined in ways not specifically defined in the claims and / or disclosed in the specification. Although each of the attached dependent claims may directly refer to only one claim, the disclosure of possible implementations includes each dependent claim in combination with every other claim in the claim group.

[0096] Unless explicitly described herein, the elements, actions or instructions used herein should not be interpreted as critical or necessary. Moreover, as used herein, the articles "a kind of" and "an" are intended to include one or more items, and can be used interchangeably with "one or more". In addition, as used herein, the term "group" is intended to include one or more items (e.g., related items, unrelated items, combinations of related items and unrelated items, etc.), and can be used interchangeably with "one or more". In the case of meaning only one item, the term "one" or similar language is used. Moreover, as used herein, the terms "have", "have", "contain" etc. are intended to be open terms. In addition, unless otherwise expressly stated, the phrase "based on" is intended to mean "based at least in part on".

Claims

1. A method for attention-based end-to-end (E2E) automatic speech recognition (ASR) training, characterized in that: The method comprises: performing cross entropy training on the model based on one or more input features of the speech signal; performing a beam search using the model to generate an n-best hypothesis list of output hypotheses; Determining the best hypothesis among the generated list of n best hypotheses; Determining a character-based gradient and a word-based gradient based on the model and a loss function that maximizes a distance between a reference sequence and the determined best hypothesis; and Back-propagation is performed on the model based on the determined character-based gradients and the determined word-based gradients to update the model.

2. The method according to claim 1, characterized in that The output hypothesis in the n-best hypothesis list has the highest posterior in the output sequence of the model.

3. The method according to claim 1, characterized in that The loss function is expressed as follows: Where (χ, s) represents the samples in the training set D, χ represents the input features, and s represents the reference samples. represents the best hypothesis, represents the distance between the reference sequence and the best hypothesis, and θ represents the model parameters.

4. The method according to claim 1, characterized in that: The loss function is expressed as follows: Where (χ, s) represents the samples in the training set D, χ represents the input features, and s represents the reference samples. represents the best hypothesis, represents the distance between the reference sequence and the best hypothesis, θ represents the model parameters, and score represents the log posterior.

5. The method according to claim 4, characterized in that Each of the character-based gradient and the word-based gradient is expressed as follows: where δ(·) represents the Kronecker delta function and γ + express 6. The method according to claim 4, characterized in that Each of the character-based gradient and the word-based gradient is expressed as follows: where δ(·) represents the Kronecker delta function, γ + express And ω represents the first wrong participle.

7. The method according to claim 1, characterized in that The method further includes: performing the beam search again using the model subjected to the back-propagation to regenerate the n-best hypothesis list.

8. An apparatus for attention-based end-to-end (E2E) automatic speech recognition (ASR) training, characterized in that The device comprises: at least one memory configured to store program code; and At least one processor, the at least one processor is configured to read the program code and operate according to the instructions of the program code, the program code comprising: a first execution code configured to cause the at least one processor to perform cross entropy training on a model based on one or more input features of a speech signal; second execution code configured to cause the at least one processor to perform a beam search using the model to generate an n-best hypothesis list of output hypotheses; first determining code configured to cause the at least one processor to determine a best hypothesis in the generated n-best hypothesis list; second determining code and third determining code configured to cause the at least one processor to determine a character-based gradient and a word-based gradient based on the model and a loss function that maximizes a distance between a reference sequence and the determined best hypothesis; and A third execution code is configured to cause the at least one processor to perform back propagation on the model based on the determined character-based gradient and the determined word-based gradient to update the model.

9. The device according to claim 8, characterized in that The output hypothesis in the n-best hypothesis list has the highest posterior in the output sequence of the model.

10. The method according to claim 8, characterized in that The loss function is expressed as follows: Where (χ, s) represents the samples in the training set D, χ represents the input features, and s represents the reference samples. represents the best hypothesis, represents the distance between the reference sequence and the best hypothesis, and θ represents the model parameters.

11. The device according to claim 8, characterized in that The loss function can also be expressed as follows: Where (χ, s) represents the samples in the training set D, χ represents the input features, and s represents the reference samples. represents the best hypothesis, represents the distance between the reference sequence and the best hypothesis, θ represents the model parameters, and score represents the log posterior.

12. The device according to claim 11, characterized in that Each of the character-based gradient and the word-based gradient is expressed as follows: where δ(·) represents the Kronecker delta function and γ + express 13. The device according to claim 11, characterized in that Each of the character-based gradient and the word-based gradient is expressed as follows: where δ(·) represents the Kronecker delta function, γ + express And ω represents the first wrong participle.

14. The device according to claim 8, characterized in that The beam search is performed again using the model that has been back-propagated to regenerate the n-best hypothesis list.

15. A non-transitory computer readable medium, characterized in that The computer readable medium stores instructions that, when executed by at least one processor of a device, cause the at least one processor to: performing cross entropy training on the model based on one or more input features of the speech signal; performing a beam search using the model to generate an n-best hypothesis list of output hypotheses; Determining the best hypothesis among the generated list of n best hypotheses; determining a character-based gradient and a word-based gradient based on the model and a loss function that maximizes a distance between a reference sequence and the determined best hypothesis; as well as Back-propagation is performed on the model based on the determined character-based gradients and the determined word-based gradients to update the model.

16. The non-transitory computer readable medium of claim 15, wherein: The output hypothesis in the n-best hypothesis list has the highest posterior in the output sequence of the model.

17. The method according to claim 15, characterized in that The loss function is expressed as follows: Where (χ, s) represents the samples in the training set D, χ represents the input features, and s represents the reference samples. represents the best hypothesis, represents the distance between the reference sequence and the best hypothesis, and θ represents the model parameters.

18. The non-transitory computer readable medium of claim 15, wherein: The loss function can also be expressed as follows: Where (χ, s) represents the sample in the training set D, χ represents the input feature, and s represents the reference sample. represents the best hypothesis, represents the distance between the reference sequence and the best hypothesis, θ represents the model parameters, and score represents the log posterior.

19. The non-transitory computer readable medium of claim 18, wherein: Each of the character-based gradient and the word-based gradient is expressed as follows: where δ(·) represents the Kronecker delta function and γ + express 20. The non-transitory computer readable medium of claim 18, wherein: Each of the character-based gradient and the word-based gradient is expressed as follows: where δ(·) represents the Kronecker delta function, γ + express And ω represents the first wrong participle.