Method, apparatus, server, and medium for calculating confidence of end-to-end speech

By extracting the acoustic features and recognition results of the speech recognition system, and combining the feature abstract model to calculate the confidence, the problem of strong coupling between the confidence module and the decoder and high storage resource consumption in the prior art is solved, and efficient and independent confidence calculation is achieved.

CN114005434BActive Publication Date: 2025-07-01BEIJING XIAOPENG AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111403940.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-24
Publication Date
2025-07-01
Estimated Expiration
2041-11-24

AI Technical Summary

Technical Problem

The confidence module of the existing speech recognition system is strongly coupled with the speech recognition decoder, resulting in the need to retrain the confidence module to adapt to different decoders, and the traditional solution consumes a lot of storage resources.

Method used

By extracting the acoustic features of the input audio and inputting them into the speech recognition decoder to obtain the recognition results, the confidence features are extracted in combination with the preset feature abstract model, and the confidence of the speech recognition result is directly calculated without relying on a specific decoder implementation.

Benefits of technology

It realizes independent optimization of confidence calculation, reduces error accumulation, reduces storage resource consumption, and has high practical value in actual business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114005434B_ABST
    Figure CN114005434B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, server, and medium for calculating the confidence of end-to-end speech in speech recognition. The recognition method includes: extracting acoustic features of each frame of input audio; inputting the acoustic features into a speech recognition decoder to obtain corresponding recognition results; extracting confidence features of each word in the recognition result according to the acoustic features, the recognition result, and a preset feature abstraction model; using the recognition result and the extracted confidence features as inputs to a confidence calculation model to predict the confidence of each word and the confidence of the sentence in the recognition result. The above method for calculating the confidence of end-to-end speech in speech recognition directly calculates the confidence of each word and sentence from acoustic features and recognition results. This confidence calculation scheme does not need to be adapted to and rely on the specific implementation of the speech recognition decoder, and has the advantages of independent optimization, high efficiency, and reduction of error accumulation, and has high practical value in actual business scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and particularly relates to a method, device, server and medium for calculating the confidence of end-to-end speech in speech recognition. Background Art

[0002] In the related art, the confidence module is a module that gives the credibility of the recognition result output by the speech recognition decoder. The recognition result combined with the confidence score is applied to downstream tasks such as dialogue systems, natural language understanding, keyword retrieval, etc. Confidence is of great significance for improving the accuracy of human-computer interaction.

[0003] The implementation of the confidence module in traditional speech recognition systems is generally calculated based on the decoded lattice graph, without the need for additional model and parameter training. In recent years, confidence algorithms based on end-to-end speech recognition systems have also emerged. Mainly, a subsequent model-based confidence module is trained using the recognition sequence generated by the decoder and the abstract features in the end-to-end acoustic model. This solution has a better precision-recall effect than the traditional lattice graph. However, the above two solutions have the following two problems:

[0004] 1) The confidence module is strongly dependent on the speech recognition decoder and has strong coupling. Especially for the model-based confidence scheme, when replacing different speech recognition decoders, different confidence modules need to be retrained to adapt.

[0005] 2) When training the confidence model after the traditional speech recognition system, a large amount of decoded results and acoustic features need to be saved, which consumes a large amount of storage resources and has poor practicability. The storage consumption is even greater in large-scale data training scenarios. Summary of the Invention

[0006] The present invention provides a method, device, server and medium for calculating the confidence of end-to-end speech in speech recognition.

[0007] A method for calculating the confidence of end-to-end speech in speech recognition according to the present invention includes:

[0008] Extracting the acoustic features of each frame of data of the input audio;

[0009] Inputting the acoustic features into a speech recognition decoder and obtaining the corresponding recognition result;

[0010] According to the acoustic features, the recognition result and a preset feature abstraction model, extracting the confidence features of each word in the recognition result;

[0011] Taking the recognition result and the extracted confidence features as the input of a confidence calculation model, and predicting the confidence of each word in the recognition result and the confidence of the sentence.

[0012] The above method for calculating the confidence of end-to-end speech in speech recognition directly calculates the confidence of each word and sentence from the acoustic features and the recognition result. This confidence calculation scheme does not need to adapt to and depend on the specific implementation of the speech recognition decoder, and has the advantages of independent optimization, high efficiency, and reduced error accumulation, and has high practical value in actual business scenarios.

[0013] According to the acoustic features, the recognition result, and a preset feature abstraction model, extract the confidence features of each word in the recognition result, including:

[0014] Pre-set a feature extraction model using the model structure of an encoder-decoder;

[0015] Train the feature extraction model;

[0016] Input the acoustic features into the encoder of the trained feature extraction model to abstract the original features;

[0017] Input the original features into the decoder of the trained feature extraction model to abstract the encoder features;

[0018] Input the original features and the recognition result into the decoder of the trained feature extraction model to abstract the decoder features.

[0019] In this way, the confidence features of each word can be obtained.

[0020] Inputting the original features into the decoder of the trained feature extraction model to abstract the encoder features includes:

[0021] Using the multi-head attention mechanism to abstract the encoder features from the original features in the decoder of the trained feature extraction model.

[0022] In this way, it can be realized that the original features output by the encoder are abstracted into encoder features in the decoder of the trained feature extraction model.

[0023] Taking the recognition result and the extracted confidence features as the input of the confidence calculation model, predicting the confidence of each word and the confidence of the sentence in the recognition result, including:

[0024] Taking the recognition result and the confidence features as the input, after feature splicing and position encoding, send them into the multi-layer Transformer Block module, and then one head generates the confidence of the word through Sigmoid, and the other head performs sentence-level abstraction through hierarchical attention and then sends it into Sigmoid to generate the confidence of the sentence.

[0025] In this way, the calculation of the confidence of words and the confidence of sentences can be realized.

[0026] The confidence calculation method includes the training stage of the confidence calculation model.

[0027] The training stage includes:

[0028] Using the recognition result and the confidence feature as inputs, the entire confidence calculation model is trained through the backpropagation algorithm.

[0029] In this way, the confidence calculation model can be trained.

[0030] Using the recognition result and the confidence feature as inputs, training the entire confidence calculation model through the backpropagation algorithm includes:

[0031] Using the recognition result and the confidence feature as inputs, through feature splicing and position encoding as the input of the confidence calculation model, and outputting the correct probability of words and the correct probability of sentences through the final Sigmoid layer;

[0032] Calculating the minimum edit distance through the correct transcription and the recognition result to obtain the word label and sentence label of the model;

[0033] Performing logistic regression loss modeling through the correct probability of words and sentences, the word label and sentence label, and training the entire confidence calculation model through the backpropagation algorithm.

[0034] In this way, the specific process of training can be realized.

[0035] The confidence calculation method includes the prediction stage of the confidence calculation model.

[0036] The prediction stage includes:

[0037] Using the recognition result and the confidence feature as inputs, through feature splicing and position encoding as the input of the confidence calculation model, feeding it into the trained confidence calculation model, outputting the correct probability of the recognized result words through one head, and outputting the correct probability of the sentence through another head for use in downstream tasks.

[0038] In this way, the calculation of the correct probability (confidence) of words and sentences can be realized.

[0039] An end-to-end speech confidence calculation device in speech recognition according to the present invention includes:

[0040] An acoustic feature extraction module for extracting acoustic features of each frame of data of the input audio.

[0041] An identification module, configured to input the acoustic features into a speech recognition decoder and obtain corresponding recognition results;

[0042] A confidence feature extraction module, configured to extract confidence features of each word in the recognition result according to the acoustic features, the recognition result, and a preset feature abstraction model; and

[0043] A confidence calculation module, configured to use the recognition result and the extracted confidence features as inputs to the confidence calculation model, and predict the confidence of each word and the confidence of the sentence in the recognition result.

[0044] A server according to the present invention includes the above-described confidence calculation device for end-to-end speech in speech recognition.

[0045] The present invention provides a non-volatile computer-readable storage medium storing computer-executable instructions, which, when executed by one or more processors, cause the processors to execute the above-described confidence calculation method for end-to-end speech in speech recognition.

[0046] The above-described confidence calculation device for end-to-end speech in speech recognition, server, and storage medium directly calculate the confidence of each word and sentence from acoustic features and recognition results. This confidence calculation solution does not need to be adapted to and depend on the specific implementation of the speech recognition decoder, and has the advantages of independent optimization, high efficiency, and reduction of error accumulation, and has high practical value in actual business scenarios.

[0047] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:

[0049] Figure 1 is a flowchart of the confidence calculation method for end-to-end speech in speech recognition according to an embodiment of the present invention;

[0050] Figure 2 is a block diagram of the confidence calculation device for end-to-end speech in speech recognition according to an embodiment of the present invention;

[0051] Figure 3 is a block diagram of the confidence feature extraction module according to an embodiment of the present invention;

[0052] Figure 4 is a block diagram of the confidence calculation module according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the embodiments of the present invention, and should not be construed as a limitation to the embodiments of the present invention.

[0054] The following disclosure provides many different embodiments or examples for implementing different structures of the embodiments of the present invention. To simplify the disclosure of the embodiments of the present invention, components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present invention. Embodiments of the present invention may repeat reference numerals and / or reference letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0055] Please refer to Figure 1 , a method for calculating the confidence of end-to-end speech in a speech recognition middle end provided by an embodiment of the present invention, includes:

[0056] Step 01, extracting acoustic features of each frame of data of the input audio;

[0057] Step 03, inputting the acoustic features into a speech recognition decoder and obtaining a corresponding recognition result;

[0058] Step 05, according to the acoustic features, the recognition result and a preset feature abstraction model, extracting the confidence features of each word in the recognition result;

[0059] Step 07, using the recognition result and the extracted confidence features as inputs of a confidence calculation model, and predicting the confidence of each word in the recognition result and the confidence of the sentence.

[0060] Please refer to Figure 2, the confidence calculation method for end-to-end speech in the above-mentioned embodiment of speech recognition can be implemented by the confidence calculation device 100 for end-to-end speech in the embodiment of the present invention. Specifically, a confidence calculation device 100 for end-to-end speech in an embodiment of the present invention includes an acoustic feature extraction module 11, a recognition module 13, a confidence feature extraction module 15, and a confidence calculation module 17. The acoustic feature extraction module 11 is configured to extract acoustic features of each frame of data of the input audio. The recognition module 13 is configured to input the acoustic features into a speech recognition decoder and obtain a corresponding recognition result. The confidence feature extraction module 15 is configured to extract confidence features of each word in the recognition result according to the acoustic features, the recognition result, and a preset feature abstraction model. The confidence calculation module 17 is configured to use the recognition result and the extracted confidence features as inputs to a confidence calculation model, and predict the confidence of each word in the recognition result and the confidence of the sentence.

[0061] The above-mentioned confidence calculation method for end-to-end speech in speech recognition and the confidence calculation device 100 for end-to-end speech in speech recognition directly calculate the confidence of each word and sentence from the acoustic features and the recognition result. This confidence calculation scheme does not need to adapt to and depend on the specific implementation of the speech recognition decoder, and has the advantages of independent optimization, high efficiency, and reduction of error accumulation, and has high practical value in actual business scenarios.

[0062] Specifically, in the embodiment of the present invention, in analyzing the problems of strong coupling between the confidence calculation scheme and the recognition decoder in the related technology and the difficulty of adapting deep learning confidence to traditional decoders, the above-mentioned end-to-end speech confidence recognition strategy independent of the speech recognition decoder is proposed. The confidence of each word and the confidence of the sentence can be applied to downstream tasks such as dialogue systems, natural language understanding, keyword retrieval, etc.

[0063] The input audio can be obtained by the first terminal to which the calculation method is applied, or can be obtained by a second terminal communicating with the first terminal and then transmitted to the first terminal. The user can input speech through the first terminal or the second terminal to generate the input audio. The first terminal and the second terminal include but are not limited to mobile phones, tablet computers, in-vehicle terminal devices, servers, etc.

[0064] Extracting the acoustic features of each frame of data of the input audio can generate an acoustic feature frame sequence. Specifically, the method for extracting the acoustic features of each frame of data of the input audio can refer to the methods in the field of speech processing in the related technology, and will not be elaborated in detail here. In this embodiment, on the one hand, the acoustic features are input to the speech recognition decoder to obtain a recognition result, and on the other hand, they can be used as inputs for confidence extraction to extract confidence features and calculate the correct probability independently of the speech recognition decoder.

[0065] In some embodiments, the speech recognition decoder includes a decoder based on an HMM speech recognition system and a decoder based on an end-to-end speech recognition system. In this way, the relevant speech recognition system can be flexibly used to obtain the recognition result.

[0066] Specifically, in one embodiment, the HMM (Hidden Markov Model) speech recognition system may include a DNN-HMM acoustic model + an Ngram language model.

[0067] In one embodiment, the end-to-end speech recognition system may include a Conformer-LSTM RNNT model, etc. The recognition result trusted by this embodiment is obtained by feeding the acoustic features into the decoder.

[0068] It can be understood that in other embodiments, other types of speech recognition decoders may also be used to obtain the recognition result, not limited to the decoder based on the HMM speech recognition system and the decoder based on the end-to-end speech recognition system.

[0069] In some embodiments, step 05 includes:

[0070] Preset a feature extraction model adopting an encoder-decoder model structure;

[0071] Train the feature extraction model;

[0072] Input the acoustic features into the encoder of the trained feature extraction model to abstract the original features;

[0073] Input the original features into the decoder of the trained feature extraction model to abstract the encoder features;

[0074] Input the original features and the recognition result into the decoder of the trained feature extraction model to abstract the decoder features.

[0075] Please refer Figure 2 , the confidence calculation method for end-to-end speech in the above-mentioned speech recognition embodiment can be implemented by the confidence calculation device 100 for end-to-end speech in the speech recognition embodiment of the present invention. Specifically, the confidence feature extraction module 15 is used for: presetting a feature extraction model adopting an encoder-decoder model structure; training the feature extraction model; inputting the acoustic features into the encoder of the trained feature extraction model to abstract the original features; inputting the original features into the decoder of the trained feature extraction model to abstract the encoder features; inputting the recognition result into the decoder of the trained feature extraction model to abstract the decoder features.

[0076] In this way, the confidence features of each word can be obtained.

[0077] Specifically, to solve the problem of strong dependence and strong coupling between the confidence calculation model and the speech recognition decoder in related technologies, the recognition method according to the embodiments of the present invention does not obtain any feature information from the decoder, but directly extracts necessary confidence features from the audio acoustic feature frames.

[0078] In one embodiment, please refer Figure 3 to. The confidence feature extraction module 15 as a whole can adopt the model structure of an encoder-decoder. The acoustic features are fed into the encoder of the trained feature extraction model to abstract the original features (such as high-dimensional features). The original features output by the encoder of the trained feature extraction model are then input into the decoder of the trained feature extraction model for processing to abstract the encoder features. The recognition results are fed into the decoder of the trained feature extraction model to abstract the decoder features (such as high-dimensional features).

[0079] In this embodiment, the confidence feature extraction module 15 can be divided into two stages: training and prediction.

[0080] Training stage: The main problem in this stage is to design a loss function that can abstract the confidence features.

[0081] The cross-entropy loss with text as the label for the speech recognition classification task can be adopted. At this time, the decoder output performs the softmax classification task.

[0082] The mean squared error loss with masked acoustic features as the label for the speech restoration pre-training task can also be adopted. At this time, the encoder output performs the mask regression task, and the decoder output performs the softmax classification task simultaneously.

[0083] Prediction stage: The encoder features finally output by the trained confidence feature extraction module 15 are obtained by the decoder of the trained feature extraction model performing multi-head attention processing on the original features output by the encoder of the trained feature extraction model, and the decoder features are directly output by the decoder of the trained feature extraction model.

[0084] In some embodiments, the encoder of the preset feature abstraction model is composed of a convolutional layer and multiple ConformerBlocks;

[0085] The decoder of the preset feature abstraction model is composed of multiple Transformer Decoder Blocks. In this way, the model structure of an encoder-decoder can be realized.

[0086] Specifically, the Conformer Block is cascaded by a normalization layer, a feed-forward layer, a multi-head attention layer, a convolutional layer, and a feed-forward layer. The Transformer Decoder Block is cascaded by a multi-head attention layer, a feed-forward layer, a multi-head attention layer, a feed-forward layer, and a normalization layer.

[0087] In some embodiments, abstracting encoder features from the original features by inputting them into the decoder of the trained feature extraction model includes:

[0088] Abstracting encoder features from the original features in the decoder of the trained feature extraction model through the multi-head attention mechanism.

[0089] Please refer to Figure 2 , the confidence calculation method for end-to-end speech in the speech recognition of the above embodiments can be implemented by the confidence calculation device 100 for end-to-end speech in the speech recognition of the embodiments of the present invention. Specifically, the confidence feature extraction module 15 is used to abstract encoder features from the original features in the decoder of the trained feature extraction model through the multi-head attention mechanism.

[0090] In this way, it can be realized that the original features output by the encoder are abstracted into encoder features in the decoder of the trained feature extraction model.

[0091] Specifically, please combine with Figure 3 , the Transformer Decoder Block of the decoder of the trained feature extraction model has a multi-head attention layer, and the original features output by the encoder of the trained feature extraction model are sent to the multi-head attention layer for processing to obtain encoder features.

[0092] In some embodiments, step 07 includes:

[0093] Taking the recognition result and the confidence feature as inputs, after feature splicing and positional encoding, they are sent into the multi-layer Transformer Block module. Then, one head generates the confidence of the word through Sigmoid, and the other head performs sentence-level abstraction through hierarchical attention and then is sent into Sigmoid to generate the confidence of the sentence.

[0094] Please refer to Figure 2, the confidence calculation method for end-to-end speech in the speech recognition of the above embodiments can be implemented by the confidence calculation device 100 for end-to-end speech in the speech recognition of the embodiments of the present invention. Specifically, the confidence calculation module 17 is used to take the recognition result and confidence features as inputs. After feature splicing and positional encoding, they are sent into the multi-layer Transformer Block module. Then, one head generates the confidence of words through Sigmoid, and the other head performs sentence-level abstraction through hierarchical attention and then is sent into Sigmoid to generate the confidence of the sentence.

[0095] In this way, the calculation of the confidence of words and the confidence of sentences can be achieved.

[0096] Specifically, taking the confidence features of each word in the recognition result obtained above as the input of the confidence calculation module 17, the confidence (correct probability) of each word and sentence is calculated.

[0097] In this embodiment, please combine Figure 4 , the confidence calculation module 17 can adopt the structure of a Transformer encoder. Taking the recognition result of the speech recognition decoder, the encoder features and decoder features extracted by the confidence feature extraction module as inputs, after feature splicing (concatenate link) and positional encoding, they are sent into the multi-layer Transformer Block module. Then, one head generates the correct probability of the word confidence through Sigmoid, and the other head performs sentence-level abstraction through hierarchical attention and then is sent into Sigmoid to generate the correct probability of the sentence confidence.

[0098] In some embodiments, the confidence calculation method includes the training stage of the confidence calculation model,

[0099] The training stage includes:

[0100] Taking the recognition result and confidence features as inputs, the entire confidence calculation model is trained through the backpropagation algorithm.

[0101] Please refer to Figure 2 , the confidence calculation method for end-to-end speech in the speech recognition of the above embodiments can be implemented by the confidence calculation device 100 for end-to-end speech in the speech recognition of the embodiments of the present invention. Specifically, the confidence calculation module 17 can have a training stage.

[0102] In this way, the confidence calculation model can be trained.

[0103] Specifically, in one embodiment, taking the recognition result and confidence features as inputs, training the entire confidence calculation model through the backpropagation algorithm includes:

[0104] Taking the recognition result and confidence feature as inputs, through feature splicing and position encoding as the input of the confidence calculation model, and outputting the correct word probability and correct sentence probability through the final Sigmoid layer;

[0105] Calculating the minimum edit distance through the correct transcription and recognition result to obtain the word label and sentence label of the model;

[0106] Conducting logistic regression loss modeling through the correct word probability, correct sentence probability, word label, and sentence label, and training the entire confidence calculation model through the backpropagation algorithm. In this way, the specific process of training can be realized.

[0107] Specifically, the method for obtaining the word label and sentence label of the model is expanded as follows:

[0108] For word correctness discrimination, the correct transcription can be aligned to each word of the recognition result through the minimum edit distance to obtain a 0-1 label for each word, and then trained through logistic regression. In an example, this alignment method can be represented by the following table:

[0109] Recognition result Dumping Dumping Percent Percent Of Five + Correct transcription Cancel Cancel Percent Percent Of + 0-1 label 0 (replacement) 0 (replacement) 1 1 1 0 (insertion) 1

[0110] For sentence correctness discrimination, it can be obtained by calculating the CER of the correct transcription and recognition result. When the CER is 0, the label is 1 (1 indicates that the sentence is correct), otherwise it is 0 (0 indicates that the sentence is incorrect), and then trained through logistic regression.

[0111] It can be understood that the label can also be represented by other numbers or symbols, not limited to 0 and 1.

[0112] In some embodiments, the confidence calculation method includes the prediction stage of the confidence calculation model,

[0113] The prediction stage includes:

[0114] Taking the recognition result and confidence feature as inputs, through feature splicing and position encoding as the input of the confidence calculation model, sending it into the trained confidence calculation model, outputting the correct probability of the recognition result word through one head, and outputting the correct probability of the sentence through another head for use in downstream tasks. In this way, the correct probability (confidence) calculation of words and sentences can be realized.

[0115] Specifically, downstream tasks include but are not limited to dialogue systems, natural language understanding, keyword retrieval, etc.

[0116] A server according to an embodiment of the present invention includes the confidence calculation device 100 for end-to-end speech of the speech recognition in the above embodiment.

[0117] The above-mentioned server directly calculates the confidence of each word and sentence from the acoustic features and recognition results. This confidence calculation scheme does not need to adapt to and depend on the specific implementation of the speech recognition decoder, and has the advantages of independent optimization, high efficiency, and reduction of error accumulation, and has high practical value in actual business scenarios.

[0118] Specifically, the input audio can be collected by the microphone of the vehicle communicating with the server and uploaded to the server by the vehicle, or can be collected by the server itself, or the user can directly input an audio file, which is not specifically limited here. Vehicles include but are not limited to fuel vehicles, range-extended electric vehicles, pure electric vehicles, hybrid vehicles, hydrogen energy vehicles, etc.

[0119] The embodiment of the present invention also provides a non-volatile computer-readable storage medium storing computer-executable instructions, which, when executed by one or more processors, cause the processors to execute the confidence calculation method for end-to-end speech in the speech recognition in any of the above embodiments.

[0120] Specifically, in one embodiment, when the computer-executable instructions are executed by the processor, the confidence calculation method for end-to-end speech in the speech recognition includes:

[0121] Step 01, extract the acoustic features of each frame of data of the input audio;

[0122] Step 03, input the acoustic features into the speech recognition decoder and obtain the corresponding recognition results;

[0123] Step 05, according to the acoustic features, recognition results and a preset feature abstraction model, extract the confidence features of each word in the recognition results;

[0124] Step 07, take the recognition results and the extracted confidence features as the input of the confidence calculation model, and predict the confidence of each word and the confidence of the sentence in the recognition results.

[0125] It can be understood that the above explanations of the embodiments and beneficial effects of the confidence calculation method for end-to-end speech in the speech recognition also apply to the computer-readable storage medium of the embodiment of the present invention. To avoid redundancy, no detailed expansion is made here.

[0126] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0127] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0128] Any process or method description shown in a flowchart or described in other ways herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a manner that is not shown or discussed in order, including in a substantially simultaneous manner according to the functions involved or in a reverse order, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0129] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for calculating the confidence of end-to-end speech in speech recognition, characterized in that, Including: Extracting the acoustic features of each frame of the input audio; Inputting the acoustic features into a speech recognition decoder and obtaining the corresponding recognition result; According to the acoustic features, the recognition result, and a preset feature abstraction model, extracting the confidence features of each word in the recognition result; Using the recognition result and the extracted confidence features as the input of a confidence calculation model to predict the confidence of each word in the recognition result and the confidence of the sentence; According to the acoustic features, the recognition result, and a preset feature abstraction model, extracting the confidence features of each word in the recognition result, including: Pre-setting a feature abstraction model adopting an encoder-decoder model structure; Training the feature abstraction model; Inputting the acoustic features into the encoder of the trained feature abstraction model to abstract the original features; Inputting the original features into the decoder of the trained feature abstraction model to abstract the encoder features; Inputting the original features and the recognition result into the decoder of the trained feature abstraction model to abstract the decoder features, where the confidence features include the encoder features and the decoder features.

2. The confidence calculation method for end-to-end speech in speech recognition according to claim 1, characterized in that, Inputting the original features into the decoder of the trained feature abstraction model to abstract the encoder features, including: Using the multi-head attention mechanism to make the original features abstract the encoder features in the decoder of the trained feature abstraction model.

3. The confidence calculation method for end-to-end speech in speech recognition according to claim 1, wherein Using the recognition result and the extracted confidence features as the input of a confidence calculation model to predict the confidence of each word in the recognition result and the confidence of the sentence, including: Taking the recognition result and the confidence features as the input, after feature splicing and position encoding, sending them into a multi-layer Transformer Block module, and then one head generates the confidence of the word through Sigmoid, and the other head performs sentence-level abstraction through hierarchical attention and then sends it into Sigmoid to generate the confidence of the sentence.

4. The confidence calculation method for end-to-end speech in speech recognition according to claim 3, characterized in that, The confidence calculation method includes the training stage of the confidence calculation model, The training stage includes: Taking the recognition result and the confidence features as the input and training the entire confidence calculation model through the backpropagation algorithm.

5. The confidence calculation method for end-to-end speech in speech recognition according to claim 4, characterized in that Taking the recognition result and the confidence features as the input and training the entire confidence calculation model through the backpropagation algorithm, including: Taking the recognition result and the confidence features as the input, using feature splicing and position encoding as the input of the confidence calculation model, and outputting the word correct probability and the sentence correct probability through the final Sigmoid layer; Calculating the minimum edit distance through the correct transcription and the recognition result to obtain the word label and the sentence label of the model; Performing logistic regression loss modeling through the word correct probability and the sentence correct probability, the word label and the sentence label, and training the entire confidence calculation model through the backpropagation algorithm.

6. The confidence calculation method for end-to-end speech in speech recognition according to claim 4, wherein The confidence calculation method includes the prediction stage of the confidence calculation model, The prediction stage includes: Taking the recognition result and the confidence feature as inputs, through feature splicing and positional encoding as the inputs of the confidence calculation model, and feeding them into the trained confidence calculation model. One head outputs the correct probability of the recognized result word, and the other head outputs the correct probability of the sentence for use in downstream tasks.

7. An end-to-end speech confidence calculation device for speech recognition, characterized in that, It includes: An acoustic feature extraction module for extracting the acoustic features of each frame of the input audio; A recognition module for inputting the acoustic features into a speech recognition decoder and obtaining the corresponding recognition result; A confidence feature extraction module for extracting the confidence feature of each word in the recognition result according to the acoustic features, the recognition result, and a preset feature abstraction model; and A confidence calculation module for taking the recognition result and the extracted confidence features as the inputs of the confidence calculation model, and predicting the confidence of each word and the confidence of the sentence in the recognition result; The confidence feature extraction module is further configured to: Pre-set a feature abstraction model adopting an encoder-decoder model structure; Train the feature abstraction model; Input the acoustic features into the encoder of the trained feature abstraction model to abstract the original features; Input the original features into the decoder of the trained feature abstraction model to abstract the encoder features; Input the original features and the recognition result into the decoder of the trained feature abstraction model to abstract the decoder features, and the confidence features include the encoder features and the decoder features.

8. A server, characterized in that, It includes the confidence calculation device for end-to-end speech in speech recognition described in claim 7.

9. A non-volatile computer-readable storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed by one or more processors, the processors are caused to execute the confidence calculation method for end-to-end speech in speech recognition described in any one of claims 1-6.

Citation Information

Patent Citations

  • End-to-end speech emotion recognition method and system

    CN110097894A

  • Speech recognition method and system, medium, computer equipment, terminal and application

    CN112712804A