Speech recognition method and device, related equipment and computer program product

By introducing shared encoder and multiple decoders into the speech recognition model, a unified solution for streaming and non-streaming speech recognition is realized, solving the problems of long model development time and waste of resources in the prior art, and improving the convenience of use.

CN120220663APending Publication Date: 2025-06-27HEFEI IFLY DIGITAL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510383600.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, separate models need to be developed for different speech recognition scenarios, resulting in long development time, waste of resources and inconvenient use.

Method used

A model that integrates streaming and non-streaming speech recognition capabilities is proposed. Through shared encoder and multiple types of decoders, it is suitable for streaming and non-streaming recognition tasks, reducing model development cycle and resource consumption.

Benefits of technology

It realizes the adaptation of streaming and non-streaming recognition tasks through a unified speech recognition model, reducing the model development cycle and resource consumption, and improving the convenience of use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220663A_ABST
    Figure CN120220663A_ABST
Patent Text Reader

Abstract

The invention discloses a speech recognition method and device, related equipment and a computer program product, and an adopted speech recognition model comprises a shared encoder, a first-class decoder supporting streaming decoding processing and a second-class decoder supporting non-streaming decoding processing. The shared encoder employs a coding network that simultaneously supports streaming and non-streaming identification tasks. The speech to be recognized can be a speech data stream composed of audio blocks in a streaming recognition task or a whole speech segment in a non-streaming recognition task, the acoustic features of the speech to be recognized are coded by the shared encoder, the coded features are sent to the target decoder to be decoded, and a final speech recognition result is obtained based on a decoding result. A target decoder in the streaming recognition task is a first-class decoder, and a target decoder in the non-streaming recognition task is a second-class decoder. Through the unified speech recognition model, the speech recognition method and device can adapt to streaming recognition tasks and non-streaming recognition tasks, and use convenience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech recognition, and more specifically, to a speech recognition method, apparatus, related device, and computer program product. Background Art

[0002] Speech recognition (Automatic Speech Recognition, ASR) includes non-streaming speech recognition and streaming speech recognition (also known as real-time speech recognition or online speech recognition). Generally, the so-called speech recognition mostly refers to non-streaming speech recognition.

[0003] With the rapid development of artificial intelligence technology, various wearable and portable intelligent devices equipped with various application software have been fully integrated into people's lives. There is a large demand for streaming speech recognition in a series of speech interaction scenarios such as common input methods, online meetings, live broadcasts, and real-time translations.

[0004] Currently, speech recognition models are generally developed separately for different speech recognition scenarios. For example, for streaming speech recognition tasks, a model supporting streaming speech recognition is developed. For non-streaming speech recognition tasks, a model supporting non-streaming speech recognition is developed. Developing different speech recognition models for different task scenarios will have problems such as long development time and wasted resources. And it is necessary to call different models for speech recognition processing according to different task scenarios, which is inconvenient to use. Summary of the Invention

[0005] In view of the above problems, the present application is proposed to provide a speech recognition method, apparatus, related device, and computer program product, so as to provide a model integrating streaming and non-streaming speech recognition capabilities, which can be applied to streaming and non-streaming recognition tasks, reduce the model development cycle and resources, and improve the convenience of use. The specific solutions are as follows:

[0006] In a first aspect, a speech recognition method is provided, including:

[0007] Extracting acoustic features of the speech to be recognized, where the speech to be recognized is a speech data stream composed of audio chunks in a streaming recognition task, or a whole speech in a non-streaming recognition task;

[0008] Feeding the acoustic features into a speech recognition model, encoding the acoustic features through a shared encoder in the model to obtain encoded features, decoding the encoded features through a target decoder in the model, and obtaining a speech recognition result based on the decoding result;

[0009] Among them, the shared encoder is an encoding network that supports both streaming and non-streaming recognition tasks. The speech recognition model includes a first type of decoder that supports streaming decoding processing and a second type of decoder that supports non-streaming decoding processing. In the streaming recognition task, the target decoder is the first type of decoder, and in the non-streaming recognition task, the target decoder at least includes the second type of decoder.

[0010] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the process of decoding the encoded features by the target decoder in the model and obtaining the speech recognition result based on the decoding result includes:

[0011] In the streaming recognition task:

[0012] Decoding the encoded features by any one of the first type of decoders in the model, and taking the decoding result as the streaming speech recognition result;

[0013] Or, decoding the encoded features by two or more of the first type of decoders in the model respectively to obtain two or more decoding results, and selecting the decoding result with a higher probability score as the streaming speech recognition result based on the two or more decoding results.

[0014] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the process of decoding the encoded features by the target decoder in the model and obtaining the speech recognition result based on the decoding result includes:

[0015] In the non-streaming recognition task:

[0016] Decoding the encoded features by any one of the second type of decoders in the model, and taking the decoding result as the non-streaming speech recognition result;

[0017] Or, decoding the encoded features by two or more of the second type of decoders in the model respectively to obtain two or more decoding results, and selecting the decoding result with a higher probability score as the non-streaming speech recognition result based on the two or more decoding results.

[0018] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the first type of decoder supports both streaming and non-streaming decoding processing, and the second type of decoder only supports non-streaming decoding processing;

[0019] The process of decoding the encoded features by the target decoder in the model and obtaining the speech recognition result based on the decoding result includes:

[0020] In the non-streaming recognition task:

[0021] Decode the encoded features through one of the first - type decoders in the model to obtain a preliminary decoding result, re - score the preliminary decoding result through the second - type decoder, and determine the non - streaming speech recognition result based on the re - scoring result;

[0022] Or,

[0023] Decode the encoded features through more than two of the first - type decoders in the model respectively to obtain more than two preliminary decoding results, select a preliminary decoding result with a higher probability score, and re - score the preliminary decoding result with a higher probability score through the second - type decoder, and determine the non - streaming speech recognition result based on the re - scoring result.

[0024] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the speech recognition model is trained in a multi - language and multi - task manner during the training phase;

[0025] The multi - tasks include a speech recognition task and at least one of the following types of tasks:

[0026] Punctuation prediction PPM task, language identification LID task, valid speech detection VAD task.

[0027] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the first - type decoder includes a CTC decoder and / or an RNN - T decoder; the second - type decoder includes an AED decoder.

[0028] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the shared encoder adopts a Convolution Transformer model Conformer, and the Conformer model internally adopts a causal convolutional structure and an attention mechanism Attention module based on the chunk form to limit the scope of action.

[0029] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the Conformer model includes a series of down - sampling modules, and the series of down - sampling modules perform progressive down - sampling operations.

[0030] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the Conformer model includes a grouped multi - head self - attention module;

[0031] When the grouped multi-head self-attention module calculates, it transforms the dimensions of the query Q, key K, and value V in the traditional multi-head self-attention mechanism from (n, d) to (n / g, d×g), then performs attention calculation, and after calculation, transforms the dimensions back to the original (n, d), where g is the group size.

[0032] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the speech recognition model is trained in the following manner:

[0033] Obtain the training data of the current batch, and allocate the non-streaming sample speech and streaming sample speech in the training data according to a set ratio;

[0034] Extract the acoustic features of the sample speech in the training data, and send them into the speech recognition model for processing to obtain the decoding results output by the first type of decoder and the second type of decoder;

[0035] Based on the decoding result output by each decoder and the recognition result label corresponding to the sample speech in the training data, calculate the loss of each decoder;

[0036] Perform weighted fusion on the losses of each decoder to obtain the total loss, and update the model parameters through the backpropagation algorithm according to the total loss until the set training end condition is reached.

[0037] In a second aspect, a speech recognition device is provided, including:

[0038] An acoustic feature extraction unit, configured to extract the acoustic features of the speech to be recognized, where the speech to be recognized is a speech data stream composed of audio chunks in a streaming recognition task, or an entire speech in a non-streaming recognition task;

[0039] A speech recognition processing unit, configured to send the acoustic features into the speech recognition model, encode the acoustic features through a shared encoder in the model to obtain encoded features, decode the encoded features through a target decoder in the model, and obtain a speech recognition result based on the decoding result;

[0040] Wherein, the shared encoder is an encoding network that supports both streaming and non-streaming recognition tasks, the speech recognition model includes a first type of decoder that supports streaming decoding processing and a second type of decoder that supports non-streaming decoding processing, and in the streaming recognition task, the target decoder is the first type of decoder, and in the non-streaming recognition task, the target decoder at least includes the second type of decoder.

[0041] In a third aspect, an electronic device is provided, including: a memory and a processor;

[0042] The memory is used to store programs;

[0043] The processor is used to execute the program to implement the steps of the speech recognition method described in any one of the foregoing first aspects of the present application.

[0044] In a fourth aspect, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the speech recognition method described in any one of the foregoing first aspects of the present application are implemented.

[0045] In a fifth aspect, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the speech recognition method described in any one of the foregoing first aspects of the present application are implemented.

[0046] By means of the above technical solution, the speech recognition method of the present application is implemented through an end-to-end unified speech recognition model. The speech recognition model integrates streaming and non-streaming speech recognition capabilities and can be applicable to a variety of different scenarios through a single speech recognition model. Specifically, the speech recognition model includes a shared encoder and two types of decoders. The shared encoder uses an encoding network that supports both streaming and non-streaming recognition tasks to meet the encoding requirements for the acoustic features of different streaming and non-streaming speech data streams. The model integrates a first type of decoder that supports streaming decoding processing and a second type of decoder that supports non-streaming decoding processing. For the speech to be recognized, it can be a speech data stream composed of audio chunks in a streaming recognition task or an entire speech in a non-streaming recognition task. After extracting the acoustic features of the speech to be recognized, they are sent into the speech recognition model, encoded by the shared encoder, and the encoded features are sent into the target decoder for decoding. Based on the decoding result, the final speech recognition result is obtained. In a streaming recognition task, the target decoder is the first type of decoder, and in a non-streaming recognition task, the target decoder is the second type of decoder. The solution of the present application ensures that through a unified speech recognition model, it can be adapted to streaming recognition tasks and non-streaming recognition tasks, thereby reducing the model development cycle and development resources, and through the unified speech recognition model, it can be adapted to a variety of task scenarios without switching different models, improving the convenience of use. Description of the Drawings

[0047] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0048] Figure 1 It is a schematic diagram of an implementation system architecture of the speech recognition method provided by an embodiment of the present application;

[0049] Figure 2 Schematic diagram of a speech recognition model structure provided by an embodiment of the present application;

[0050] Figure 3 Schematic diagram of a decoding strategy of a decoder network under a streaming recognition task provided by an embodiment of the present application;

[0051] Figure 4 Schematic diagram of a decoding strategy of a decoder network under a non-streaming recognition task provided by an embodiment of the present application;

[0052] Figure 5 Schematic diagram of a Conformer structure provided by an embodiment of the present application;

[0053] Figure 6 Illustrates Figure 5 A schematic diagram of a structure of a convolutional downsampling block in

[0054] Figure 7 Schematic diagram of a grouped multi-head self-attention mechanism provided by an embodiment of the present application;

[0055] Figure 8 Schematic diagram of a speech recognition model training process provided by an embodiment of the present application;

[0056] Figure 9 Block diagram of a multi-task, multi-language training format encoding provided by an embodiment of the present application;

[0057] Figure 10a Illustrates an encoding format without timestamps;

[0058] Figure 10b Illustrates an encoding format with timestamps;

[0059] Figure 11 Schematic diagram of a speech recognition device structure provided by an embodiment of the present application;

[0060] Figure 12 Schematic diagram of a structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0061] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0062] The present application provides a speech recognition method, which can be applied to a system architecture as shown in Figure 1 . The system may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 taking one server as an example for illustration).

[0063] Either the terminal 100 or the server 200 can be used alone to execute the speech recognition method provided in the embodiments of the present application. In addition, the terminal 100 and the server 200 can also be used in cooperation to execute the speech recognition method provided in the embodiments of the present application.

[0064] Next, the product form of the terminal 100 will be described Figure 1 in

[0065] The terminal 100 in the embodiments of the present application can be a mobile phone, a tablet computer, a translator, a learning machine, a wearable device, a vehicle-mounted device, a conference terminal, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiments of the present application do not make any restrictions on this.

[0066] The speech recognition method provided by the present application, through an end-to-end unified speech recognition model, can be applicable to a variety of different task scenarios at the same time. For example, it can be applicable to various types of streaming recognition task scenarios (such as online meetings, live broadcasts, real-time translations, etc.), and can also be applicable to various types of non-streaming recognition task scenarios.

[0067] The embodiments of the present application provide a speech recognition method. Taking the application of this method to a computer device as an example, the computer device can specifically be Figure 1 the terminal 100 in Figure 2 or a system composed of the terminal 100 and the server 200. Combining

[0068] Step S100: Extract the acoustic features of the speech to be recognized.

[0069] Wherein, the speech to be recognized is a speech data stream composed of audio chunks in a streaming recognition task, or an entire speech in a non-streaming recognition task.

[0070] That is to say, the speech recognition method of the present application supports both streaming recognition tasks and non-streaming recognition tasks.

[0071] When the speech recognition method is applied to a streaming recognition task, the speech to be recognized is a chunk speech data stream, that is, at this time, a streaming inference decoding method is adopted chunk by chunk. When extracting acoustic features, the acoustic features of the speech data stream of the set decoding window decoder_window can be extracted. Among them, decoder_window is related to the chunk size, and the chunk size is the number of audio frames included in a single chunk. Exemplarily, decode_window = (chunk - 1) × 2 + 2 + 1.

[0072] The following example describes the chunk-by-chunk decoding process with a 15-second speech of 1500 frames and chunk = 12, that is, decode_window = 25:

[0073] current = 0, end = 25

[0074] current = 24, end = 29

[0075] ……

[0076] current = 1464, end = 1489

[0077] current = 1488, end = 1500

[0078] It can be seen that for a 1500-frame speech data stream with a chunk size of 12 and a decoding window decode_windows of 25, it is divided into 63 chunks for decoding, and the last window is less than 25, which is 12 frames. There can be an overlap of 2 frames (or other numbers of frames) between the decoding windows, which can improve the recognition effect while ensuring the latency.

[0079] Optionally, in the streaming recognition task, the context information of the historical chunk speech data stream can also be added based on the current chunk information for decoding, so as to obtain a better streaming recognition effect.

[0080] When the speech recognition method is applied to a non-streaming recognition task, the speech to be recognized is a whole speech. The acoustic features of the whole speech can be extracted.

[0081] In this step, when extracting the acoustic features of the speech to be recognized, the speech to be recognized can be preprocessed first. The preprocessing process includes but is not limited to: processing the speech to be recognized into a fixed frequency, fixed sampling rate, fixed number of channels, and fixed coding format. For example, the speech to be recognized is uniformly processed into a format of 16000Hz, 16bit, single channel, and PCM encoding.

[0082] In this step, the process of extracting the acoustic features of the speech to be recognized can be implemented by a configured feature extractor. The extracted acoustic features include, but are not limited to, filter bank features Fbank, Mel spectrum features Mel, etc.

[0083] Step S110: Send the acoustic features into the speech recognition model, encode the acoustic features through the shared encoder in the model to obtain encoded features, decode the encoded features through the target decoder in the model, and obtain the speech recognition result based on the decoding result.

[0084] Combined Figure 2 As shown, the speech recognition model may include a shared encoder and a decoder network. The decoder network includes two types of decoders, which are respectively defined as the first type of decoder and the second type of decoder. The number of the first type of decoders is one or more, and the number of the second type of decoders is one or more.

[0085] To achieve the unification of the speech recognition model in streaming and non-streaming recognition tasks, in terms of the network structure, the shared encoder in this application is an encoding network that supports both streaming and non-streaming recognition tasks.

[0086] The shared encoder in this embodiment may adopt a Transformer or a Conformer model, or an encoding network with other structures. Since the Conformer has stronger performance in speech recognition tasks, in an optional example, the Conformer model can be used as the shared encoder. To ensure that the Conformer model can support both streaming and non-streaming recognition tasks at the same time, a causal convolutional structure and an attention mechanism Attention module based on the chunk form to limit the scope of action can be adopted inside the Conformer model in this embodiment, ensuring that the Conformer model can meet the encoding requirements of both streaming and non-streaming recognition tasks.

[0087] The first type of decoder in the speech recognition model supports streaming decoding processing, and the second type of decoder supports non-streaming decoding processing. By configuring two types of decoders, the decoding requirements of streaming and non-streaming recognition tasks can be met.

[0088] When the speech recognition method is applied to a streaming recognition task (that is, the speech to be recognized is a speech data stream collected under a streaming recognition task), the target decoder for decoding processing in step S110 is specifically the first type of decoder.

[0089] When the speech recognition method is applied to a non-streaming recognition task (that is, the speech to be recognized is the entire speech collected under a non-streaming recognition task), the target decoder for decoding processing in step S110 includes at least the second type of decoder.

[0090] The speech recognition method provided by the embodiments of the present application is implemented through an end-to-end unified speech recognition model. This speech recognition model integrates streaming and non-streaming speech recognition capabilities and can be applicable to multiple different scenarios through a single speech recognition model. Specifically, the speech recognition model includes a shared encoder and two types of decoders. The shared encoder uses an encoding network that supports both streaming and non-streaming recognition tasks to meet the encoding requirements for the acoustic features of different streaming and non-streaming speech data streams. The model integrates a first type of decoder that supports streaming decoding processing and a second type of decoder that supports non-streaming decoding processing. For the speech to be recognized, it can be a speech data stream composed of audio chunks in a streaming recognition task or the entire speech in a non-streaming recognition task. After extracting the acoustic features of the speech to be recognized, they are sent into the speech recognition model, encoded by the shared encoder, and the encoded features are sent into the target decoder for decoding. Based on the decoding result, the final speech recognition result is obtained. In a streaming recognition task, the target decoder is the first type of decoder, and in a non-streaming recognition task, the target decoder is the second type of decoder. The solution of the present application ensures that through a unified speech recognition model, it can be adapted to both streaming and non-streaming recognition tasks, thereby reducing the model development cycle and development resources. Moreover, through the unified speech recognition model, it can be adapted to multiple task scenarios without switching different models, improving the convenience of use.

[0091] The number of the first type of decoders in the speech recognition model can be one or more. The more the number of the first type of decoders, the more decoding strategies can be adopted during the decoding operation. For example, different first type of decoders can be selected for decoding processing according to different task requirements, or the decoding results of multiple different first type of decoders can be combined simultaneously to improve the accuracy of the recognition result. When the number of the first type of decoders is multiple, decoders that support streaming recognition tasks with multiple different network structures can be selected. The following are several optional structures of the first type of decoders: CTC decoder, RNN-T decoder, etc. In some possible examples, the first type of decoder supports both streaming and non-streaming recognition tasks while supporting the streaming recognition task. The two first type of decoders in the above examples support both streaming and non-streaming recognition tasks.

[0092] For the CTC decoder:

[0093] The CTC decoder consists of a fully connected layer and a Softmax layer. It converts the output of the shared encoder into CTC activations, models the frame-level alignment information between the acoustic features and the text modeling units, and obtains the probabilities of the modeling units.

[0094] For the RNN-T decoder:

[0095] The RNN-T is a more advanced streaming decoder. It concatenates the acoustic features learned by the encoder with the prediction vectors output by the Pred. prediction network through the Joint network, and then transmits the result to the Softmax layer to obtain the probability of the corresponding modeling unit.

[0096] 1) Pred. Network: The role of this network is to make predictions based on the previous non-empty symbol sequence to obtain a higher-level feature representation. It includes a word embedding layer that maps a morpheme vocabulary vector containing m units to an n-dimensional space vector (n < m), followed by a one-way LSTM with 2×n units and a fully connected layer with n units.

[0097] 2) Joint Network: Combines the high-level feature representations output by the encoder and the Pred. prediction network, and then sends them to the Softmax layer to obtain the probability of the corresponding modeling unit.

[0098] The number of the second type of decoders in the speech recognition model can be one or more. The more the number of the second type of decoders, the more decoding strategies can be adopted during the decoding operation. For example, different second type of decoders can be selected for decoding according to different task requirements, or the decoding results of multiple different second type of decoders can be combined simultaneously to improve the accuracy of the recognition result. In this embodiment, an optional network structure of the second type of decoder is provided, and the AED (Attention Encoder-Decoder) decoder can be adopted.

[0099] The AED decoder consists of an N d layer Transformer Decoder decoder unit block. It inputs the acoustic features learned by the encoder and the encoded annotation text into the decoder network, and then through a fully connected layer and a Softmax layer, obtains the probability of the corresponding modeling unit.

[0100] Next, for the streaming recognition task and the non-streaming recognition task, the specific implementation processes of decoding the encoded features by the target decoder in step S110 and obtaining the speech recognition result based on the decoding result are introduced respectively.

[0101] 1. Streaming Recognition Task

[0102] The number of the first type of decoders in the speech recognition model can be one or more. As shown in Figure 3 This embodiment provides two optional implementation schemes for the streaming recognition task:

[0103] Scheme 1: Any one of the first type of decoders in the speech recognition model can be used to decode the encoded features, and the decoding result is used as the streaming speech recognition result.

[0104] Taking the example that the speech recognition model includes two first - type decoders, namely the CTC decoder and the RNNT decoder. The decoding result of the CTC decoder can be used as the streaming speech recognition result, or the decoding result of the RNNT decoder can be used as the streaming speech recognition result.

[0105] Since the first - type decoders support streaming recognition tasks, any one of the first - type decoders can be selected to decode the encoded features to obtain the streaming speech recognition result. This processing method is simpler and has a lower computational cost.

[0106] Solution 2: Two or more first - type decoders in the speech recognition model can be used to separately decode the encoded features to obtain two or more decoding results. PK is performed on the two or more decoding results, and the streaming speech recognition result is determined based on the PK result.

[0107] Specifically, different decoding results include the probability scores of the modeling units. The decoding result with a higher probability score can be selected from the two or more decoding results as the streaming speech recognition result.

[0108] Taking the example that the speech recognition model includes two first - type decoders, namely the CTC decoder and the RNNT decoder. The decoding result with a higher probability score can be obtained in the form of PK between the CTC decoder and the RNNT decoder as the streaming speech recognition result.

[0109] By performing PK on two or more decoding results and selecting the decoding result with a higher probability score as the streaming speech recognition result, the decoding capabilities of different first - type decoders can be fully utilized, improving the accuracy and stability of the streaming speech recognition result, and it is not easy to have the problem of abnormal repeated decoding.

[0110] 2. Non - streaming recognition task

[0111] The number of second - type decoders in the speech recognition model can be one or more. As shown in Figure 4 , several optional implementation solutions for non - streaming recognition tasks are provided in this embodiment:

[0112] Solution 1: Any one of the second - type decoders in the model can be used to decode the encoded features, and the decoding result is used as the non - streaming speech recognition result.

[0113] Since the second - type decoders support non - streaming recognition tasks, any one of the second - type decoders can be selected to decode the encoded features to obtain the non - streaming speech recognition result. This processing method is simpler and has a lower computational cost.

[0114] Solution 2: More than two second - type decoders in the speech recognition model can be used to decode the encoded features respectively to obtain more than two decoding results. Then, perform PK on the more than two decoding results, and determine the non - streaming speech recognition result based on the PK result.

[0115] Specifically, different decoding results include the probability scores of the modeling units. The decoding result with a higher probability score can be selected from the more than two decoding results as the non - streaming speech recognition result.

[0116] By performing PK on more than two decoding results and selecting the decoding result with a higher probability score as the non - streaming speech recognition result, the decoding capabilities of different second - type decoders can be fully utilized, improving the accuracy and stability of the non - streaming speech recognition result.

[0117] In some possible examples, the first - type decoder supports both streaming and non - streaming decoding processes, while the second - type decoder only supports non - streaming decoding processes. For example, the first - type decoder includes any one or more of the CTC decoder and the RNN - T decoder. The second - type decoder can adopt the AED decoder or other decoders that only support non - streaming decoding processes.

[0118] In this case, for the process of decoding the encoded features through the target decoder in the model and obtaining the speech recognition result based on the decoding result in the non - streaming recognition task, the present embodiment further provides two other alternative implementation solutions:

[0119] Solution 3: Decode the encoded features through a first - type decoder in the model to obtain a preliminary decoding result, and then re - score the preliminary decoding result through a second - type decoder, and determine the non - streaming speech recognition result based on the re - scoring result.

[0120] Specifically, since the first - type decoder supports both streaming and non - streaming decoding processes, any first - type decoder can be called first to decode the encoded features to obtain a preliminary decoding result. On this basis, in order to improve the accuracy and stability of the final speech recognition result, the preliminary decoding result can be further sent to any second - type decoder to re - score the preliminary decoding result through the second - type decoder, and determine the non - streaming speech recognition result based on the re - scoring result.

[0121] By combining the first - type decoder and the second - type decoder and adopting the re - scoring mechanism, the recognition effect of the non - streaming recognition task can be improved, and the recognition stability can also be improved, that is, the problem of abnormal repeated decoding is not likely to occur.

[0122] Solution 4: Decode the encoded features through two or more first - type decoders in the model to obtain two or more preliminary decoding results. Select one preliminary decoding result with a higher probability score, and re - score the one preliminary decoding result with a higher probability score through the second - type decoder. Determine the non - streaming speech recognition result based on the re - scoring result.

[0123] Compared with the above Solution 3, in Solution 4, two or more first - type decoders are used to decode the encoded features to obtain two or more preliminary decoding results, and then PK is performed to select one preliminary decoding result with a higher probability score and send it to the second - type decoder for re - scoring, which can further improve the recognition effect of the non - streaming recognition task.

[0124] Taking the case where the first - type decoders in the speech recognition model include a CTC decoder and an RNNT decoder, and the second - type decoder includes an AED decoder as an example. First, the CTC decoder and the RNNT decoder can be used to decode the encoded features, perform PK on the preliminary decoding results, select one preliminary decoding result with a higher probability score, and send it to the AED decoder for two - pass re - scoring. Determine the non - streaming speech recognition result based on the re - scoring result. This method can reuse the capabilities of various types of decoders. Through the PK mechanism and the re - scoring mechanism, the recognition effect of the non - streaming speech recognition result can be improved, and it will also be more stable, that is, it is not easy to have problems such as abnormal repeated decoding.

[0125] In some embodiments of the present application, the shared encoder Conformer in the speech recognition model is described.

[0126] In order to unify the non - streaming and streaming recognition tasks in the encoder structure, a causal convolutional structure that only considers the previous context and a chunk - based form are adopted inside Conformer to limit the scope of action of Attention.

[0127] Furthermore, in order to improve the stability of the encoded features output by Conformer, a progressive down - sampling strategy is provided in this embodiment. Specifically, combined with Figure 5 the Conformer structure shown, it can be seen that:

[0128] Conformer includes a number of down - sampling modules connected in series, and the number of down - sampling modules performs progressive down - sampling operations.

[0129] Figure 5The downsampling module includes a pre-downsampling module and two downsampling modules, stage1 and stage2, connected in series. Different downsampling modules perform progressive downsampling operations according to the set downsampling multiples. Compared with the traditional Conformer that only performs a one-time downsampling operation through a pre-downsampling module, the Conformer provided in this embodiment performs progressive downsampling operations through multiple downsampling modules, which can improve the stability of the encoded features output by the Conformer.

[0130] In addition, the traditional Conformer performs 1 / 4 times downsampling through the pre-downsampling module. Through experiments and verification in this case, it is found that when the overall downsampling multiple of the Conformer is 1 / 8, faster computing efficiency and better recognition effects can be obtained. On this basis, through a number of downsampling modules connected in series, progressive downsampling can be carried out step by step until after the last downsampling module performs the downsampling operation, the Conformer can overall reach 1 / 8 times downsampling. Taking Figure 5 as an example, the downsampling multiple of the pre-downsampling module can be set to 1 / 2, and the downsampling multiples of the two downsampling modules of stage1 and stage2 are each 1 / 2, so that an overall downsampling of 1 / 8 times can be achieved.

[0131] In some possible implementations, in order to improve the computing efficiency of the Conformer and reduce the required computing resources, a grouped multi-head self-attention mechanism is provided in this embodiment, and the Conformer model can include a grouped multi-head self-attention module.

[0132] When the grouped multi-head self-attention module is calculating, the dimensions of the query Q, key K, and value V in the traditional multi-head self-attention mechanism are transformed from (n, d) to (n / g, d×g), and then attention calculation is performed. After the calculation, the dimensions are transformed back to the original (n, d), where g is the group size.

[0133] Combined with Figure 5 as shown, in the stage1 and stage2 modules, a downsampling block can be stacked after N Conformer Blocks to perform downsampling along the time dimension. Figure 5 An optional network structure of the downsampling block is illustrated on the right:

[0134] It includes a grouped multi-head self-attention module. The output of the grouped multi-head self-attention module is sent to a convolutional downsampling block to perform downsampling operation in the time dimension.

[0135] Figure 6 Illustrated Figure 5 An optional network structure of the convolutional downsampling block in

[0136] It includes a Layer Norm module, a Pointwise Conv module, a Glu activation function module, a Strided Depthwise Conv module, a Bath Norm module, a Swish activation function module, a Pointwise Conv module, and a Dropout output module connected in series. Among them, the stride of the Strided Depthwise Conv module can be set to a value greater than 1 to achieve temporal downsampling. Since the dimension of the output after downsampling is smaller than that of the input, the residual module needs to synchronously add a Pointwise Projection module to perform downsampling and map the input and output to the same dimension.

[0137] Among them, the Pointwise Projection module can be implemented in two ways: convolution sampling and Pooling downsampling. Considering that the Pooling effect is better and the streaming inference is more convenient, the Pooling downsampling method can be preferentially used.

[0138] In the traditional Multi-head Self-Attention (MHSA) module, the sizes of Q, K, and V are (n, d), where (n, d) is the dimension of the K and V matrices. The computational complexity of the traditional multi-head self-attention module is O(n 2 d). In this embodiment, in order to reduce the computational complexity and accelerate the training and inference speed, a grouped multi-head self-attention module is proposed.

[0139] Combined with Figure 7 the schematic diagram of the grouped multi-head self-attention shown, when the grouped multi-head self-attention module calculates, it transforms the dimensions of the query Q, key K, and value V in the traditional multi-head self-attention mechanism from (n, d) to (n / g, d×g), then performs attention calculation, and after the calculation, transforms the dimensions back to the original (n, d), where g is the group size.

[0140] After the transformation, the computational complexity of the grouped multi-head self-attention module can be reduced to O(n 2 d / g).

[0141] In summary, by constructing a Conformer shared encoder with progressive downsampling and grouped multi-head self-attention, this application can improve the computational efficiency of the Conformer, reduce the required computing resources, and ensure the stability of the output encoded features, thus improving the final speech recognition effect.

[0142] In some embodiments of the present application, considering the overall system framework of the ASR task, in addition to the most core speech recognition task, it generally includes other various types of tasks. Examples include the preposed Voice Activity Detection (VAD) task, Language IDentification (LID) task, and the postposed Punctuation Prediction Model (PPM) task, case normalization task, etc.

[0143] The VAD task is to determine the start point and end point of speech from a signal containing speech. The LID task is a technology for automatically determining the language type of a speech signal. The PPM task is the process of automatically adding correct punctuation marks to the recognized result text. The PPM task and the case normalization task can improve the readability and applicability of the recognized result text.

[0144] In the current overall system framework of the ASR task, for different types of tasks, corresponding task models are generally trained separately, such as separately training an effective speech detection model, a language identification model, a speech recognition model, a punctuation prediction model, etc.

[0145] In addition, for different languages, some existing technologies also train specific task models separately for a single language. Examples of different languages include Chinese, English, Japanese, Korean, Russian, German, French, etc.

[0146] The traditional ASR overall system has problems such as single-task and single-language training consuming a large amount of time and human resources, being unable to support streaming, and limited recognition effects. Therefore, in this embodiment, a multi-task and multi-language training method is provided, which saves a large amount of time and human resources and simultaneously meets non-streaming and streaming recognition tasks.

[0147] For the speech recognition model introduced in the foregoing embodiments, it is trained in a multi-language and multi-task manner during the training stage. The multi-tasks include the speech recognition task ASR and at least one of the following types of tasks:

[0148] Punctuation prediction PPM task, language identification LID task, effective speech detection VAD task, etc.

[0149] By adopting the above training strategy, not only can a speech recognition model with higher efficiency, more recognition tasks, and stronger generalization ability be obtained, seamless adaptation between different tasks can be achieved, and a large amount of time and human resources consumed in establishing single-task and single-language models can be reduced, but also the recognition effect will be better, the problem of easy abnormal repeated decoding can be solved, and more stable and robust recognition results can be obtained.

[0150] The embodiments of this application provide a training solution for a speech recognition model. In terms of the training method, a dynamic chunk training method can be adopted, and different batches use dynamic chunk sizes during training. The dynamic chunk size ranges from 1 to a uniform distribution of the maximum utterance length.

[0151] The model training method may specifically include the following steps:

[0152] S1. Obtain the training data of the current batch, and the non-streaming sample speech and streaming sample speech in the training data are allocated according to a set ratio.

[0153] In a possible example, to ensure the balance of non-streaming and streaming chunk data, in the training data of each batch, the non-streaming chunks (the maximum utterance length complete chunks without streaming transmission) and the streaming chunks each account for half of the ratio, that is, the above set ratio is 1:1.

[0154] S2. Extract the acoustic features of the sample speech in the training data and send them into the speech recognition model for processing to obtain the decoding results output by the first type of decoder and the second type of decoder.

[0155] Specifically, there can be multiple decoding strategies. For details, refer to the relevant introduction in the previous model inference process and will not be elaborated here.

[0156] S3. Calculate the loss of each decoder based on the decoding result output by each decoder and the recognition result label corresponding to the sample speech in the training data.

[0157] S4. Perform weighted fusion on the losses of each decoder to obtain the total loss, and update the model parameters through the backpropagation algorithm according to the total loss until the set training end condition is reached.

[0158] In the training process of the speech recognition model introduced in this embodiment, by fusing streaming sample speech and non-streaming sample speech in the training data of each batch, it can be ensured that the trained model can be applied to both streaming and non-streaming recognition tasks.

[0159] Combined with Figure 8 As shown, it exemplifies a schematic diagram of the training process of a speech recognition model under a multi-language and multi-task training strategy.

[0160] Figure 8 It is illustrated by taking the first type of decoder including a CTC decoder and an RNN-T decoder, and the second type of decoder including an AED decoder as an example.

[0161] The multi-tasks include a punctuation prediction PPM task, a language identification LID task, a valid speech detection VAD task, and a speech recognition ASR task.

[0162] First, it is necessary to uniformly encode the annotated text corresponding to the sample speech in the multi-task training format, referring to Figure 9 An example of a multi-task and multi-language training format encoding block diagram is shown.

[0163] The annotation text encoding is to convert the text input into a numerical input that can be accepted by the model. This converter can be called a tokenizer, which mainly includes functions such as word segmentation, special processing, dictionary construction, and digitization, such as GPT2Tokenizer, SentencePiece Tokenizer, etc. Among them, the dictionary not only contains multi-language word segmentation text tokens for languages such as Chinese, English, Japanese, Korean, Russian, German, French, Spanish, etc., but also contains corresponding symbol tokens (punc tokens) for each language. In addition, it also contains special tokens such as <|startoftranscript|>, <|endoftranscript|>, <|en|>, <|zh|>,... (language tags)..., <|ja|>, <|transcribe|>, <|startofprev|>, <|nospeech|>, <|notimestamps|>. Among them, <|en|>, <|zh|>,... (language tags)..., <|ja|> are Figure 9 the language tags in

[0164] (1) The encoding format without timestamps (no timestamps) is:

[0165] <|startoftranscript|> + language tag + <|transcribe|> + <|notimestamps|> + text / punc tokens + <|endoftranscript|>.

[0166] For example Figure 10aAn encoding format without timestamps is exemplified. Among them, "SOT" at the beginning represents |startoftranscript|, "EOT" at the end represents |endoftranscript|, and "ZH" represents the Chinese language tag.

[0167] (2) The encoding format with timestamps is as follows:

[0168] <|startoftranscript|> + language tag + <|transcribe|> + begin time + text / punc tokens + end time +... + begin time + text / punc tokens + end time + <|endoftranscript|>.

[0169] Such as Figure 10b An encoding format with timestamps is exemplified. Among them, the start time and end time are included before and after each segmented word or punctuation mark.

[0170] The above embodiments of this application provide a unified encoding strategy for multi-task and multi-language training formats, and the annotated text corresponding to the sample speech can be encoded according to this encoding strategy.

[0171] During training, the combined loss of the CTC decoder, RNN-T decoder, and AED decoder is used as the training loss function.

[0172] Calculate respectively , , The three losses, and use different weights to combine them as the loss of the entire network, and then continuously update the network parameters through the backpropagation algorithm until the training reaches model convergence, and finally the speech recognition model can be obtained.

[0173]

[0174] Among them, are , , The weight ratios of different losses, the parameters are conveniently adjustable, and satisfy relationship.

[0175] Through a large amount of training corpus, a speech recognition model can be trained to obtain a robust speech recognition model. On this basis, some training corpus in target vertical fields can be further used for fine-tuning and customized training of the model to obtain a vertical speech recognition model with better speech recognition effect in the target vertical field.

[0176] In summary, for the problems existing in the current overall ASR system, such as single-task and single-language training consuming a large amount of time and human resources, being unable to support streaming, and limited recognition effect, etc., this application proposes a multi-task and multi-language non-streaming & streaming end-to-end unified speech recognition solution based on an improved Conformer as the encoder network and a combination of CTC, RNN-T, and AED as the decoder network. From the unified encoding of the multi-task and multi-language training format of speech and annotation text, an improved Conformer shared encoder is constructed with progressive downsampling and grouped multi-head self-attention to learn acoustic features, a hybrid decoder network is formed by combining CTC, RNN-T, and AED, to the end-to-end unified speech recognition model training and model finetune for directional optimization, a speech recognition model with higher efficiency, more recognition tasks, and stronger generalization ability is obtained. In the model inference stage, through a non-streaming inference method with the entire sentence of speech as the input, the CTC and RNN-T decoders are used in a PK form to obtain a result with a higher probability score as the non-streaming one-pass decoding result, and then the AED decoder is used to re-score the one-pass decoding result as the final result to construct a non-streaming speech recognition system; through a streaming inference method with a per-chunk speech data stream, the CTC and RNN-T are used in a PK form to obtain a result with a higher probability score as the final streaming recognition result to construct a streaming speech recognition system. This can not only achieve seamless adaptation between tasks, reduce the large amount of time and human resources consumed by establishing single-task and single-language models, but also improve the recognition effect, solve the problem of abnormal repeated decoding that is likely to occur, and obtain more stable and robust recognition results.

[0177] In addition, CTC, RNN-T, and AED can also be used alone as decoders to decode the recognition results for the needs of different efficiency application scenarios.

[0178] Next, the speech recognition device provided by the embodiments of the present application will be described. The speech recognition device described below can be correspondingly referred to the speech recognition method described above.

[0179] See Figure 11 , Figure 11 which is a schematic structural diagram of a speech recognition device disclosed in the embodiments of the present application.

[0180] As Figure 11 shown, the device may include:

[0181] An acoustic feature extraction unit 11 is configured to extract acoustic features of a speech to be recognized, where the speech to be recognized is a speech data stream composed of audio chunks in a streaming recognition task, or a whole speech in a non-streaming recognition task;

[0182] A speech recognition processing unit 12 is configured to send the acoustic features into a speech recognition model, encode the acoustic features through a shared encoder in the model to obtain encoded features, decode the encoded features through a target decoder in the model, and obtain a speech recognition result based on the decoding result;

[0183] Wherein, the shared encoder is an encoding network that supports both streaming and non-streaming recognition tasks, and the speech recognition model includes a first type of decoder that supports streaming decoding processing and a second type of decoder that supports non-streaming decoding processing. In the streaming recognition task, the target decoder is the first type of decoder, and in the non-streaming recognition task, the target decoder includes at least the second type of decoder.

[0184] In a possible implementation, the process in which the speech recognition processing unit decodes the encoded features through a target decoder in the model and obtains a speech recognition result based on the decoding result includes:

[0185] In the streaming recognition task:

[0186] Decode the encoded features through any one of the first type of decoders in the model, and use the decoding result as the streaming speech recognition result;

[0187] Or, decode the encoded features through two or more of the first type of decoders in the model respectively to obtain two or more decoding results, and select the decoding result with a higher probability score as the streaming speech recognition result based on the two or more decoding results.

[0188] In a possible implementation, the process in which the speech recognition processing unit decodes the encoded features through a target decoder in the model and obtains a speech recognition result based on the decoding result includes:

[0189] In the non-streaming recognition task:

[0190] Decode the encoded features through any one of the second type of decoders in the model, and use the decoding result as the non-streaming speech recognition result;

[0191] Or, decode the encoded features through two or more of the second type of decoders in the model respectively to obtain two or more decoding results, and select the decoding result with a higher probability score as the non-streaming speech recognition result based on the two or more decoding results.

[0192] In one possible implementation, the first type of decoder supports both streaming and non-streaming decoding processes, and the second type of decoder only supports non-streaming decoding processes. On this basis, the process of the speech recognition processing unit decoding the encoded features through the target decoder in the model and obtaining the speech recognition result based on the decoding result includes:

[0193] In the non-streaming recognition task:

[0194] Decoding the encoded features through one of the first type of decoders in the model to obtain a preliminary decoding result, re-scoring the preliminary decoding result through the second type of decoder, and determining the non-streaming speech recognition result based on the re-scoring result;

[0195] Or,

[0196] Decoding the encoded features through more than two of the first type of decoders in the model respectively to obtain more than two preliminary decoding results, selecting a preliminary decoding result with a higher probability score, and re-scoring the preliminary decoding result with a higher probability score through the second type of decoder, and determining the non-streaming speech recognition result based on the re-scoring result.

[0197] In one possible implementation, the speech recognition model called by the speech recognition processing unit is trained in a multi-language and multi-task manner during the training phase;

[0198] The multi-tasks include a speech recognition task and at least one of the following types of tasks:

[0199] Punctuation prediction PPM task, language identification LID task, valid speech detection VAD task.

[0200] In one possible implementation, the first type of decoder in the speech recognition model called by the speech recognition processing unit includes a CTC decoder and / or an RNN-T decoder; the second type of decoder includes an AED decoder.

[0201] In one possible implementation, the shared encoder in the speech recognition model called by the speech recognition processing unit adopts a Convolutional Transformer model Conformer, and the Conformer model internally adopts a causal convolutional structure and an Attention module based on a chunk form to limit the scope of action.

[0202] In one possible implementation, the Conformer model in the speech recognition model called by the speech recognition processing unit includes a plurality of downsampling modules connected in series, and the plurality of downsampling modules perform progressive downsampling operations.

[0203] In one possible implementation, the Conformer model in the speech recognition model called by the speech recognition processing unit includes a grouped multi-head self-attention module;

[0204] When the grouped multi-head self-attention module calculates, it transforms the dimensions of the query Q, key K, and value V in the traditional multi-head self-attention mechanism from (n, d) to (n / g, d×g), and then performs attention calculation. After the calculation, the dimensions are transformed back to the original (n, d), where g is the group size.

[0205] In one possible implementation, the device of the present application may further include a model training unit for training the speech recognition model. The training process includes:

[0206] Obtain the training data of the current batch, and allocate the non-streaming sample speech and streaming sample speech in the training data according to a set ratio;

[0207] Extract the acoustic features of the sample speech in the training data and send them to the speech recognition model for processing to obtain the decoding results output by the first type of decoder and the second type of decoder;

[0208] Based on the decoding result output by each decoder and the recognition result label corresponding to the sample speech in the training data, calculate the loss of each decoder;

[0209] Perform weighted fusion on the losses of each decoder to obtain the total loss, and update the model parameters through the backpropagation algorithm according to the total loss until the set training end condition is reached.

[0210] An electronic device is also provided in an embodiment of the present application. Refer to Figure 12 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, tablet computers, learning machines, wearable devices, and the like. Figure 12 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiment of the present application.

[0211] As Figure 12As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603, so as to implement the voice recognition method of the foregoing embodiments of the present application. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0212] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 12 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0213] An embodiment of the present application also provides a computer program product including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any one of the voice recognition methods provided by the embodiments of the present application.

[0214] An embodiment of the present application also provides a computer-readable storage medium, which carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, can enable the electronic device to implement any one of the voice recognition methods provided by the embodiments of the present application.

[0215] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationships between the modules indicate that they have communication connections, which may be specifically implemented as one or more communication buses or signal lines.

[0216] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can easily be implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for the present application, in more cases, software program implementation is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0217] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0218] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0219] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

Claims

1. A speech recognition method, characterized in that: include: Extracting acoustic features of a speech to be recognized, wherein the speech to be recognized is a speech data stream composed of audio chunks in a streaming recognition task, or is a whole segment of speech in a non-streaming recognition task; The acoustic features are sent to a speech recognition model to encode the acoustic features through a shared encoder in the model to obtain encoded features, the encoded features are decoded through a target decoder in the model, and a speech recognition result is obtained based on the decoding result; The shared encoder is a coding network that supports streaming and non-streaming recognition tasks, the speech recognition model includes a first type of decoder that supports streaming decoding processing and a second type of decoder that supports non-streaming decoding processing, the target decoder in the streaming recognition task is the first type of decoder, and the target decoder in the non-streaming recognition task includes at least the second type of decoder.

2. The method according to claim 1, characterized in that The process of decoding the encoded features by a target decoder in the model and obtaining a speech recognition result based on the decoding result includes: In the streaming recognition task: Decoding the encoded features by any one of the first-class decoders in the model, and using the decoding result as the streaming speech recognition result; Or, the coding features are decoded respectively by more than two of the first-type decoders in the model to obtain more than two decoding results, and a decoding result with a higher probability score is selected as the streaming speech recognition result based on the more than two decoding results.

3. The method according to claim 1, characterized in that The process of decoding the encoded features by a target decoder in the model and obtaining a speech recognition result based on the decoding result includes: In the non-streaming recognition task: Decoding the encoded features by any one of the second-class decoders in the model, and using the decoding result as the non-streaming speech recognition result; Or, the encoding features are decoded respectively by more than two of the second-type decoders in the model to obtain more than two decoding results, and a decoding result with a higher probability score is selected as the non-streaming speech recognition result based on the more than two decoding results.

4. The method according to claim 1, characterized in that The first type of decoder supports both streaming and non-streaming decoding processing, and the second type of decoder only supports non-streaming decoding processing; The process of decoding the encoded features by a target decoder in the model and obtaining a speech recognition result based on the decoding result includes: In the non-streaming recognition task: Decoding the encoded features by a first-class decoder in the model to obtain a preliminary decoding result, rescoring the preliminary decoding result by a second-class decoder, and determining a non-streaming speech recognition result based on the rescoring result; or, The coding features are decoded respectively by more than two of the first-class decoders in the model to obtain more than two preliminary decoding results, a preliminary decoding result with a higher probability score is selected, and the preliminary decoding result with a higher probability score is re-scored by the second-class decoder, and the non-streaming speech recognition result is determined based on the re-scoring result.

5. The method according to claim 1, characterized in that The speech recognition model is trained in a multi-language and multi-task manner during the training phase; The multiple tasks include a speech recognition task and at least one of the following types of tasks: Punctuation prediction PPM task, language identification LID task, and effective speech detection VAD task.

6. The method according to any one of claims 1 to 5, characterized in that: The first type of decoders includes CTC decoders and / or RNN-T decoders; the second type of decoders includes AED decoders.

7. The method according to any one of claims 1 to 5, characterized in that: The shared encoder adopts a convolution transformer model Conformer, and the Conformer model internally adopts a causal convolution structure and an Attention module based on a chunk-based scope-limiting attention mechanism.

8. A speech recognition device, characterized in that: include: An acoustic feature extraction unit, used to extract acoustic features of a speech to be recognized, wherein the speech to be recognized is a speech data stream composed of audio chunks in a streaming recognition task, or is a whole segment of speech in a non-streaming recognition task; A speech recognition processing unit, configured to input the acoustic features into a speech recognition model, encode the acoustic features through a shared encoder in the model to obtain encoded features, decode the encoded features through a target decoder in the model, and obtain a speech recognition result based on the decoding result; The shared encoder is a coding network that supports streaming and non-streaming recognition tasks, the speech recognition model includes a first type of decoder that supports streaming decoding processing and a second type of decoder that supports non-streaming decoding processing, the target decoder in the streaming recognition task is the first type of decoder, and the target decoder in the non-streaming recognition task includes at least the second type of decoder.

9. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the speech recognition method as described in any one of claims 1 to 7.

10. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the speech recognition method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, each step of the speech recognition method as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Language identification method, device and equipment

    CN121306096A