Voice processing method and electronic device

By combining the same text sequences in the one-pass decoding result in end-to-end speech recognition, the problem of the N-best decoding algorithm being responsive is solved, and the speech recognition performance and decoding efficiency are improved without increasing the N value.

CN115410577BActive Publication Date: 2025-06-24LENOVO (BEIJING) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211066538.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-01
Publication Date
2025-06-24
Estimated Expiration
2042-09-01

AI Technical Summary

Technical Problem

In the existing end-to-end speech recognition technology, the N-best decoding algorithm has a small receptive field, which leads to the fact that the first-pass decoding result is not improved in the second-pass decoding, affecting the final speech recognition performance.

Method used

By merging the same text sequences under different recognition paths in the first sub-block recognition result corresponding to the current sub-block recognition result to recognize the speech, the matching recognition probability of the same text sequence in multiple recognition paths is improved, and the second sub-block recognition result of the current sub-block is determined based on the merge processing result.

Benefits of technology

Without increasing the N value, the effective recognition results in the N-best information are added to improve the decoding efficiency and ensure the speech recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115410577B_ABST
    Figure CN115410577B_ABST
Patent Text Reader

Abstract

The present application discloses a voice processing method and an electronic device. The method includes: obtaining a current target voice block to be recognized of the voice to be recognized, recognizing the text information corresponding to the target voice block to obtain a voice block recognition result of the target voice block, based on the voice block recognition result, determining a first sub-block recognition result corresponding to the current sub-block formed by the starting voice block to the target voice block of the voice to be recognized, and performing a merging process on the same text sequences corresponding to different recognition paths included in the first sub-block recognition result, and determining a second sub-block recognition result of the current sub-block based on the merging process result, and determining a text recognition result of the voice to be recognized based on the second sub-block recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of speech recognition, and particularly relates to a speech processing method and an electronic device. Background Art

[0002] End-to-end speech recognition refers to the process of directly obtaining a text sequence based on a sequence of speech feature vectors using a deep neural network.

[0003] Currently, mainstream and relatively effective end-to-end speech recognition generally adopts a 2-pass (i.e., two-pass decoding) module design. It obtains an online streaming first-pass decoding result through one-pass decoding processing, and improves speech recognition performance by re-scoring / re-ranking the N-best information (i.e., the N text results corresponding to the probabilities belonging to the top N) included in the first-pass decoding result through two-pass decoding processing.

[0004] The N-best decoding algorithm has a relatively small receptive field, resulting in fewer actual effective recognition results corresponding to the N-best information provided to the two-pass decoding process. This causes the improvement of the first-pass decoding result in the two-pass decoding to be not high, thus affecting the final speech recognition performance. Traditional techniques increase the value of N to ensure that the N-best information fed into the two-pass decoding contains more different effective recognition results. However, the larger N is, the slower the decoding speed and the more complex the decoding, which will lead to a significant decrease in decoding efficiency. Summary of the Invention

[0005] Therefore, the present application discloses the following technical solutions:

[0006] A speech processing method, the method comprising:

[0007] Obtain the current target speech block to be recognized of the speech to be recognized;

[0008] Recognize the text information corresponding to the target speech block to obtain a speech block recognition result;

[0009] Based on the speech block recognition result, determine a first sub-block recognition result corresponding to the current sub-block formed by the speech to be recognized from the starting speech block to the target speech block;

[0010] Perform a merging process on the same text sequences corresponding to different recognition paths in the first sub-block recognition result; the merging process can improve the recognition probability of the same text sequence matching any one of the multiple recognition paths corresponding thereto;

[0011] Determine a second sub-block recognition result of the current sub-block based on the merging process result, and determine a text recognition result of the speech to be recognized based on the second sub-block recognition result.

[0012] Optionally, determining the first sub-block recognition result corresponding to the current sub-block formed by the speech to be recognized from the starting speech block to the target speech block based on the speech block recognition result includes:

[0013] Performing splicing processing on the multiple text information included in the speech block recognition result and the multiple different previous text sequences included in the previous recognition result corresponding to the previous speech block of the target speech block in the speech to be recognized, respectively, to obtain the multiple text sequences corresponding to the current sub-block, so as to determine the first sub-block recognition result.

[0014] Optionally, determining the first sub-block recognition result corresponding to the current sub-block formed by the speech to be recognized from the starting speech block to the target speech block based on the speech block recognition result further includes:

[0015] Fusing the recognition probabilities corresponding to the currently spliced text information and the previous text sequences respectively, and taking the fused probability as the recognition probability of the spliced text sequence; so as to determine the first sub-block recognition result of the current sub-block based on the multiple text sequences corresponding to the current sub-block and the recognition probabilities corresponding to the multiple text sequences respectively;

[0016] Among them, the recognition probability corresponding to the previous text sequence of the target speech block is: on the recognition path corresponding to each previous speech block of the target speech block, when the recognition of one speech block is completed, the recognition probability of the text information corresponding to the currently completed recognized speech block is fused with the recognition probability of the previous text sequence corresponding to the speech block currently corresponding, until the fused result obtained by fusing the recognition probability of the previous adjacent speech block of the target speech block is completed.

[0017] Optionally, performing splicing processing on the multiple text information included in the speech block recognition result and the multiple different previous text sequences included in the previous recognition result corresponding to the previous speech block of the target speech block in the speech to be recognized, respectively, includes:

[0018] Determining multiple text information whose corresponding recognition probabilities in the speech block recognition result belong to the top N in the recognition probability descending sequence;

[0019] Separately splicing each of the multiple text information of the top N to the tail of the multiple different previous text sequences of the target speech block.

[0020] Optionally, the speech block recognition result includes the corresponding relationship between each text information in the text recognition space and the corresponding recognition probability, and the recognition probability in the corresponding relationship includes: the conditional probability that the target speech block respectively matches each text information in the text recognition space under the corresponding preconditions; the text recognition space includes multiple different text information provided by the speech recognition model for speech recognition;

[0021] The preconditions corresponding to the target speech block include: using the speech block recognition results of each previous speech block corresponding to the target speech block in the speech to be recognized as known conditions;

[0022] The fusion of the recognition probabilities corresponding to the currently concatenated text information and the previous text sequence respectively includes: fusing the conditional probability of the currently concatenated text information and the recognition probability of the previous text sequence.

[0023] Optionally, the merging process for the same text sequences corresponding to different recognition paths in the first sub-block recognition result includes:

[0024] Determine whether there are the same text sequences corresponding to different recognition paths in the first sub-block recognition result;

[0025] If so, fuse the recognition probabilities of the same text sequences respectively matched to the different recognition paths, and use the fused probability as the recognition probability of the same text sequence.

[0026] Optionally, determining the second sub-block recognition result of the current sub-block based on the merging process result, and determining the text recognition result of the speech to be recognized based on the second sub-block recognition result includes:

[0027] Based on the merging process result, determine that the recognition probabilities corresponding to the current sub-block belong to the top N text sequences in the descending order of recognition probabilities as the second sub-block recognition result of the current sub-block;

[0028] Let the top N text sequences and their corresponding recognition probabilities participate in the processing of the next speech block of the target speech block. When the processing of the last speech block of the speech to be recognized is completed, use the second sub-block recognition result of the sub-block corresponding to the last speech block as the first-stage recognition result of the speech to be recognized;

[0029] Determine the second-stage recognition result of the speech to be recognized according to the first-stage recognition result and the speech features of the speech to be recognized;

[0030] Determine the text recognition result of the speech to be recognized according to the first-stage recognition result and the second-stage recognition result.

[0031] Optionally, the recognition of the text information corresponding to the target speech block to obtain the speech block recognition result includes:

[0032] Determine the speech features of the target speech block;

[0033] According to the speech features, recognize the text information corresponding to the target speech block to obtain the speech block recognition result.

[0034] Optionally, the voice feature includes the acoustic feature and the language feature of the target voice block;

[0035] The process of determining the acoustic feature and the language feature of the target voice block includes:

[0036] Encoding the target voice block by using an encoding unit of a speech recognition model, and using the obtained voice feature vector after the encoding process as the acoustic feature of the target voice block;

[0037] Performing a prediction process on the basis of the language information corresponding to the target voice block by using a prediction unit in a first decoding unit of the speech recognition model to obtain the language feature of the target voice block;

[0038] Wherein, the language information corresponding to the target voice block is the language information obtained by extracting information from the language context environment where the target voice block is located.

[0039] An electronic device includes:

[0040] A memory for storing at least one set of computer instruction sets;

[0041] A processor for implementing the speech processing method as described in any one of the above by calling and executing the instruction sets stored in the memory.

[0042] As can be seen from the above solutions, for the speech processing method and the electronic device disclosed in this application, the to-be-recognized target voice block of the to-be-recognized speech is obtained, the text information corresponding to the target voice block is recognized to obtain the speech block recognition result of the target voice block, based on this speech block recognition result, the first sub-block recognition result corresponding to the current sub-block formed by the starting voice block to the target voice block of the to-be-recognized speech is determined, and the same text sequences corresponding to different recognition paths included in the first sub-block recognition result are merged, and based on the result of the merging process, the second sub-block recognition result of the current sub-block is determined, and the text recognition result of the to-be-recognized speech is determined based on the second sub-block recognition result.

[0043] In this application, by merging the same text sequences under different recognition paths in the first sub-block recognition result corresponding to the current sub-block of the to-be-recognized speech, and based on the result of the merging process, the second sub-block recognition result of the current sub-block and the text recognition result of the to-be-recognized speech based on this are determined, which can avoid the phenomenon that the receptive field of the N-best decoding algorithm becomes smaller due to the existence of the same text recognition results in different paths. Correspondingly, without increasing the value of N, the N-best information can include more different valid recognition results, thereby improving the decoding efficiency and ensuring the speech recognition performance. Description of the Drawings

[0044] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0045] Figure 1 It is a schematic flowchart of the voice processing method provided by the present application;

[0046] Figure 2 It is a composition structure diagram of the voice decoding model provided by the present application;

[0047] Figure 3 It is a partial model structure diagram of the voice decoding model after introducing the prediction unit provided by the present application;

[0048] Figure 4 It is an example where different recognition paths correspond to the same text sequence provided by the present application;

[0049] Figure 5 It is the decoding process of performing one-pass decoding on the voice based on the voice recognition model provided by the present application;

[0050] Figure 6 It is a composition structure diagram of the electronic device provided by the present application. Detailed implementation manners

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0052] The embodiments of the present application disclose a voice processing method and an electronic device, which are applicable to streaming or non-streaming voice recognition scenarios, and are used to improve the problem of the relatively small receptive field of the decoding algorithm in streaming or non-streaming voice recognition scenarios without increasing the N value of the N-Best decoding algorithm, so as to improve the decoding efficiency and ensure the voice recognition performance. The voice processing method of the present application can be, but is not limited to, applied to electronic devices in many general or special computing device environments or configurations, such as: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor devices, and so on.

[0053] Refer to Figure 1 the flowchart of the voice processing method provided, and the voice processing method provided by the embodiments of the present application includes the following processing procedures:

[0054] Step 101: Obtain the target speech chunk to be recognized in the current speech to be recognized.

[0055] The speech to be recognized can be a complete speech sentence / speech segment in a streaming or non-streaming speech recognition scenario, or can also be a segment obtained by splitting a complete speech sentence / speech segment, without limitation. The speech chunk of the speech to be recognized can be a speech frame or multiple consecutive speech frames of the speech to be recognized, such as a speech sentence / speech segment.

[0056] In a speech recognition scenario, according to the actual recognition progress, the target speech chunk to be recognized in the current speech to be recognized can be continuously input into the speech recognition model, and the speech recognition model performs recognition processing on the continuously input target speech chunks.

[0057] Optionally, the embodiment of the present application uses a speech recognition model designed based on a 2-pass (i.e., two-pass decoding) module for speech recognition. That is to say, the decoding stage in the speech recognition process includes two-pass decoding processing. Correspondingly, the speech recognition model includes an encoding unit and two decoding units: a first decoding unit and a second decoding unit.

[0058] The composition structure of the speech recognition model is as Figure 2 shown, where Shared encoder represents the encoding unit, which can be shared by two decoding units, so it can also be called a shared encoding unit or a shared encoder; first-pass decoder represents the first decoding unit, which can also be called a one-pass decoding unit or a one-pass decoder, and is used to perform one-pass decoding processing of speech to obtain a one-pass decoding result, such as obtaining an online streaming one-pass decoding result. Second-pass decoder represents the second decoding unit, which can also be called a two-pass decoding unit or a two-pass decoder, and is responsible for performing two-pass decoding processing of speech. Specifically, it improves the speech recognition performance by re-scoring the N-best (i.e., the N text results corresponding to the probabilities belonging to the top N, where N is an integer greater than 1) in the one-pass decoding result output by the first-pass decoder. Score merge represents a score fusion unit, which is used to fuse the recognition probabilities of the N-best texts in the one-pass decoding result and the recognition probabilities of the N-best texts in the two-pass decoding result to obtain the final recognition result of the speech to be recognized.

[0059] Based on the above composition structure of the speech recognition model, this step 101 can specifically obtain the target speech chunk to be recognized of the input speech to be recognized in the encoding unit of the speech recognition model, such as the shared encoder Shared encoder.

[0060] Step 102: Recognize the text information corresponding to the target speech chunk to obtain a speech chunk recognition result.

[0061] After obtaining the target speech chunk, the speech feature of the target speech chunk can be determined. According to the included speech feature, the text information corresponding to the target speech chunk is recognized, and the speech chunk recognition result of the target speech chunk is obtained accordingly. It should be noted that the text information here can include Chinese text, text in other languages, and character information, etc.

[0062] The speech feature of the target speech chunk includes at least the acoustic feature of the target speech chunk.

[0063] Optionally, specifically, the encoding unit of the speech recognition model can be used to encode the target speech chunk, and the obtained speech feature vector after encoding is used as the acoustic feature of the target speech chunk. For Figure 2 the model structure shown, the shared encoder of the model can be used to encode the input target speech chunk to obtain the acoustic feature of the target speech chunk.

[0064] In other embodiments, in addition to the acoustic feature, the speech feature of the target speech chunk can also include a language feature, so as to improve the accuracy of speech recognition by combining the acoustic feature and the language feature of the speech.

[0065] The language feature of the target speech chunk can be obtained by processing the language information corresponding to the target speech chunk. The language information corresponding to the target speech chunk can be obtained by extracting information from the language context where the target speech chunk is located.

[0066] Exemplarily, specifically, each previous speech chunk corresponding to the target speech chunk in the speech to be recognized can be used as the context information of the target speech chunk, and the language information corresponding to the target speech chunk can be obtained based on the context information represented by each previous speech chunk. In this case, optionally, the language information corresponding to the target speech chunk is empty or can be the speech chunk recognition result corresponding to its (the target speech chunk) previous adjacent speech chunk that has recognized text information. Among them, when the target speech chunk is the first speech chunk of the speech to be recognized, or when no text information has been recognized for each previous speech chunk corresponding to the target speech chunk in the speech to be recognized, the language information corresponding to the target speech chunk is empty. When text information has been recognized for the previous speech chunk corresponding to the target speech chunk in the speech to be recognized, the language information corresponding to the target speech chunk is the speech chunk recognition result corresponding to its previous adjacent speech chunk that has recognized text information, and specifically can be the N-best information in the recognition result corresponding to the previous adjacent speech chunk that has recognized text information.

[0067] Optionally, a prediction unit is introduced into the first decoding unit of the speech recognition model. By using the introduced prediction unit to perform prediction processing on the language information corresponding to the target speech block, the features of the target speech block at the language level are predicted, and the language features of the target speech block are correspondingly obtained.

[0068] See Figure 3 , which further shows a partial composition structure of the speech recognition model after the introduction of the prediction unit. Among them, Encoder represents the encoding unit of the model, that is, the shared encoder, which is used to encode the input speech block to obtain the acoustic features of the speech block. The remaining parts, namely predictor, net, and softmax, are the components of the first decoding unit, jointly constituting a first-pass decoder. Predictor represents the introduced prediction unit, and net and softmax respectively represent the decoding network and the normalization unit of the first-pass decoder. The prediction unit predictor is responsible for processing the language information of the currently to-be-recognized speech block to obtain the language features of the currently to-be-recognized speech block, facilitating subsequent speech decoding by combining acoustic features and language features.

[0069] After determining the speech features of the target speech block, the determined speech features can be sent to the decoding network of the first decoding unit for decoding processing. Through the decoding processing, the text information corresponding to the target speech block is recognized according to the speech features of the target speech block.

[0070] Among them, if the determined speech features include the acoustic features of the target speech block, the acoustic features of the target speech block, that is, the speech feature vector obtained by the encoding unit encoding the target speech block, are transmitted to the decoding network of the first decoding unit for decoding processing.

[0071] If the determined speech features include the acoustic features and language features of the target speech block, then the acoustic and language features of the target speech block are sent to the decoding network of the first decoding unit so that the decoding network can jointly perform speech decoding by combining acoustic features and language features.

[0072] As Figure 3 shown, assume that the target speech block is speech block x t , that is, the t-th speech block of the speech to be recognized. After the encoding unit Encoder (that is, the shared encoder) encodes x t , the obtained feature vector is used as the acoustic feature of x t and input into the decoding network net of the first-pass decoder of the first decoding unit. After the prediction unit predictor performs prediction processing on the corresponding language information y t of x u-1 , the prediction result As the speech chunk x t The linguistic features are input into the decoding network net of the first-pass decoder of the first decoding unit. After introducing the predictor, the decoding network net becomes the joint network, namely joint net, which is responsible for jointly decoding the acoustic features t of the speech chunk x and the linguistic features to implement the decoding of the speech chunk x t so as to identify the recognition probabilities respectively corresponding to different text information in the text recognition space of the model. In t it is denoted as z Figure 3 . The softmax is a normalization unit, which is used to map the probabilities respectively corresponding to different text information in the text recognition space of the speech chunk to the range of [0, 1] to obtain the output of the softmax, that is t,u P(y Figure 3 |x u , y 1:t ) in 1:u-1 .

[0073] The above-mentioned text recognition space includes multiple different text information provided by the speech recognition model for speech recognition, such as multiple different keywords, key phrases, etc.

[0074] In this step, the speech chunk recognition result obtained by recognizing the target speech chunk according to the speech features of the target speech chunk is the result obtained after performing the above decoding process on the speech features of the target speech chunk by the first decoding unit. This result can specifically be Figure 3 the speech chunk decoding result output by the joint net in

[0075] or the result output after further softmax normalization processing, without limitation.

[0076] The speech chunk recognition result corresponding to the target speech chunk specifically includes the corresponding relationship between each text information in the text recognition space and the corresponding recognition probability. Optionally, the recognition probability in the corresponding relationship includes: the conditional probability that the target speech chunk matches each text information in the text recognition space respectively under the corresponding preconditions; further, the preconditions corresponding to the target speech chunk include: taking the speech chunk recognition results of each previous speech chunk corresponding to the target speech chunk in the speech to be recognized as known conditions.

[0077] Among them, the speech chunk recognition result corresponding to the target speech chunk is allowed to be empty, which is specifically reflected in that each probability in the above corresponding relationship is empty, indicating that for the target speech chunk, no valid text information can be recognized (for example, in the case where the target speech chunk is the speech frame corresponding to the gap between different speeches in a sentence). Figure 3P(y) output for the speech block u |x 1:t ,y 1:u-1 ) as an example, P(y u |x 1:t ,y 1:u-1 ) can be empty, and if it is not empty, it means the speech block x t The conditional probability under the condition that its previous speech blocks are formed, specifically, under the known condition that the recognition probabilities of the first t-1 speech blocks to be recognized and the first u-1 speech blocks that recognize text information are formed, x t The conditional probabilities corresponding to each text information (such as each keyword / word) contained in the text recognition space of the speech recognition model. It is easy to understand that the speech blocks other than the u-1 speech blocks in the first t-1 speech blocks are speech blocks whose text information cannot be recognized (such as speech frames corresponding to the gaps between different speech in a sentence), and their corresponding speech block recognition results are correspondingly empty.

[0078] Step 103: Based on the speech block recognition result, determine a first sub-block recognition result corresponding to a current sub-block of the speech to be recognized, which is formed from the starting speech block to the target speech block.

[0079] The first sub-block recognition result includes: text sequences corresponding to different recognition paths of the current sub-block and recognition probabilities corresponding to each text sequence.

[0080] The text sequence can specifically be a text string. For example, for the speech stream corresponding to "wo ai zu guo" to be recognized, and the speech block to be recognized corresponding to "zu", one of the multiple text sequences corresponding to the current sub-block can be the text string "我爱祖".

[0081] After obtaining the speech block recognition result of the target speech block, the speech block recognition result of the target speech block and the preceding recognition result corresponding to the preceding speech block in the speech to be recognized can be processed to obtain the first sub-block recognition result corresponding to the current sub-block composed of the speech to be recognized from the starting speech block to the target speech block. This process includes processing in two aspects: text splicing and probability fusion, which can be specifically implemented as follows:

[0082] 11) The multiple text information included in the speech block recognition result of the target speech block and the multiple different preceding text sequences included in the preceding recognition result corresponding to the preceding speech block of the target speech block in the speech to be recognized are respectively concatenated to obtain the multiple text sequences corresponding to the current sub-block to determine the first sub-block recognition result.

[0083] Specifically, the speech block recognition results of the target speech block can be pruned to determine the corresponding recognition probabilities belonging to the top N text information in the recognition probability descending sequence, and each of the top N text information can be separately spliced ​​to the end of each preceding text sequence included in the preceding recognition result corresponding to the preceding speech block of the target speech block in the speech to be recognized, and each text sequence obtained by splicing is the text sequence corresponding to the current sub-block.

[0084] For the N-best decoding algorithm, the preceding text sequence recognized based on the preceding speech block of the target speech block also retains its corresponding N-best results through pruning, that is, the preceding text sequence with the recognition probability belonging to TOPN among all preceding text sequences of the target speech block is retained.

[0085] Accordingly, the N-Best text information in the speech block recognition result of the target speech block can be individually spliced ​​to the end of each sequence in the N-best preceding text sequence of the target speech block, and a total of N*N splicing results can be obtained.

[0086] Taking the speech stream corresponding to "wo ai zu guo" as the speech to be recognized as an example, assuming that the current target speech block to be recognized is the speech frame corresponding to "zu", and assuming that the N-Best in the corresponding speech block recognition result is 6-Best (i.e., N=6) text information, then each text information in the 6-Best information needs to be separately spliced ​​into the 6-Best preceding text sequence corresponding to the recognized "wo ai" speech stream, and a total of 36 splicing results are obtained, such as one of the splicing results can be "I love ancestor".

[0087] 12) The recognition probabilities corresponding to the currently concatenated text information and the previous text sequence are fused, and the fused probability is used as the recognition probability of the concatenated text sequence.

[0088] At the same time, the conditional probability corresponding to the text information of the target speech block and the recognition probability corresponding to the preceding text sequence concatenated with the text information are fused, and the obtained fused probability is used as the recognition probability of the current sub-block. The fusion process may be, but is not limited to, multiplying the probabilities corresponding to the two.

[0089] Each text sequence corresponding to the current sub-block and the recognition probability corresponding to each text sequence constitute the first sub-block recognition result corresponding to the current sub-block.

[0090] Among them, the recognition probability corresponding to the preceding text sequence of the target speech block is: every time the recognition of a speech block is completed on the recognition path corresponding to each preceding speech block of the target speech block, the recognition probability of the text information corresponding to the speech block currently completed is fused with the recognition probability of the preceding text sequence currently corresponding to the speech block, until the fusion result obtained by fusion of the recognition probability of the previous adjacent speech block of the target speech block is completed.

[0091] For example, when the text information "祖" in the N-best text recognition result of the target speech block "zu" is concatenated with the preceding text sequence "我爱" corresponding to its preceding speech block into "我爱祖", the conditional probability corresponding to "祖" and the recognition probability corresponding to "我爱" are fused at the same time, and the fused probability is used as the recognition probability of the current sub-block "我爱祖".

[0092] The concatenation process of the text information corresponding to the target speech block and the corresponding preceding text sequence can be performed in the decoding network of the first decoding unit of the speech recognition model (e.g., Figure 3 The decoding network decodes and recognizes the continuously input speech blocks on the one hand, and on the other hand, continuously concatenates the recognized text information with the recognized previous text sequence.

[0093] Step 104: merge the same text sequences corresponding to different recognition paths contained in the sub-block recognition results; the merge process can improve the recognition probability that the same text sequence matches any recognition path in the corresponding multiple recognition paths.

[0094] In traditional technology, the first sub-block recognition result obtained after concatenating the text information of the last speech block is used as the first-pass recognition result of the speech to be recognized, and N-Best information (that is, each text sequence with the recognition probability belonging to the top N in the first-pass recognition result) is screened and input into the second encoding and decoding unit for second-pass decoding processing.

[0095] The applicant has found that the N-best information in the decoding result often has the same decoding result, but the corresponding decoding paths are different, which will reduce the receptive field of the N-best decoding algorithm. Figure 4 In the example provided, the voice stream corresponding to "team" is the voice to be recognized, and each voice frame in the voice stream of "team" is a voice block to be recognized. Figure 4As shown, although the recognition result is the same, namely "team", there are three different decoding paths (arrows of the same gray scale belong to the same path). If this recognition result, i.e., "team", belongs to the N-best in the first-pass decoding result, it will lead to the actual N-best containing only N - 2 candidate text results, not meeting the required number N of the N-best decoding algorithm, with a smaller receptive field, resulting in a limited improvement of the first-pass decoding result in the second-pass decoder and affecting the final speech recognition performance.

[0096] In view of the above situation, the embodiments of the present application propose a technical idea of merging the same text sequences corresponding to each time step after the decoding of each time step to improve the problem of the smaller receptive field of the N-best decoding algorithm. Among them, the decoding process of each time step refers to the decoding process of each speech block of the speech to be recognized in the first-pass decoding stage, and the decoding of each speech block is regarded as corresponding to one time step.

[0097] Based on the above technical idea, after obtaining the speech block recognition result of the target speech block, it means that the decoding of the current time step is completed, and accordingly, the text sequences in the first sub-block recognition result of the current sub-block at the current time step can be merged. The current sub-block at the current time step is the sub-block formed by the starting speech block to the target speech block of the speech to be recognized.

[0098] The above merging process can be implemented as follows:

[0099] Determine whether there are the same text sequences corresponding to different recognition paths in the first sub-block recognition result of the current sub-block. If so, fuse the recognition probabilities of the same text sequence matching different recognition paths, and use the fused probability as the recognition probability of the same text sequence. Otherwise, if not, do not fuse.

[0100] Fusing the recognition probabilities of the same text sequence matching different recognition paths can be, but is not limited to, summing the recognition probabilities of the same text sequence matching different recognition paths, and using the obtained probability sum value as the recognition probability of the same text sequence. It should be noted that in practical applications, other fusion algorithms can also be used, without limitation. For example, weighted summation operations, etc., as long as the fused probability is higher than the recognition probability of the same text sequence matching any path among the multiple corresponding recognition paths, it belongs to the protection scope of the embodiments of the present application.

[0101] For example, with reference to Figure 3, assuming that the TOPN recognition results corresponding to the "team" speech stream contain the text sequence "team", and there are 3 different recognition paths corresponding to the text sequence "team", then the recognition probabilities corresponding to "team" in the 3 different recognition paths can be fused, such as summing, etc., and the fused probability is used as the final recognition probability of "team".

[0102] Step 105, determine the second sub-block recognition result of the current sub-block based on the merging processing result, and determine the text recognition result of the speech to be recognized based on the second sub-block recognition result.

[0103] After that, based on the merging processing result, the second sub-block recognition result of the current sub-block can be continuously determined. Specifically, multiple text sequences with recognition probabilities belonging to the top N in the recognition probability descending sequence in the merging processing result corresponding to the current sub-block can be determined and used as the second sub-block recognition result of the current sub-block.

[0104] Subsequently, the second sub-block recognition result of the current sub-block will participate in the processing of the next speech block of the target speech block, such as splicing and probability fusion with the N-best text information of the next speech block. Until the processing of the last speech block of the speech to be recognized is completed, the second sub-block recognition result of the sub-block corresponding to the last speech block is used as the first-stage recognition result of the speech to be recognized, and this first-stage recognition result is also the one-pass decoding result of the speech to be recognized.

[0105] For the speech to be recognized, after decoding is completed at each time step of the present application, that is, after each speech block recognition result is obtained, the merging processing of the same text sequences in different paths in the recognition result of the sub-block at the current time step is introduced, which can make the N-Best text sequences obtained based on the merging processing different from each other, thus ensuring the receptive field of the N-Best decoding at each time step, and correspondingly ensuring the receptive field of the finally output one-pass decoding result.

[0106] For example, for the speech stream to be recognized "wo ai zu guo", it is assumed that the 6-Best text information corresponding to the speech block "zu" (such as "祖", "足", "组", "租"...) is spliced ​​with the 6-Best text sequence corresponding to the recognized "wo ai" speech stream (such as "我爱", "我挨"...), and 36 splicing results are obtained. And it is assumed that among the 6-Best text sequences of the 36 splicing results, there is a text sequence corresponding to 3 recognition paths. Then the traditional technology will cause the 6-Best output of the splicing result to actually contain only 4 text sequences, which does not reach the value of N required by the N-Best decoding algorithm, that is, less than 6, and the receptive field is small. The present application merges the same text sequences under different paths, and performs 6-Best screening based on the merged results, so that the 6-Best text sequences finally screened are different from each other and can make up for the required value of N.

[0107] After obtaining the first-stage recognition results of the speech to be recognized, the N-best recognition results in the first stage, that is, the N recognition results whose corresponding recognition probabilities belong to TOPN, are further input into the second decoding unit of the speech recognition model, and the speech features of the speech to be recognized are input into the second decoding unit. The second decoding unit performs a second-pass decoding process on the speech to be recognized based on the input information. In the second-pass decoding process, the second decoding unit, such as the second-pass decoder, re-scores / re-sorts the N-best results in the first-pass decoding results according to the speech features of the speech to be recognized.

[0108] Optionally, the speech features of the speech to be recognized input into the second decoding unit may be acoustic features of the speech to be recognized, for example, specifically may be acoustic features obtained after the encoding unit (such as a shared encoder) of the speech recognition model encodes the speech blocks of the speech to be recognized.

[0109] Finally, the recognition probabilities of the N-Best text results in the first-stage recognition results and the recognition probabilities of the N-Best text results after rescoring / sorting in the second-stage recognition results can be combined to determine the final recognition probabilities corresponding to the N-Best text results of the speech to be recognized, so as to output the results. For example, based on the final recognition probability, the recognition result with the highest probability is selected from the N-Best text results (such as selecting the highest probability recognition result "I love my motherland" corresponding to "wo ai zu guo") and output it as the optimal result of the speech to be recognized.

[0110] See also Figure 2, specifically, the N-best information in the output result of the first-pass decoder and the N-best information output by the second-pass decoder can be sent to the score fusion unit score merge. The score merge fuses the N-best information of the two decoders to obtain the final N-best recognition result of the speech to be recognized.

[0111] Since the same text sequences under different paths are eliminated in the first-pass decoding stage, the N text recognition results (i.e., N text sequences) included in the N-best information sent to the second-pass decoder are different from each other. Therefore, there is no need to increase the value of N to ensure the receptive field, and the decoding speed of the whole process is faster and the efficiency is higher.

[0112] In summary, the method of the present application merges the same text sequences under different recognition paths in the first sub-block recognition result corresponding to the current sub-block of the speech to be recognized, and determines the second sub-block recognition result of the current sub-block and the text recognition result of the speech to be recognized based on the merged result, which can avoid the phenomenon that the receptive field of the N-best decoding algorithm becomes smaller due to the same text recognition results in different paths. Correspondingly, without increasing the value of N, the N-best information can contain more different valid recognition results, thereby improving the decoding efficiency and ensuring the speech recognition performance.

[0113] To facilitate a clear understanding of the first-pass decoding process in the method of the present application, an example is provided below for illustration.

[0114] In this example, the normal training method can be adopted when training the speech recognition model, that is, the process of merging the same text under different paths can be not introduced in the first-pass decoding process during the model training stage. However, this is not limited thereto, and this merging process can also be introduced during the training stage.

[0115] See Figure 5 , in this example, based on the speech recognition model obtained by training, the process of performing first-pass decoding on the speech to achieve speech recognition includes:

[0116] 21) Initialize the state of the predictor; initialize the current token set.

[0117] Optionally, both the state of the predictor and the token information in the token set are initialized to blank (empty), that is, initially, the input of the predictor is blank, and the token information in the token set is empty. Tokens are used to record and transfer the recognition state and the predictor state on the speech recognition path during the speech recognition process of the speech to be recognized.

[0118] Among them, the recognition status recorded in the token set includes, after completing one pass of decoding at the current time step, under the recognition progress of this time step, the N-best recognition results of the recognized text sequences corresponding to each recognition path. Here, one token corresponds to one recognition path; the predictor status recorded in the token set includes the N-best recognition results corresponding to the speech block adjacent to the previous recognized text information on the current speech block.

[0119] 22) Detect whether the encoder has an output. If "no", output the N-best results of the current token set; if "yes", the predictor obtains the output of the predictor according to the predictor status in the token set.

[0120] The encoder having no output means that no current speech block is input to the encoder, which correspondingly means that one pass of decoding of the last speech block of the speech to be recognized has been completed. Thus, the N-best results recorded in the current token set can be directly output as the N-best of the one-pass decoding result of the speech to be recognized, and will be sent to the two-pass decoder for two-pass decoding processing later (or after normalization).

[0121] Conversely, if the encoder has an output, it means that there are still speech blocks of the speech to be recognized that need to be processed. Correspondingly, the output of the predictor can be obtained according to the predictor status in the token set for use as the language feature of the current speech block to be recognized.

[0122] 23) The Joint net obtains the probability vector of the current speech block according to the output of the predictor and the output of the encoder, and performs the first pruning.

[0123] The pruning here refers to screening out the text information with recognition probabilities belonging to TOPN from the recognition results corresponding to each text information of the current speech block in the model text recognition space.

[0124] 24) Set the token set at the next moment as an empty list.

[0125] 25) Traverse the current token set: perform two operations on each token traversed:

[0126] a. Multiply the probability of this token (specifically referring to the recognition probability of the currently recognized text sequence recorded in this token) by the blank probability in the probability vector, and add it as a new token to the token set at the next moment;

[0127] b. Traverse the probability vector after the first pruning: Concatenate the text information corresponding to each probability (conditional probability) in the probability vector after the first pruning to the text sequence corresponding to this token, and multiply the probabilities of the two as the new token to be added to the token set at the next moment.

[0128] 26) Merge the tokens with the same text sequence in the token set at the next moment, and add the probabilities.

[0129] 27) Prune the token set at the next moment and set it as the current token set. Continue to detect whether the encoder has an output.

[0130] The pruning here refers to the N-Best text sequence screening based on the merged result on the basis of merging the probabilities of the same text sequence.

[0131] The embodiment of the present application also discloses an electronic device, and its composition structure is as Figure 6 shown, including at least:

[0132] A memory 10 for storing a computer instruction set;

[0133] The computer instruction set can be implemented in the form of a computer program.

[0134] A processor 20 for implementing the voice processing method disclosed in any of the above method embodiments by executing the computer instruction set.

[0135] The processor 20 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, etc.

[0136] The electronic device is equipped with a display device and / or has a display interface and can be externally connected to a display device.

[0137] Optionally, the electronic device further includes a camera component and / or is connected to an external camera component.

[0138] In addition, the electronic device may further include components such as a communication interface and a communication bus. The memory, the processor, and the communication interface complete communication with each other through the communication bus.

[0139] The communication interface is used for communication between the electronic device and other devices. The communication bus can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0140] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.

[0141] For the sake of convenience in description, when describing the above system or device, it is divided into various modules or units according to functions for separate description. Of course, when implementing the present application, the functions of each unit can be realized in the same or multiple software and / or hardware.

[0142] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.

[0143] Finally, it should also be noted that in this article, relational terms such as first, second, third, and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.

[0144] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A speech processing method, the method comprising: Obtaining a to-be-recognized target speech block of the to-be-recognized speech currently; Recognizing text information corresponding to the target speech block to obtain a speech block recognition result; Based on the speech block recognition result, determining a first sub-block recognition result corresponding to a current sub-block formed by the to-be-recognized speech from a starting speech block to the target speech block; Performing a merging process on the same text sequences corresponding to different recognition paths in the first sub-block recognition result; the merging process can increase the recognition probability of the same text sequence matching any one of the multiple recognition paths corresponding thereto; Determining a second sub-block recognition result of the current sub-block based on the merging process result, and determining a text recognition result of the to-be-recognized speech based on the second sub-block recognition result.

2. The method according to claim 1, wherein the determining a first sub-block recognition result corresponding to a current sub-block formed by the to-be-recognized speech from a starting speech block to the target speech block based on the speech block recognition result comprises: Performing a splicing process on multiple text information included in the speech block recognition result and multiple different previous text sequences included in a previous recognition result corresponding to a previous speech block of the target speech block in the to-be-recognized speech respectively, to obtain multiple text sequences corresponding to the current sub-block, so as to determine the first sub-block recognition result.

3. The method according to claim 2, wherein the determining a first sub-block recognition result corresponding to a current sub-block formed by the to-be-recognized speech from a starting speech block to the target speech block based on the speech block recognition result further comprises: Fusing recognition probabilities corresponding to the currently spliced text information and the previous text sequences respectively, and using the fused probability as the recognition probability of the spliced text sequence; Determining the first sub-block recognition result of the current sub-block based on the multiple text sequences corresponding to the current sub-block and the recognition probabilities corresponding to the multiple text sequences respectively; wherein, the recognition probability corresponding to the previous text sequence of the target speech block is: when recognizing each speech block on the recognition paths corresponding to the previous speech blocks of the target speech block, fusing the recognition probability of the text information corresponding to the currently recognized speech block with the recognition probability of the previous text sequence corresponding to the speech block currently, until the fused result of fusing the recognition probabilities of the previous adjacent speech block of the target speech block is completed.

4. The method according to claim 2 or 3, wherein the performing a splicing process on multiple text information included in the speech block recognition result and multiple different previous text sequences included in a previous recognition result corresponding to a previous speech block of the target speech block in the to-be-recognized speech respectively comprises: Determining multiple text information whose corresponding recognition probabilities in the speech block recognition result belong to the top N in a recognition probability descending sequence; Separately splicing each of the multiple text information of the top N to the tails of the multiple different previous text sequences of the target speech block.

5. According to the method described in claim 3, the speech block recognition result includes the correspondence between each text information in the text recognition space and the corresponding recognition probability, and the recognition probability in the correspondence includes: The conditional probabilities of the target speech block respectively matching each text information in the text recognition space under corresponding preconditions; The text recognition space includes multiple different text information for speech recognition provided by a speech recognition model; The preconditions corresponding to the target speech block include: taking the speech block recognition results of each previous speech block corresponding to the target speech block in the speech to be recognized as known conditions; The fusion of the recognition probabilities corresponding to the currently spliced text information and the previous text sequence respectively includes: fusing the conditional probability of the currently spliced text information and the recognition probability of the previous text sequence.

6. The method according to claim 1, wherein the merging process of the same text sequences corresponding to different recognition paths in the recognition result of the first sub-block includes: Determining whether there are the same text sequences corresponding to different recognition paths in the recognition result of the first sub-block; If so, fusing the recognition probabilities of the same text sequences respectively matched to the different recognition paths, and taking the fused probability as the recognition probability of the same text sequence.

7. The method according to claim 1, wherein determining the recognition result of the second sub-block of the current sub-block based on the merging process result, and determining the text recognition result of the speech to be recognized based on the recognition result of the second sub-block includes: Based on the merging process result, determining that the recognition probabilities corresponding to the current sub-block belong to the top N text sequences in the recognition probability descending sequence as the recognition result of the second sub-block of the current sub-block; Participating the top N text sequences and the corresponding recognition probabilities in the processing of the next speech block of the target speech block, until when the processing of the last speech block of the speech to be recognized is completed, taking the recognition result of the second sub-block of the sub-block corresponding to the last speech block as the first-stage recognition result of the speech to be recognized; Determining the second-stage recognition result of the speech to be recognized according to the first-stage recognition result and the speech feature of the speech to be recognized; Determining the text recognition result of the speech to be recognized according to the first-stage recognition result and the second-stage recognition result.

8. The method according to claim 1, wherein recognizing the text information corresponding to the target speech block to obtain a speech block recognition result includes: Determining the speech feature of the target speech block; According to the speech feature, recognizing the text information corresponding to the target speech block to obtain a speech block recognition result.

9. The method according to claim 8, wherein, The speech feature includes the acoustic feature and the language feature of the target speech block; The process of determining the acoustic feature and the language feature of the target speech block includes: Encoding the target speech block by using an encoding unit of a speech recognition model, and taking the obtained speech feature vector as the acoustic feature of the target speech block; Performing prediction processing on the language information corresponding to the target speech block by using a prediction unit in a first decoding unit of the speech recognition model to obtain the language feature of the target speech block; Wherein, the language information corresponding to the target speech block is the language information obtained by extracting information from the language context environment where the target speech block is located.

10. An electronic device, comprising: A memory for storing at least a set of computer instruction sets; A processor, configured to implement the voice processing method according to any one of claims 1-9 by calling and executing the instruction set stored in the memory.

Citation Information

Patent Citations

  • Speech recognition method and device, electronic equipment and storage medium

    CN113066480A

  • Model training method and device, voice recognition method and device, medium and equipment

    CN113436620A