Speech Recognition Method, Apparatus, Electronic Device, and Computer-Readable Storage Medium

By adopting a parallel dual speech recognition method in intelligent products, combining posterior probability and category probability, the problems of increasing new words and high training costs are solved, and more efficient and accurate speech recognition is achieved.

CN114495945BActive Publication Date: 2025-05-30ALIBABA GROUP HOLDING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011265991.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-12
Publication Date
2025-05-30
Estimated Expiration
2040-11-12

AI Technical Summary

Technical Problem

In the prior art, when recognizing speech in smart products, there are missing new words in the increase in the number of words, and the changes in the vocabulary cause the workload and time cost of data training, which affects the accuracy and efficiency of recognition.

Method used

The parallel dual speech recognition method based on posterior probability and category-based speech recognition is adopted to achieve the minimum cost of new words adding by category-based speech recognition, and the target speech recognition results are obtained through mixed sorting.

Benefits of technology

Under the condition of ensuring the comprehensiveness of new words, the training workload and time cost of speech recognition are reduced, thereby improving the accuracy and efficiency of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495945B_ABST
    Figure CN114495945B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a voice recognition method, apparatus, electronic device, and computer-readable storage medium. The method includes: performing first voice recognition based on posterior probability on voice data to be recognized to obtain one or more first voice recognition results and their corresponding first voice recognition evaluation scores; performing second voice recognition based on categories on the voice data to be recognized to obtain one or more second voice recognition results and their corresponding second voice recognition evaluation scores; and performing a hybrid ranking on the first voice recognition results and the second voice recognition results based on the first voice recognition evaluation scores and the second voice recognition evaluation scores to obtain a target voice recognition result. This technical solution can greatly reduce the training workload and the required time cost of voice recognition to a certain extent while ensuring the comprehensiveness of new word addition, thereby facilitating the improvement of the accuracy and efficiency of voice recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of speech recognition technology, and more particularly to a speech recognition method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] With the advancement of science and technology, numerous smart products have emerged. To enhance the user experience and increase their intelligence, many smart products support voice interaction, which requires advanced speech recognition capabilities. However, in real-world applications, the speech content that needs to be recognized spans a wide range of fields, including music, movies, games, place names, food, everyday expressions, actions, objects, and more, with tens of millions of related entity words. If entity words are added or modified, for example, a singer releases a new song, a performer launches a new movie or TV series, a developer develops a new game, a smart express locker adds a new function, etc., in the existing technology, the training corpus vocabulary based on which the smart product voice recognition is based needs to be pronounced, updated or expanded, and the voice recognition model based on which the smart product voice recognition is based needs to be retrained. However, due to the limited capacity of the language recognition model vocabulary, the addition of new words requires certain screening, which results in certain omissions in the addition of new words. At the same time, due to changes in the vocabulary, all corpora need to be re-segmented and trained, which will greatly increase the data training workload and the time cost required, which is not conducive to improving the accuracy and efficiency of voice recognition. Summary of the Invention

[0003] Embodiments of the present disclosure provide a speech recognition method, apparatus, electronic device, and computer-readable storage medium.

[0004] In a first aspect, an embodiment of the present disclosure provides a speech recognition method.

[0005] Specifically, the speech recognition method includes:

[0006] Performing first speech recognition based on posterior probability on the speech data to be recognized, obtaining one or more first speech recognition results and their corresponding first speech recognition evaluation scores;

[0007] performing category-based second speech recognition on the to-be-recognized speech data to obtain one or more second speech recognition results and their corresponding second speech recognition evaluation scores;

[0008] The first speech recognition result and the second speech recognition result are mixed and sorted based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a target speech recognition result.

[0009] In conjunction with the first aspect, in a first implementation of the first aspect of the embodiment of the present disclosure, performing first speech recognition based on posterior probability on the speech data to be recognized to obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores includes:

[0010] Extracting acoustic features of the speech data to be recognized;

[0011] Inputting the acoustic features into a pre-trained posterior probability prediction model to obtain a posterior probability matrix, wherein the matrix elements in the posterior probability matrix are the posterior probabilities corresponding to the acoustic features;

[0012] A beam search is performed based on the posterior probability matrix to obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores, wherein the first speech recognition evaluation scores are calculated based on the posterior probability.

[0013] In combination with the first aspect and the first implementation of the first aspect, in the second implementation of the first aspect of the embodiment of the present disclosure, performing category-based second speech recognition on the to-be-recognized speech data to obtain one or more second speech recognition results and their corresponding second speech recognition evaluation scores includes:

[0014] The speech data to be recognized is input into a pre-trained category speech recognition model, and one or more second speech recognition results and their corresponding second speech recognition evaluation scores are obtained based on beam search, wherein the second speech recognition evaluation scores are calculated based on the category probability and the posterior probability.

[0015] In combination with the first aspect, the first implementation manner of the first aspect, and the second implementation manner of the first aspect, in a third implementation manner of the first aspect of the present disclosure, performing category-based second speech recognition on the to-be-recognized speech data to obtain one or more second speech recognition results and their corresponding second speech recognition evaluation scores includes:

[0016] The speech data to be recognized is input into a pre-trained main category speech recognition network. During speech recognition based on the main category speech recognition network, when it is detected that the speech unit to be recognized is a preset category content, the pre-trained auxiliary category speech recognition network is called to perform speech recognition. When recognition is completed or fails based on the auxiliary category speech recognition network, the speech recognition is continued by jumping back to the main category speech recognition network until one or more second speech recognition results and their corresponding second speech recognition evaluation scores are obtained, wherein the second speech recognition evaluation scores are calculated based on the category probability and the posterior probability.

[0017] In combination with the first aspect, the first implementation method of the first aspect, the second implementation method of the first aspect and the third implementation method of the first aspect, in the fourth implementation method of the first aspect of the present disclosure, when training the category speech recognition model, the category labels corresponding to the speech training data and the relevant words are used as input, and the speech recognition results corresponding to the speech training data and the corresponding mixed posterior probabilities are used as output to train the category speech recognition model, wherein the category labels corresponding to the words are obtained by querying a pre-generated category dictionary.

[0018] In combination with the first aspect, the first implementation of the first aspect, the second implementation of the first aspect, the third implementation of the first aspect and the fourth implementation of the first aspect, in the fifth implementation of the first aspect of the present disclosure, the category label is also provided with a word frequency weight to weight the category probability.

[0019] In combination with the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, and the fifth implementation manner of the first aspect, in a sixth implementation manner of the first aspect of the present disclosure, the mixed sorting of the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a target speech recognition result includes:

[0020] performing mixed sorting of the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score;

[0021] The speech recognition result with the highest speech recognition evaluation score is taken as the target speech recognition result.

[0022] In combination with the first aspect, the first implementation of the first aspect, the second implementation of the first aspect, the third implementation of the first aspect, the fourth implementation of the first aspect, the fifth implementation of the first aspect and the sixth implementation of the first aspect, in the seventh implementation of the first aspect of the present disclosure, the first speech recognition and the second speech recognition are executed in parallel and cross-wise.

[0023] In combination with the first aspect, the first implementation manner of the first aspect, the second implementation manner of the first aspect, the third implementation manner of the first aspect, the fourth implementation manner of the first aspect, the fifth implementation manner of the first aspect, the sixth implementation manner of the first aspect, and the seventh implementation manner of the first aspect, in an eighth implementation manner of the first aspect of the present disclosure, the mixed sorting of the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score includes:

[0024] When the first speech recognition and the second speech recognition are performed in parallel, when the semantic completeness of the second speech recognition result obtained by the second speech recognition meets a preset completeness condition, the current first speech recognition result and the current second speech recognition result are mixed and sorted based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a preset number of first intermediate speech recognition results with the highest speech recognition evaluation scores;

[0025] The first speech recognition and the second speech recognition continue to be performed in parallel, and when the semantic completeness of the second speech recognition result obtained by the second speech recognition meets a preset completeness condition, the current first speech recognition result and the current second speech recognition result are mixed and sorted based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a preset number of second intermediate speech recognition results with the highest speech recognition evaluation scores, and the first intermediate speech recognition result is updated using the second intermediate speech recognition results;

[0026] Repeat the steps of generating and updating the intermediate speech recognition result until the first speech recognition and the second speech recognition are completed.

[0027] In a second aspect, an embodiment of the present disclosure provides a speech recognition device.

[0028] Specifically, the speech recognition device includes:

[0029] A first recognition module is configured to perform first speech recognition based on posterior probability on the speech data to be recognized, and obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores;

[0030] a second recognition module configured to perform category-based second speech recognition on the speech data to be recognized, and obtain one or more second speech recognition results and their corresponding second speech recognition evaluation scores;

[0031] The mixing module is configured to perform mixed sorting on the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a target speech recognition result.

[0032] In conjunction with the second aspect, in a first implementation of the second aspect of the embodiment of the present disclosure, the first identification module is configured as follows:

[0033] Extracting acoustic features of the speech data to be recognized;

[0034] Inputting the acoustic features into a pre-trained posterior probability prediction model to obtain a posterior probability matrix, wherein the matrix elements in the posterior probability matrix are the posterior probabilities corresponding to the acoustic features;

[0035] A beam search is performed based on the posterior probability matrix to obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores, wherein the first speech recognition evaluation scores are calculated based on the posterior probability.

[0036] In combination with the second aspect and the first implementation of the second aspect, in a second implementation of the second aspect of the embodiment of the present disclosure, the second identification module is configured as follows:

[0037] The speech data to be recognized is input into a pre-trained category speech recognition model, and one or more second speech recognition results and their corresponding second speech recognition evaluation scores are obtained based on beam search, wherein the second speech recognition evaluation scores are calculated based on the category probability and the posterior probability.

[0038] In combination with the second aspect, the first implementation of the second aspect, and the second implementation of the second aspect, in a third implementation of the second aspect of the present disclosure, the second identification module is configured as follows:

[0039] The speech data to be recognized is input into a pre-trained main category speech recognition network. During speech recognition based on the main category speech recognition network, when it is detected that the speech unit to be recognized is a preset category content, the pre-trained auxiliary category speech recognition network is called to perform speech recognition. When recognition is completed or fails based on the auxiliary category speech recognition network, the speech recognition is continued by jumping back to the main category speech recognition network until one or more second speech recognition results and their corresponding second speech recognition evaluation scores are obtained, wherein the second speech recognition evaluation scores are calculated based on the category probability and the posterior probability.

[0040] In combination with the second aspect, the first implementation method of the second aspect, the second implementation method of the second aspect and the third implementation method of the second aspect, in the fourth implementation method of the second aspect of the present disclosure, when training the category speech recognition model, the category labels corresponding to the speech training data and the relevant words are used as input, and the speech recognition results corresponding to the speech training data and the corresponding mixed posterior probabilities are used as output to train the category speech recognition model, wherein the category labels corresponding to the words are obtained by querying a pre-generated category dictionary.

[0041] In combination with the second aspect, the first implementation of the second aspect, the second implementation of the second aspect, the third implementation of the second aspect and the fourth implementation of the second aspect, in the fifth implementation of the second aspect of the present disclosure, the category label is also provided with a word frequency weight to weight the category probability.

[0042] In combination with the second aspect, the first implementation manner of the second aspect, the second implementation manner of the second aspect, the third implementation manner of the second aspect, the fourth implementation manner of the second aspect, and the fifth implementation manner of the second aspect, in a sixth implementation manner of the second aspect of the present disclosure, the mixing module is configured as follows:

[0043] performing mixed sorting of the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score;

[0044] The speech recognition result with the highest speech recognition evaluation score is taken as the target speech recognition result.

[0045] In combination with the second aspect, the first implementation of the second aspect, the second implementation of the second aspect, the third implementation of the second aspect, the fourth implementation of the second aspect, the fifth implementation of the second aspect and the sixth implementation of the second aspect, in the seventh implementation of the second aspect of the present disclosure, the first speech recognition and the second speech recognition are executed in parallel and cross-wise.

[0046] In combination with the second aspect, the first implementation manner of the second aspect, the second implementation manner of the second aspect, the third implementation manner of the second aspect, the fourth implementation manner of the second aspect, the fifth implementation manner of the second aspect, the sixth implementation manner of the second aspect, and the seventh implementation manner of the second aspect, in an eighth implementation manner of the second aspect of the present disclosure, the part that performs mixed sorting of the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score is configured as follows:

[0047] When the first speech recognition and the second speech recognition are performed in parallel, when the semantic completeness of the second speech recognition result obtained by the second speech recognition meets a preset completeness condition, the current first speech recognition result and the current second speech recognition result are mixed and sorted based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a preset number of first intermediate speech recognition results with the highest speech recognition evaluation scores;

[0048] The first speech recognition and the second speech recognition continue to be performed in parallel, and when the semantic completeness of the second speech recognition result obtained by the second speech recognition meets a preset completeness condition, the current first speech recognition result and the current second speech recognition result are mixed and sorted based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a preset number of second intermediate speech recognition results with the highest speech recognition evaluation scores, and the first intermediate speech recognition result is updated using the second intermediate speech recognition results;

[0049] The generation and updating of the intermediate speech recognition result are repeatedly performed until the first speech recognition and the second speech recognition are completed.

[0050] In a third aspect, embodiments of the present disclosure provide an electronic device comprising a memory and a processor, wherein the memory is configured to store one or more computer instructions that enable a speech recognition device to perform the above-described speech recognition method, and the processor is configured to execute the computer instructions stored in the memory. The speech recognition device may also include a communication interface for communicating with other devices or a communication network.

[0051] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium for storing computer instructions used by a speech recognition device, which includes computer instructions involved in the speech recognition device for executing the above-mentioned speech recognition method.

[0052] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:

[0053] This technical solution performs parallel dual speech recognition based on posterior probability and category-based speech recognition on the speech data to be recognized. This technology utilizes category-based speech recognition to minimize the cost of adding new words and enhance the recognition of words in certain domains. This technical solution significantly reduces the training workload and time required for speech recognition, while ensuring the comprehensiveness of new word addition, thereby improving the accuracy and efficiency of speech recognition.

[0054] It should be understood that the foregoing general description and the following detailed description are merely exemplary and explanatory and are not restrictive of the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Other features, objectives and advantages of the embodiments of the present disclosure will become more apparent through the following detailed description of non-limiting embodiments in conjunction with the accompanying drawings. In the accompanying drawings:

[0056] Figure 1 A flowchart of a speech recognition method according to an embodiment of the present disclosure is shown;

[0057] Figure 2FIG. 1 is an overall flow chart of a speech recognition method according to an embodiment of the present disclosure;

[0058] Figure 3 A structural block diagram of a speech recognition device according to an embodiment of the present disclosure is shown;

[0059] Figure 4 It is a structural diagram of a computer system suitable for implementing the speech recognition method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0060] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for the sake of clarity, parts not related to the description of the exemplary embodiments are omitted in the accompanying drawings.

[0061] In the embodiments of the present disclosure, it should be understood that terms such as "including" or "having" are intended to indicate the existence of features, numbers, steps, behaviors, components, parts, or a combination thereof disclosed in this specification, and are not intended to exclude the possibility of one or more other features, numbers, steps, behaviors, components, parts, or a combination thereof existing or being added.

[0062] It should also be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present disclosure can be combined with each other. The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0063] The technical solution provided by the disclosed embodiments performs parallel dual speech recognition based on posterior probability and category on the speech data to be recognized. This technology utilizes category-based speech recognition to minimize the cost of adding new words and enhance the recognition of words in certain domains. This technical solution can significantly reduce the training workload and time cost of speech recognition while ensuring the comprehensiveness of new word addition, thereby improving the accuracy and efficiency of speech recognition.

[0064] Figure 1 A flow chart of a speech recognition method according to an embodiment of the present disclosure is shown as follows: Figure 1 As shown, the speech recognition method includes the following steps S101-S103:

[0065] In step S101, first speech recognition based on posterior probability is performed on the speech data to be recognized, and one or more first speech recognition results and their corresponding first speech recognition evaluation scores are obtained;

[0066] In step S102, category-based second speech recognition is performed on the speech data to be recognized to obtain one or more second speech recognition results and their corresponding second speech recognition evaluation scores;

[0067] In step S103, the first speech recognition result and the second speech recognition result are mixed and sorted based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a target speech recognition result.

[0068] As mentioned above, with the advancement of science and technology, numerous smart products have emerged. To enhance the user experience and increase their intelligence, many smart products support voice interaction, which requires these smart products to have advanced speech recognition capabilities. However, in actual application scenarios, the speech content that needs to be recognized covers a wide range of fields, such as music, movies, games, place names, food, daily expressions, actions, objects, and so on, with the number of related entity words reaching into the tens of millions. If entity words are added or modified, for example, a singer releases a new song, a performer launches a new movie or TV series, a developer develops a new game, a smart express locker adds a new function, etc., in the existing technology, the training corpus vocabulary based on which the smart product voice recognition is based needs to be pronounced, updated or expanded, and the voice recognition model based on which the smart product voice recognition is based needs to be retrained. However, due to the limited capacity of the language recognition model vocabulary, the addition of new words requires certain screening, which results in certain omissions in the addition of new words. At the same time, due to changes in the vocabulary, all corpora need to be re-segmented and trained, which will greatly increase the data training workload and the time cost required, which is not conducive to improving the accuracy and efficiency of voice recognition.

[0069] In light of the above issues, this embodiment proposes a speech recognition method that performs parallel dual speech recognition based on posterior probability and category-based speech recognition on the speech data to be recognized. This method utilizes category-based speech recognition to minimize the cost of adding new words and enhance the recognition of words in certain domains. This technical solution significantly reduces the training workload and time cost of speech recognition while ensuring the comprehensiveness of new word addition, thereby improving the accuracy and efficiency of speech recognition.

[0070] In one embodiment of the present disclosure, the speech recognition method may be applicable to a terminal computer, a computing device, an electronic device, a server, a service cluster, etc. that can perform speech recognition processing.

[0071] In one embodiment of the present disclosure, the to-be-recognized voice data refers to voice data that needs to be recognized, has a certain voice data format, and is input via a voice input device such as a microphone or a receiver.

[0072] In one embodiment of the present disclosure, the first speech recognition based on posterior probability refers to a speech recognition method that uses the size of the posterior probability corresponding to the speech recognition result as a judgment criterion. In the first speech recognition based on posterior probability, one or more first speech recognition results and their corresponding first speech recognition evaluation scores, that is, the posterior probability values, will be obtained. The first speech recognition result corresponding to the highest first speech recognition evaluation score will be considered to be the final speech recognition result obtained by speech recognition based on posterior probability.

[0073] In one embodiment of the present disclosure, the category-based second speech recognition refers to a speech recognition method that uses the category corresponding to a certain word as a recognition unit to recognize it together with other independent words. The category-based second speech recognition not only considers the posterior probability contributed by the independent words in the speech recognition result, but also considers the posterior probability contributed by the word category in the speech recognition result, that is, the category probability. In the category-based second speech recognition, one or more second speech recognition results and their corresponding second speech recognition evaluation scores will be obtained, wherein the second speech recognition evaluation score is obtained by comprehensive calculation based on the category probability and the posterior probability. The second speech recognition result corresponding to the highest second speech recognition evaluation score will be considered as the final speech recognition result obtained by category-based speech recognition.

[0074] In the above embodiment, after performing a first speech recognition based on posterior probability on the speech data to be recognized to obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores, and performing a second speech recognition based on category on the speech data to be recognized to obtain one or more second speech recognition results and their corresponding second speech recognition evaluation scores, the first speech recognition results and the second speech recognition results can be mixedly sorted based on the obtained first speech recognition evaluation scores and second speech recognition evaluation scores, that is, the first speech recognition results and the second speech recognition results are sorted as a whole, and the speech recognition result corresponding to the highest speech recognition evaluation score can be considered as the final result obtained by performing speech recognition on the speech data to be recognized, that is, the target speech recognition result.

[0075] In one embodiment of the present disclosure, step S101, i.e., performing first speech recognition based on posterior probability on the speech data to be recognized to obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores, may include the following steps:

[0076] Extracting acoustic features of the speech data to be recognized;

[0077] Inputting the acoustic features into a pre-trained posterior probability prediction model to obtain a posterior probability matrix, wherein the matrix elements in the posterior probability matrix are the posterior probabilities corresponding to the acoustic features;

[0078] A beam search is performed based on the posterior probability matrix to obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores, wherein the first speech recognition evaluation scores are calculated based on the posterior probability.

[0079] In this embodiment, when performing the first speech recognition based on posterior probability on the speech data to be recognized, first, the acoustic features of the speech data to be recognized are extracted. For example, for the speech data to be recognized, one or more acoustic features can be extracted according to a fixed feature extraction frequency. The acoustic features can be, for example, Fbank features.

[0080] Then, the extracted acoustic features are input into the pre-trained posterior probability prediction model to obtain a posterior probability matrix, wherein the matrix elements in the posterior probability matrix are the posterior probability values corresponding to the acoustic features. Assuming that 10 acoustic features are extracted according to a fixed feature extraction frequency, these 10 acoustic features are input into the posterior probability prediction model. After the posterior probability prediction model, the posterior probability corresponding to each acoustic feature is obtained. Assuming that the recognition result vocabulary includes 100 words, the posterior probability corresponding to each acoustic feature is a probability vector with a dimension of 100. At this time, the posterior probability matrix is a 10*100 probability matrix, and the matrix elements in the probability matrix are the probabilities of each acoustic feature corresponding to each word in the recognition result vocabulary. Wherein, during training, the posterior probability prediction model uses the acoustic features of the speech training data as input and the speech recognition results and their corresponding posterior probabilities as output for training.

[0081] Finally, perform beam search based on the posterior probability matrix to obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores. Among them, the first speech recognition evaluation score is calculated based on the posterior probability. Assuming that the speech data to be recognized is the speech of "Who am I", the first speech recognition evaluation score is the product of the probabilities of recognizing the speech of "Who am I" as different characters in sequence according to the pronunciation order. For example, the posterior probability of recognizing the speech of "我 (wǒ)" as "我 (wǒ)" is 0.9, and the posterior probability of recognizing it as "卧 (wò)" is 0.1. Subsequently, the posterior probability of recognizing "是 (shì)" as "是 (shì)" is 0.9, and the posterior probability of recognizing it as "室 (shì)" is 0.1. The posterior probability of recognizing "谁 (shuí)" as "谁 (shuí)" is 1. Then the posterior probability of the speech recognition result "我 (wǒ)""是 (shì)""谁 (shuí)" is 0.9 * 0.9 * 1 = 0.81, while the posterior probability of the speech recognition result "我 (wǒ)""室 (shì)""谁 (shuí)" is 0.9 * 0.1 * 1 = 0.09. The posterior probability of the speech recognition result "卧 (wò)""是 (shì)""谁 (shuí)" is 0.1 * 0.9 * 1 = 0.09, and the posterior probability of the speech recognition result "卧 (wò)""室 (shì)""谁 (shuí)" is 0.1 * 0.1 * 1 = 0.01. Assuming that the speech data to be recognized is the speech of "Pick up the express delivery", the first speech recognition evaluation score is the product of the probabilities of recognizing the speech of "Pick up the express delivery" as different characters in sequence according to the pronunciation order. For example, the posterior probability of recognizing "取 (qǔ)" as "取 (qǔ)" is 0.6, and the posterior probability of recognizing it as "去 (qù)" is 0.4. The posterior probability of recognizing "快 (kuài)" as "快 (kuài)" is 0.9, and the posterior probability of recognizing it as "筷 (kuài)" is 0.1. The posterior probability of recognizing "递 (dì)" as "递 (dì)" is 0.8, and the posterior probability of recognizing it as "第 (dì)" is 0.2. Then the posterior probability of the speech recognition result "取 (qǔ)""快 (kuài)""递 (dì)" is 0.6 * 0.9 * 0.8 = 0.43, while the posterior probability of the speech recognition result "取 (qǔ)""快 (kuài)""第 (dì)" is 0.6 * 0.9 * 0.2 = 0.11. The posterior probability of the speech recognition result "取 (qǔ)""筷 (kuài)""递 (dì)" is 0.6 * 0.1 * 0.8 = 0.05, and the posterior probability of the speech recognition result "取 (qǔ)""筷 (kuài)""第 (dì)" is 0.6 * 0.1 * 0.2 = 0.01. The posterior probability of the speech recognition result "去 (qù)""快 (kuài)""递 (dì)" is 0.4 * 0.9 * 0.8 = 0.29,...... It should be noted that the beam search algorithm is a commonly used algorithm for object recognition in the prior art. Based on the scoring strategy, it can achieve parallel search and merging of multiple search spaces. The implementation principle of the beam search algorithm will not be elaborated in this disclosure. This disclosure uses the beam search algorithm for speech recognition. By using the posterior probability matrix as the input of the beam search algorithm, one or more speech recognition results corresponding to the posterior probability matrix and the posterior probability value corresponding to each speech recognition result, that is, the first speech recognition evaluation score, can be obtained.

[0082] In one embodiment of the present disclosure, step S102, the step of performing category-based second speech recognition on the speech data to be recognized to obtain one or more second speech recognition results and their corresponding second speech recognition evaluation scores, may include the following steps:

[0083] The speech data to be recognized is input into a pre-trained category speech recognition model, and one or more second speech recognition results and their corresponding second speech recognition evaluation scores are obtained based on beam search, wherein the second speech recognition evaluation scores are calculated based on the category probability and the posterior probability.

[0084] In this embodiment, when performing category-based second speech recognition on the speech data to be recognized, unlike the first speech recognition, there is no need to extract the acoustic features of the speech data to be recognized. Instead, the speech data to be recognized is directly input into a pre-trained category speech recognition model, and at the same time, combined with the waveform search algorithm, one or more second speech recognition results and their corresponding second speech recognition evaluation scores are obtained.

[0085] In another embodiment of the present disclosure, step S102, the step of performing category-based second speech recognition on the speech data to be recognized to obtain one or more second speech recognition results and their corresponding second speech recognition evaluation scores, may include the following steps:

[0086] The speech data to be recognized is input into a pre-trained main category speech recognition network. During speech recognition based on the main category speech recognition network, when it is detected that the speech unit to be recognized is a preset category content, the pre-trained auxiliary category speech recognition network is called to perform speech recognition. When recognition is completed or fails based on the auxiliary category speech recognition network, the speech recognition is continued by jumping back to the main category speech recognition network until one or more second speech recognition results and their corresponding second speech recognition evaluation scores are obtained, wherein the second speech recognition evaluation scores are calculated based on the category probability and the posterior probability.

[0087] Among them, the main category speech recognition network refers to a speech recognition model obtained by training using words labeled with general categories, and is combined with a beam search algorithm to achieve speech recognition. The auxiliary category speech recognition network refers to a speech recognition model obtained by training using words labeled with categories that require enhanced recognition, and is combined with a beam search algorithm to achieve speech recognition. There can be one or more auxiliary category speech recognition networks, and their number is related to the number of categories that require enhanced recognition, that is, each category that requires enhanced recognition corresponds to an auxiliary category speech recognition network. Based on the main category speech recognition network or the auxiliary category speech recognition network, corresponding speech recognition results and their corresponding speech recognition evaluation scores can be obtained, and the speech recognition evaluation scores are also calculated based on category probabilities and posterior probabilities.

[0088] In this embodiment, when performing category-based second speech recognition on the speech data to be recognized, the speech data to be recognized is first input into a pre-trained main category speech recognition network for speech recognition. During the speech recognition based on the main category speech recognition network, if it is detected that the speech unit to be recognized is a preset category content, that is, a category content that requires enhanced recognition, the pre-trained auxiliary category speech recognition network corresponding to the preset category content is called to perform speech recognition. When the recognition is completed based on the auxiliary category speech recognition network or it is confirmed that the speech recognition fails, it jumps back to the main category speech recognition network to continue speech recognition. The above process can be repeated multiple times until one or more final second speech recognition results and their corresponding second speech recognition evaluation scores are obtained.

[0089] In one embodiment of the present disclosure, when training the categorical speech recognition model, the categorical speech recognition model uses speech training data and category labels corresponding to related words as input, and uses the speech recognition results corresponding to the speech training data and their corresponding mixed posterior probabilities as output to train the categorical speech recognition model. The category labels corresponding to the words are obtained by querying a pre-generated category dictionary. In the category dictionary, entity words that may be used or may appear in speech recognition are classified, i.e., category annotation. For example, words in categories such as music, film and television, games, names of people, and food are classified and labeled as MUSIC, VIDEO, GAME, NAME, and FOOD, respectively. Words in categories such as actions and objects can also be classified and labeled as ACTION, GOOD, etc., respectively. In another embodiment of the present disclosure, when labeling, a word frequency weight is also attached to the category label to weight the category probability. In this way, the posterior probability corresponding to categories that appear more frequently in historical speech data is larger, further achieving enhanced categorical speech recognition. Among them, the mixed posterior probability corresponding to the speech recognition result refers to the contribution of both the labeled category and the independent words to the posterior probability of the speech recognition result in the speech recognition result, so that the category speech recognition model and the subsequent beam search can accurately express the probability distribution of the category label. For example, for the speech data "I want to listen to Dongfeng Po", the mixed posterior probability corresponding to the final speech recognition result "I want to listen to Dongfeng Po" is the mixed posterior probability of the independent words "I", "I want" and "listen" connected with the labeled category "MUSIC: Dongfeng Po". For another example, for the speech data "I want to pick up a courier", the mixed posterior probability corresponding to the final speech recognition result "I want to pick up a courier" is the mixed posterior probability of the independent words "I" and "I want" connected with the labeled categories "ACTION: pick up" and "GOOD: courier".

[0090] In one embodiment of the present disclosure, step S103, wherein the step of performing mixed sorting on the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a target speech recognition result, may include the following steps:

[0091] performing mixed sorting of the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score;

[0092] The speech recognition result with the highest speech recognition evaluation score is taken as the target speech recognition result.

[0093] As mentioned above, after obtaining the one or more first speech recognition results and the second speech recognition results and their corresponding first speech recognition evaluation scores and the second speech recognition evaluation scores, the first speech recognition results and the second speech recognition results can be mixed and sorted based on the first speech recognition evaluation scores and the second speech recognition evaluation scores, that is, the first speech recognition results and the second speech recognition results are sorted as a whole, and the speech recognition result corresponding to the highest speech recognition evaluation score can be considered as the final result obtained by speech recognition of the speech data to be recognized, that is, the target speech recognition result.

[0094] In one embodiment of the present disclosure, the first speech recognition and the second speech recognition are executed in parallel and cross-wise, that is, the first speech recognition and the second speech recognition are executed in parallel, but the speech recognition results generated by them can participate in mixed sorting at any time, and the current target speech recognition result is generated immediately. As the first speech recognition and the second speech recognition continue to be executed, the speech recognition results continue to be mixed sorted, and the target speech recognition result can be updated at any time according to the mixed sorting result until the first speech recognition and the second speech recognition are all executed. At this time, the target speech recognition result obtained is the final target speech recognition result.

[0095] That is, in this embodiment, the step of performing mixed sorting of the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score may include the following steps:

[0096] When the first speech recognition and the second speech recognition are performed in parallel, when the semantic completeness of the second speech recognition result obtained by the second speech recognition meets a preset completeness condition, the current first speech recognition result and the current second speech recognition result are mixed and sorted based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a preset number of first intermediate speech recognition results with the highest speech recognition evaluation scores;

[0097] The first speech recognition and the second speech recognition continue to be performed in parallel, and when the semantic completeness of the second speech recognition result obtained by the second speech recognition meets a preset completeness condition, the current first speech recognition result and the current second speech recognition result are mixed and sorted based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a preset number of second intermediate speech recognition results with the highest speech recognition evaluation scores, and the first intermediate speech recognition result is updated using the second intermediate speech recognition results;

[0098] Repeat the steps of generating and updating the intermediate speech recognition result until the first speech recognition and the second speech recognition are completed.

[0099] In this embodiment, the first speech recognition and the second speech recognition are performed in parallel. When it is determined that the semantic completeness of the second speech recognition result obtained by the second speech recognition meets the preset completeness condition, the second speech recognition result is considered to have a certain semantic completeness, and can be output to participate in a mixed sorting with the first speech recognition result. Then, based on the first speech recognition evaluation score and the second speech recognition evaluation score, the current first speech recognition result and the current second speech recognition result are mixed and sorted to obtain a preset number of first intermediate speech recognition results with the highest current speech recognition evaluation score, wherein the preset completeness condition can be determined according to the semantic completeness calculation method in the prior art, and the present disclosure does not make any further introduction to it; wherein, the preset number can be determined according to the needs of actual application and the characteristics of speech data, and the present disclosure does not make any specific limitation on its specific value. Afterwards, the first speech recognition and the second speech recognition continue to be executed in parallel. When the semantic completeness of the second speech recognition result obtained by the second speech recognition once again meets the preset completeness condition, the current first speech recognition result and the current second speech recognition result are mixed and sorted again based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a preset number of second intermediate speech recognition results with the highest speech recognition evaluation score. At this time, the second intermediate speech recognition result will replace the first intermediate speech recognition result. The above steps of speech recognition, mixed sorting, intermediate speech recognition result generation and update can be repeated and iteratively executed until the first speech recognition and the second speech recognition are completed. At this time, the intermediate speech recognition result obtained can be considered as the final target speech recognition result.

[0100] Figure 2 FIG. 1 shows an overall flow chart of a speech recognition method according to an embodiment of the present disclosure. Figure 2 As shown, first, the acoustic features of the speech data to be recognized are extracted, and the acoustic features are input into the pre-trained posterior probability prediction model to obtain a posterior probability matrix; beam search is performed based on the posterior probability matrix to obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores; the speech data to be recognized is input into the pre-trained category speech recognition model, and category-based second speech recognition is performed in combination with beam search to obtain one or more second speech recognition results and their corresponding second speech recognition evaluation scores; the first speech recognition results and the second speech recognition results are mixed and sorted based on the first speech recognition evaluation scores and the second speech recognition evaluation scores to obtain the target speech recognition result.

[0101] The following are embodiments of the apparatus disclosed herein, which can be used to execute embodiments of the method disclosed herein.

[0102] Figure 3FIG1 shows a structural block diagram of a speech recognition device according to an embodiment of the present disclosure. The device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. Figure 3 As shown, the speech recognition device includes:

[0103] The first recognition module 301 is configured to perform first speech recognition based on posterior probability on the speech data to be recognized, and obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores;

[0104] The second recognition module 302 is configured to perform category-based second speech recognition on the speech data to be recognized, and obtain one or more second speech recognition results and their corresponding second speech recognition evaluation scores;

[0105] The mixing module 303 is configured to perform mixed sorting on the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a target speech recognition result.

[0106] As mentioned above, with the advancement of science and technology, numerous smart products have emerged. To enhance the user experience and increase their intelligence, many smart products support voice interaction, which requires these smart products to have advanced speech recognition capabilities. However, in actual application scenarios, the speech content that needs to be recognized spans a wide range of fields, such as music, film, games, place names, food, and daily expressions, and the number of related entity words is in the tens of millions. If entity words are added or modified, for example, a singer releases a new song, a performer launches a new movie or TV series, a developer develops a new game, etc., in the existing technology, the training corpus vocabulary based on which the smart product voice recognition is based needs to be pronounced, updated or expanded, and the voice recognition model based on which the smart product voice recognition is based needs to be retrained. However, due to the limited capacity of the language recognition model vocabulary, the addition of new words requires certain screening, which results in certain omissions in the addition of new words. At the same time, due to changes in the vocabulary, all corpora need to be re-segmented and trained, which will greatly increase the data training workload and the time cost required, which is not conducive to improving the accuracy and efficiency of voice recognition.

[0107] In light of the above issues, this embodiment proposes a speech recognition device that performs parallel dual speech recognition based on posterior probability and category-based speech recognition on the speech data to be recognized. This device utilizes category-based speech recognition to minimize the cost of adding new words and enhance the recognition of words in certain domains. This technical solution significantly reduces the training workload and time cost of speech recognition while ensuring the comprehensiveness of new word addition, thereby improving the accuracy and efficiency of speech recognition.

[0108] In one embodiment of the present disclosure, the speech recognition apparatus may be implemented as a terminal computer, a computing device, an electronic device, a server, a service cluster, etc. that can perform speech recognition processing.

[0109] In one embodiment of the present disclosure, the to-be-recognized voice data refers to voice data that needs to be recognized, has a certain voice data format, and is input via a voice input device such as a microphone or a receiver.

[0110] In one embodiment of the present disclosure, the first speech recognition based on posterior probability refers to a speech recognition method that uses the size of the posterior probability corresponding to the speech recognition result as a judgment criterion. In the first speech recognition based on posterior probability, one or more first speech recognition results and their corresponding first speech recognition evaluation scores, that is, the posterior probability values, will be obtained. The first speech recognition result corresponding to the highest first speech recognition evaluation score will be considered to be the final speech recognition result obtained by speech recognition based on posterior probability.

[0111] In one embodiment of the present disclosure, the category-based second speech recognition refers to a speech recognition method that uses the category corresponding to a certain word as a recognition unit to recognize it together with other independent words. The category-based second speech recognition not only considers the posterior probability contributed by the independent words in the speech recognition result, but also considers the posterior probability contributed by the word category in the speech recognition result, that is, the category probability. In the category-based second speech recognition, one or more second speech recognition results and their corresponding second speech recognition evaluation scores will be obtained, wherein the second speech recognition evaluation score is obtained by comprehensive calculation based on the category probability and the posterior probability. The second speech recognition result corresponding to the highest second speech recognition evaluation score will be considered as the final speech recognition result obtained by category-based speech recognition.

[0112] In the above embodiment, after performing a first speech recognition based on posterior probability on the speech data to be recognized to obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores, and performing a second speech recognition based on category on the speech data to be recognized to obtain one or more second speech recognition results and their corresponding second speech recognition evaluation scores, the first speech recognition results and the second speech recognition results can be mixedly sorted based on the obtained first speech recognition evaluation scores and second speech recognition evaluation scores, that is, the first speech recognition results and the second speech recognition results are sorted as a whole, and the speech recognition result corresponding to the highest speech recognition evaluation score can be considered as the final result obtained by performing speech recognition on the speech data to be recognized, that is, the target speech recognition result.

[0113] In one embodiment of the present disclosure, the first identification module 301 may be configured as follows:

[0114] Extracting acoustic features of the speech data to be recognized;

[0115] Inputting the acoustic features into a pre-trained posterior probability prediction model to obtain a posterior probability matrix, wherein the matrix elements in the posterior probability matrix are the posterior probabilities corresponding to the acoustic features;

[0116] A beam search is performed based on the posterior probability matrix to obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores, wherein the first speech recognition evaluation scores are calculated based on the posterior probability.

[0117] In this embodiment, when performing the first speech recognition based on posterior probability on the speech data to be recognized, first, the acoustic features of the speech data to be recognized are extracted. For example, for the speech data to be recognized, one or more acoustic features can be extracted according to a fixed feature extraction frequency. The acoustic features can be, for example, Fbank features.

[0118] Then, the extracted acoustic features are input into the pre-trained posterior probability prediction model to obtain a posterior probability matrix, wherein the matrix elements in the posterior probability matrix are the posterior probability values corresponding to the acoustic features. Assuming that 10 acoustic features are extracted according to a fixed feature extraction frequency, these 10 acoustic features are input into the posterior probability prediction model. After the posterior probability prediction model, the posterior probability corresponding to each acoustic feature is obtained. Assuming that the recognition result vocabulary includes 100 words, the posterior probability corresponding to each acoustic feature is a probability vector with a dimension of 100. At this time, the posterior probability matrix is a 10*100 probability matrix, and the matrix elements in the probability matrix are the probabilities of each acoustic feature corresponding to each word in the recognition result vocabulary. Wherein, during training, the posterior probability prediction model uses the acoustic features of the speech training data as input and the speech recognition results and their corresponding posterior probabilities as output for training.

[0119] Finally, beam search is performed based on the posterior probability matrix to obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores. Among them, the first speech recognition evaluation score is calculated based on the posterior probability. Assuming that the speech data to be recognized is the speech of "Who am I", the first speech recognition evaluation score is the product of the probabilities of recognizing the speech of "Who am I" as different characters in sequence according to the pronunciation order. For example, the posterior probability of recognizing the speech of "我 (wǒ)" as "我 (wǒ)" is 0.9, and the posterior probability of recognizing it as "卧 (wò)" is 0.1. Subsequently, the posterior probability of recognizing "是 (shì)" as "是 (shì)" is 0.9, and the posterior probability of recognizing it as "室 (shì)" is 0.1. The posterior probability of recognizing "谁 (shuí)" as "谁 (shuí)" is 1. Then, the posterior probability of the speech recognition result "我 (wǒ)""是 (shì)""谁 (shuí)" is 0.9 * 0.9 * 1 = 0.81, while the posterior probability of the speech recognition result "我 (wǒ)""室 (shì)""谁 (shuí)" is 0.9 * 0.1 * 1 = 0.09. The posterior probability of the speech recognition result "卧 (wò)""是 (shì)""谁 (shuí)" is 0.1 * 0.9 * 1 = 0.09, and the posterior probability of the speech recognition result "卧 (wò)""室 (shì)""谁 (shuí)" is 0.1 * 0.1 * 1 = 0.01. Assuming that the speech data to be recognized is the speech of "Pick up the express delivery", the first speech recognition evaluation score is the product of the probabilities of recognizing the speech of "Pick up the express delivery" as different characters in sequence according to the pronunciation order. For example, the posterior probability of recognizing "取 (qǔ)" as "取 (qǔ)" is 0.6, and the posterior probability of recognizing it as "去 (qù)" is 0.4. The posterior probability of recognizing "快 (kuài)" as "快 (kuài)" is 0.9, and the posterior probability of recognizing it as "筷 (kuài)" is 0.1. The posterior probability of recognizing "递 (dì)" as "递 (dì)" is 0.8, and the posterior probability of recognizing it as "第 (dì)" is 0.2. Then, the posterior probability of the speech recognition result "取 (qǔ)""快 (kuài)""递 (dì)" is 0.6 * 0.9 * 0.8 = 0.43, while the posterior probability of the speech recognition result "取 (qǔ)""快 (kuài)""第 (dì)" is 0.6 * 0.9 * 0.2 = 0.11. The posterior probability of the speech recognition result "取 (qǔ)""筷 (kuài)""递 (dì)" is 0.6 * 0.1 * 0.8 = 0.05, and the posterior probability of the speech recognition result "取 (qǔ)""筷 (kuài)""第 (dì)" is 0.6 * 0.1 * 0.2 = 0.01. The posterior probability of the speech recognition result "去 (qù)""快 (kuài)""递 (dì)" is 0.4 * 0.9 * 0.8 = 0.29,...... It should be noted that the beam search algorithm is a commonly used algorithm for object recognition in the prior art. Based on the scoring strategy, parallel search and merging of multiple search spaces can be achieved. The implementation principle of the beam search algorithm will not be elaborated in this disclosure. This disclosure uses the beam search algorithm for speech recognition. By using the posterior probability matrix as the input of the beam search algorithm, one or more speech recognition results corresponding to the posterior probability matrix and the posterior probability value corresponding to each speech recognition result, that is, the first speech recognition evaluation score, can be obtained.

[0120] In an embodiment of the present disclosure, the second recognition module 302 may be configured to:

[0121] The speech data to be recognized is input into a pre-trained category speech recognition model, and one or more second speech recognition results and their corresponding second speech recognition evaluation scores are obtained based on beam search, wherein the second speech recognition evaluation scores are calculated based on the category probability and the posterior probability.

[0122] In this embodiment, when performing category-based second speech recognition on the speech data to be recognized, unlike the first speech recognition, there is no need to extract the acoustic features of the speech data to be recognized. Instead, the speech data to be recognized is directly input into a pre-trained category speech recognition model, and at the same time, combined with the waveform search algorithm, one or more second speech recognition results and their corresponding second speech recognition evaluation scores are obtained.

[0123] In another embodiment of the present disclosure, the second identification module 302 may also be configured to:

[0124] The speech data to be recognized is input into a pre-trained main category speech recognition network. During speech recognition based on the main category speech recognition network, when it is detected that the speech unit to be recognized is a preset category content, the pre-trained auxiliary category speech recognition network is called to perform speech recognition. When recognition is completed or fails based on the auxiliary category speech recognition network, the speech recognition is continued by jumping back to the main category speech recognition network until one or more second speech recognition results and their corresponding second speech recognition evaluation scores are obtained, wherein the second speech recognition evaluation scores are calculated based on the category probability and the posterior probability.

[0125] Among them, the main category speech recognition network refers to a speech recognition model obtained by training using words labeled with general categories, and is combined with a beam search algorithm to achieve speech recognition. The auxiliary category speech recognition network refers to a speech recognition model obtained by training using words labeled with categories that require enhanced recognition, and is combined with a beam search algorithm to achieve speech recognition. There can be one or more auxiliary category speech recognition networks, and their number is related to the number of categories that require enhanced recognition, that is, each category that requires enhanced recognition corresponds to an auxiliary category speech recognition network. Based on the main category speech recognition network or the auxiliary category speech recognition network, corresponding speech recognition results and their corresponding speech recognition evaluation scores can be obtained, and the speech recognition evaluation scores are also calculated based on category probabilities and posterior probabilities.

[0126] In this embodiment, when performing category-based second speech recognition on the speech data to be recognized, the speech data to be recognized is first input into a pre-trained main category speech recognition network for speech recognition. During the speech recognition based on the main category speech recognition network, if it is detected that the speech unit to be recognized is a preset category content, that is, a category content that requires enhanced recognition, the pre-trained auxiliary category speech recognition network corresponding to the preset category content is called to perform speech recognition. When the recognition is completed based on the auxiliary category speech recognition network or it is confirmed that the speech recognition fails, it jumps back to the main category speech recognition network to continue speech recognition. The above process can be repeated multiple times until one or more final second speech recognition results and their corresponding second speech recognition evaluation scores are obtained.

[0127] In one embodiment of the present disclosure, when training the categorical speech recognition model, the categorical speech recognition model uses speech training data and category labels corresponding to related words as input, and uses the speech recognition results corresponding to the speech training data and their corresponding mixed posterior probabilities as output to train the categorical speech recognition model. The category labels corresponding to the words are obtained by querying a pre-generated category dictionary. In the category dictionary, entity words that may be used or may appear in speech recognition are classified, i.e., category annotation. For example, words in categories such as music, film and television, games, names of people, and food are classified and labeled as MUSIC, VIDEO, GAME, NAME, and FOOD, respectively. Words in categories such as actions and objects can also be classified and labeled as ACTION, GOOD, etc., respectively. In another embodiment of the present disclosure, when labeling, a word frequency weight is also attached to the category label to weight the category probability. In this way, the posterior probability corresponding to categories that appear more frequently in historical speech data is larger, further achieving enhanced categorical speech recognition. Among them, the mixed posterior probability corresponding to the speech recognition result refers to the contribution of both the labeled category and the independent words to the posterior probability of the speech recognition result in the speech recognition result, so that the category speech recognition model and the subsequent beam search can accurately express the probability distribution of the category label. For example, for the speech data "I want to listen to Dongfeng Po", the mixed posterior probability corresponding to the final speech recognition result "I want to listen to Dongfeng Po" is the mixed posterior probability of the independent words "I", "I want" and "listen" connected with the labeled category "MUSIC: Dongfeng Po". For another example, for the speech data "I want to pick up a courier", the mixed posterior probability corresponding to the final speech recognition result "I want to pick up a courier" is the mixed posterior probability of the independent words "I" and "I want" connected with the labeled categories "ACTION: pick up" and "GOOD: courier".

[0128] In one embodiment of the present disclosure, the mixing module 303 may be configured as follows:

[0129] performing mixed sorting of the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score;

[0130] The speech recognition result with the highest speech recognition evaluation score is taken as the target speech recognition result.

[0131] As mentioned above, after obtaining the one or more first speech recognition results and the second speech recognition results and their corresponding first speech recognition evaluation scores and the second speech recognition evaluation scores, the first speech recognition results and the second speech recognition results can be mixed and sorted based on the first speech recognition evaluation scores and the second speech recognition evaluation scores, that is, the first speech recognition results and the second speech recognition results are sorted as a whole, and the speech recognition result corresponding to the highest speech recognition evaluation score can be considered as the final result obtained by speech recognition of the speech data to be recognized, that is, the target speech recognition result.

[0132] In one embodiment of the present disclosure, the first speech recognition and the second speech recognition are executed in parallel and cross-wise, that is, the first speech recognition and the second speech recognition are executed in parallel, but the speech recognition results generated by them can participate in mixed sorting at any time, and the current target speech recognition result is generated immediately. As the first speech recognition and the second speech recognition continue to be executed, the speech recognition results continue to be mixed sorted, and the target speech recognition result can be updated at any time according to the mixed sorting result until the first speech recognition and the second speech recognition are all executed. At this time, the target speech recognition result obtained is the final target speech recognition result.

[0133] That is, in this embodiment, the step of performing mixed sorting of the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score may include the following steps:

[0134] When the first speech recognition and the second speech recognition are performed in parallel, when the semantic completeness of the second speech recognition result obtained by the second speech recognition meets a preset completeness condition, the current first speech recognition result and the current second speech recognition result are mixed and sorted based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a preset number of first intermediate speech recognition results with the highest speech recognition evaluation scores;

[0135] The first speech recognition and the second speech recognition continue to be performed in parallel, and when the semantic completeness of the second speech recognition result obtained by the second speech recognition meets a preset completeness condition, the current first speech recognition result and the current second speech recognition result are mixed and sorted based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a preset number of second intermediate speech recognition results with the highest speech recognition evaluation scores, and the first intermediate speech recognition result is updated using the second intermediate speech recognition results;

[0136] The generation and updating of the intermediate speech recognition result are repeatedly performed until the first speech recognition and the second speech recognition are completed.

[0137] In this embodiment, the first speech recognition and the second speech recognition are performed in parallel. When it is determined that the semantic completeness of the second speech recognition result obtained by the second speech recognition meets the preset completeness condition, the second speech recognition result is considered to have a certain semantic completeness, and can be output to participate in a mixed sorting with the first speech recognition result. Then, based on the first speech recognition evaluation score and the second speech recognition evaluation score, the current first speech recognition result and the current second speech recognition result are mixed and sorted to obtain a preset number of first intermediate speech recognition results with the highest current speech recognition evaluation score, wherein the preset completeness condition can be determined according to the semantic completeness calculation method in the prior art, and the present disclosure does not make any further introduction to it; wherein, the preset number can be determined according to the needs of actual application and the characteristics of speech data, and the present disclosure does not make any specific limitation on its specific value. Afterwards, the first speech recognition and the second speech recognition continue to be executed in parallel. When the semantic completeness of the second speech recognition result obtained by the second speech recognition once again meets the preset completeness condition, the current first speech recognition result and the current second speech recognition result are mixed and sorted again based on the first speech recognition evaluation score and the second speech recognition evaluation score to obtain a preset number of second intermediate speech recognition results with the highest speech recognition evaluation scores. At this time, the second intermediate speech recognition result will replace the first intermediate speech recognition result. The above-mentioned speech recognition, mixed sorting, generation and update of intermediate speech recognition results can be repeated and iteratively executed until the first speech recognition and the second speech recognition are completed. At this time, the intermediate speech recognition result obtained can be considered as the final target speech recognition result.

[0138] The present disclosure also discloses an electronic device, which includes a memory and a processor; wherein:

[0139] The memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement any of the above method steps.

[0140] Figure 4 It is a structural diagram of a computer system suitable for implementing the speech recognition method according to an embodiment of the present disclosure.

[0141] like Figure 4 As shown, computer system 400 includes a processing unit 401, which can execute various processes in the above-mentioned embodiments according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage unit 408 into a random access memory (RAM) 403. Various programs and data required for the operation of system 400 are also stored in RAM 403. Processing unit 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to bus 404.

[0142] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read therefrom can be installed into the storage section 408 as needed. Among them, the processing unit 401 can be implemented as a processing unit such as a CPU, a GPU, a TPU, an FPGA, an NPU, etc.

[0143] In particular, according to embodiments of the present disclosure, the method described above can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program tangibly embodied on a computer-readable medium, the computer program comprising program code for executing the speech recognition method. In such embodiments, the computer program can be downloaded and installed from a network via the communication portion 409 and / or installed from the removable medium 411.

[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the diagram or block diagram can represent a module, program segment or part of the code, and the module, program segment or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, as well as the combination of boxes in the block diagram and / or flow chart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.

[0145] The units or modules described in the embodiments of the present disclosure may be implemented in software or hardware. The units or modules described may also be provided in a processor, and the names of these units or modules do not, in certain circumstances, limit the units or modules themselves.

[0146] As another aspect, embodiments of the present disclosure further provide a computer-readable storage medium. This computer-readable storage medium may be included in the apparatus described in the above embodiments, or may be a standalone computer-readable storage medium not incorporated into the apparatus. The computer-readable storage medium stores one or more programs, which are used by one or more processors to execute the methods described in the embodiments of the present disclosure.

[0147] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also encompass other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A speech recognition method, comprising: performing first speech recognition based on posterior probability on the speech data to be recognized, obtaining one or more first speech recognition results and their corresponding first speech recognition evaluation scores; performing second speech recognition based on categories on the speech data to be recognized, obtaining one or more second speech recognition results and their corresponding second speech recognition evaluation scores; performing hybrid sorting on the first speech recognition results and the second speech recognition results based on the first speech recognition evaluation scores and the second speech recognition evaluation scores, obtaining a target speech recognition result; wherein, the performing hybrid sorting on the first speech recognition results and the second speech recognition results based on the first speech recognition evaluation scores and the second speech recognition evaluation scores includes: when the first speech recognition and the second speech recognition are executed in parallel, when the semantic integrity of the second speech recognition result obtained by the second speech recognition meets a preset integrity condition, performing hybrid sorting on the current first speech recognition result and the current second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score, obtaining a preset number of first intermediate speech recognition results with the highest speech recognition evaluation scores; continuing to execute the first speech recognition and the second speech recognition in parallel, when the semantic integrity of the second speech recognition result obtained by the second speech recognition meets the preset integrity condition, performing hybrid sorting on the current first speech recognition result and the current second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score, obtaining a preset number of second intermediate speech recognition results with the highest speech recognition evaluation scores, and updating the first intermediate speech recognition results with the second intermediate speech recognition results; repeating the steps of generating and updating intermediate speech recognition results until the first speech recognition and the second speech recognition are completed.

2. The method according to claim 1, wherein the performing first speech recognition based on posterior probability on the speech data to be recognized, obtaining one or more first speech recognition results and their corresponding first speech recognition evaluation scores, comprising: extracting acoustic features of the speech data to be recognized; inputting the acoustic features into a pre-trained posterior probability prediction model, obtaining a posterior probability matrix, wherein matrix elements in the posterior probability matrix are posterior probabilities corresponding to the acoustic features; performing beam search based on the posterior probability matrix, obtaining one or more first speech recognition results and their corresponding first speech recognition evaluation scores, wherein the first speech recognition evaluation scores are calculated based on the posterior probabilities.

3. The method according to claim 1 or 2, wherein the performing second speech recognition based on categories on the speech data to be recognized, obtaining one or more second speech recognition results and their corresponding second speech recognition evaluation scores, comprising: Input the speech data to be recognized into a pre-trained category speech recognition model, and obtain one or more second speech recognition results and their corresponding second speech recognition evaluation scores based on beam search, where the second speech recognition evaluation score is calculated based on the category probability and the posterior probability.

4. The method according to claim 1 or 2, wherein for the speech data to be recognized, perform category-based second speech recognition to obtain one or more second speech recognition results and their corresponding second speech recognition evaluation scores. including: Input the speech data to be recognized into a pre-trained main category speech recognition network. During the process of speech recognition based on the main category speech recognition network, when it is detected that the speech unit to be recognized is preset category content, call a pre-trained auxiliary category speech recognition network to perform speech recognition. When the recognition is completed or fails based on the auxiliary category speech recognition network, jump back to the main category speech recognition network to continue the speech recognition until one or more second speech recognition results and their corresponding second speech recognition evaluation scores are obtained, where the second speech recognition evaluation score is calculated based on the category probability and the posterior probability.

5. The method according to claim 3, when training the category speech recognition model, use speech training data and category labels corresponding to relevant words as input, and use the speech recognition results corresponding to the speech training data and their corresponding mixed posterior probabilities as output to train the category speech recognition model. wherein, The category labels corresponding to the words are obtained by querying a pre-generated category dictionary.

6. The method according to claim 4, when training the category speech recognition model, use speech training data and category labels corresponding to relevant words as input, and use the speech recognition results corresponding to the speech training data and their corresponding mixed posterior probabilities as output to train the category speech recognition model. wherein, The category labels corresponding to the words are obtained by querying a pre-generated category dictionary.

7. The method according to claim 5 or 6, wherein the category label is also attached with a word frequency weight to weight the category probability.

8. The method according to claim 1, wherein based on the first speech recognition evaluation score and the second speech recognition evaluation score, perform a mixed sorting on the first speech recognition result and the second speech recognition result to obtain a target speech recognition result. including: Perform a mixed sorting on the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score; Take the speech recognition result with the highest speech recognition evaluation score as the target speech recognition result.

9. A speech recognition device including: A first recognition module configured to perform first speech recognition based on the posterior probability on the speech data to be recognized, and obtain one or more first speech recognition results and their corresponding first speech recognition evaluation scores; A second recognition module, configured to perform second voice recognition based on categories on the voice data to be recognized, and obtain one or more second voice recognition results and their corresponding second voice recognition evaluation scores; A mixing module, configured to perform mixed sorting on the first voice recognition results and the second voice recognition results based on the first voice recognition evaluation score and the second voice recognition evaluation score, to obtain a target voice recognition result; Wherein, the part of performing mixed sorting on the first voice recognition results and the second voice recognition results based on the first voice recognition evaluation score and the second voice recognition evaluation score is configured to: When the first voice recognition and the second voice recognition are executed in parallel, when the semantic integrity of the second voice recognition result obtained by the second voice recognition meets a preset integrity condition, perform mixed sorting on the current first voice recognition result and the current second voice recognition result based on the first voice recognition evaluation score and the second voice recognition evaluation score, to obtain a preset number of first intermediate voice recognition results with the highest voice recognition evaluation scores; The first voice recognition and the second voice recognition continue to be executed in parallel. When the semantic integrity of the second voice recognition result obtained by the second voice recognition meets the preset integrity condition, perform mixed sorting on the current first voice recognition result and the current second voice recognition result based on the first voice recognition evaluation score and the second voice recognition evaluation score, to obtain a preset number of second intermediate voice recognition results with the highest voice recognition evaluation scores, and use the second intermediate voice recognition results to update the first intermediate voice recognition results; Repeat the generation and update of the intermediate voice recognition results until the first voice recognition and the second voice recognition are completed.

10. The apparatus according to claim 9, wherein the first recognition module is configured to: Extract acoustic features of the voice data to be recognized; Input the acoustic features into a posterior probability prediction model obtained by pre-training to obtain a posterior probability matrix, Wherein, The matrix elements in the posterior probability matrix are the posterior probabilities corresponding to the acoustic features; Perform beam search based on the posterior probability matrix to obtain one or more first voice recognition results and their corresponding first voice recognition evaluation scores, wherein the first voice recognition evaluation score is calculated based on the posterior probability.

11. The apparatus according to claim 9 or 10, wherein the second recognition module is configured to: Input the voice data to be recognized into a category voice recognition model obtained by pre-training, and obtain one or more second voice recognition results and their corresponding second voice recognition evaluation scores based on beam search, Wherein, The second voice recognition evaluation score is calculated based on category probabilities and the posterior probability.

12. The apparatus according to claim 9 or 10, wherein the second recognition module is configured to: Input the speech data to be recognized into a pre-trained main-category speech recognition network. During the process of speech recognition based on the main-category speech recognition network, when it is detected that the speech unit to be recognized is content of a preset category, call a pre-trained secondary-category speech recognition network to perform speech recognition. When the recognition is completed or fails based on the secondary-category speech recognition network, jump back to the main-category speech recognition network to continue speech recognition until one or more second speech recognition results and their corresponding second speech recognition evaluation scores are obtained. Wherein, The second speech recognition evaluation score is calculated based on the category probability and the posterior probability.

13. The apparatus according to claim 11, when the category speech recognition model is trained, using speech training data and category labels corresponding to relevant words as inputs, and using the speech recognition results corresponding to the speech training data and their corresponding mixed posterior probabilities as outputs to train the category speech recognition model. Wherein, The category labels corresponding to the words are obtained by querying a pre-generated category dictionary.

14. The apparatus according to claim 12, when the category speech recognition model is trained, using speech training data and category labels corresponding to relevant words as inputs, and using the speech recognition results corresponding to the speech training data and their corresponding mixed posterior probabilities as outputs to train the category speech recognition model. Wherein, The category labels corresponding to the words are obtained by querying a pre-generated category dictionary.

15. The apparatus according to claim 13 or 14, wherein the category label is also attached with a word frequency weight to weight the category probability.

16. The apparatus according to claim 9, wherein the mixing module is configured to: Perform a mixed sorting on the first speech recognition result and the second speech recognition result based on the first speech recognition evaluation score and the second speech recognition evaluation score; Use the speech recognition result with the highest speech recognition evaluation score as the target speech recognition result.

17. An electronic device, comprising a memory and a processor; Wherein, The memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to any one of claims 1-8.

18. A computer-readable storage medium, on which computer instructions are stored, Wherein, When the computer instructions are executed by a processor, the method according to any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Voice search method and device and voice recognition system

    CN108899013A

  • Speech recognition method, device and equipment, and computer readable storage medium

    CN110534095A

  • Voice recognition method and device and storage medium

    CN110797026A