Speech recognition method, apparatus, storage medium, and device

By introducing an error index label training set into the mixed-language speech recognition model, summarizing error types, and performing prediction tasks, the accuracy problem of mixed-language speech recognition models is solved, and the accuracy and performance of mixed-language speech recognition are improved.

CN115810348BActive Publication Date: 2025-12-19BEIJING YUANLI WEILAI SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111070864.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-13
Publication Date
2025-12-19
Estimated Expiration
2041-09-13

AI Technical Summary

Technical Problem

The recognition performance of mixed-language speech recognition models is worse than that of single-language models, making it difficult to meet the needs of mixed-language communication scenarios. This is mainly due to the scarcity of training data, which leads to poor recognition accuracy.

Method used

A pre-defined speech recognition model, including at least one independent target decoding module, is adopted. It is trained based on a training set of error index identifiers. By summarizing error types and introducing a prediction task of error index sequences into the decoding module, the speech recognition accuracy is improved.

Benefits of technology

Despite the scarcity of training data, this study significantly improved the recognition accuracy of the mixed-language speech recognition model, resolved recognition errors caused by the scarcity of mixed-language data, and enhanced overall recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115810348B_ABST
    Figure CN115810348B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a speech recognition method and device, a storage medium and equipment, and mainly aims to solve the problem of poor accuracy of speech recognition in the current speech recognition process. The method comprises the following steps: acquiring speech information, wherein the speech information comprises speech in at least two languages; identifying the speech information based on a preset speech recognition model to determine a speech recognition result, wherein the preset speech recognition model comprises at least one independent target decoding module, and the target decoding module is obtained based on at least a decoding module trained by a training set comprising an error index identifier.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of speech recognition, and in particular to a speech recognition method and device, a storage medium and an apparatus. BACKGROUND

[0002] With the intensification of globalization, mixed language communication is reflected in politics, economy, sports, culture and other aspects, from official talks to daily oral communication. Speech recognition technology converts audio signals into more intuitive text records, and has rapidly developed in recent years, giving rise to a large number of applications, such as video captioning, automatic meeting minutes, etc., greatly facilitating people's life. At the same time, it further improves the requirements for the recognition performance of mixed language speech recognition systems.

[0003] However, due to the flexibility of mixed language communication and the high cost of mixed language audio transcription, training data is scarce. Therefore, the recognition effect of the mixed language speech recognition model is often poorer than that of the single language speech recognition model, and it is difficult to meet the existing mixed language communication scenarios. SUMMARY

[0004] In view of the above problems, embodiments of the present application provide a speech recognition method and device, a storage medium and an apparatus, the main purpose of which is to solve the problem of poor accuracy of speech recognition in the current speech recognition process.

[0005] To solve the above technical problems, in a first aspect, embodiments of the present application provide a speech recognition method, which can include:

[0006] obtaining speech information, the speech information including speech of at least two languages;

[0007] identifying the speech information based on a preset speech recognition model to determine a speech recognition result, wherein the preset speech recognition model includes at least one independent target decoding module, and the target decoding module is obtained based on a decoding module trained by a training set including error index labels.

[0008] In a first possible implementation of the first aspect, the number of error index labels is determined based on the number of language types.

[0009] In a second possible implementation of the first aspect, the error type of the error index label is obtained by comparing the initial recognition result of the initial speech recognition model with the training data.

[0010] In a third possible implementation of the first aspect, the error type of the error index label includes same language recognition error of the target language, cross language recognition error of the target language, insertion error of the target language and deletion error of the target language.

[0011] In a fourth possible implementation form of the first aspect as such, the target decoding module is obtained based on at least a training set comprising error index identification and language identification.

[0012] In a fifth possible implementation form of the first aspect as such, the preset speech recognition model further comprises at least one basic decoding module, and the method further comprises:

[0013] training the weights of each decoding module in the preset speech recognition model based on a preset data training set;

[0014] the decoding module obtains a decoding result of each decoding module respectively, and determines a speech recognition result based on the obtained decoding result of each decoding module and the corresponding weight.

[0015] In a sixth possible implementation form of the first aspect as such, the training set comprises a single-language data set and a mixed-language data set, and the single-language data set is less than or equal to ten orders of magnitude of the mixed-language data set.

[0016] In a seventh possible implementation form of the first aspect as such, the method further comprises:

[0017] selecting data with a word frequency less than or equal to an initial word frequency threshold in the single-language corpus data to add to the single-language data set;

[0018] if the single-language data set is still less than the preset multiple of the mixed-language data set after traversing the single-language corpus data, increasing the initial word frequency threshold to continue traversing the single-language corpus data to increase the single-language data set.

[0019] In a second aspect, the embodiments of the present application further provide a speech recognition device, which can comprise: an obtaining unit, an identifying unit and a determining unit,

[0020] The obtaining unit can be configured to obtain speech information, and the speech information comprises speech of at least two languages.

[0021] The identifying unit can be configured to identify the speech information based on a preset speech recognition model to determine a speech recognition result, wherein the preset speech recognition model comprises at least one independent target decoding module, and the target decoding module is a decoding module obtained based on at least a training set comprising error index identification.

[0022] In a first possible implementation form of the second aspect as such, the number of error index identifications is determined based on the number of language types.

[0023] In a second possible implementation manner of the second aspect, the error type identified by the error index is obtained based on a comparison between an initial recognition result of the initial speech recognition model and the training data.

[0024] In a third possible implementation manner of the second aspect, the error type identified by the error index includes a same-language recognition error of the target language, a cross-language recognition error of the target language, an insertion error of the target language, and a deletion error of the target language.

[0025] In a fourth possible implementation manner of the second aspect, the target decoding module is trained based on at least a training set including the error index and the language recognition index.

[0026] In a fifth possible implementation manner of the second aspect, the preset speech recognition model further includes at least one basic decoding module, and the determination unit is further configured to:

[0027] determine the weight of each decoding module based on a training result of the preset speech recognition model based on a preset data training set;

[0028] obtain a decoding result of each decoding module respectively, and determine the speech recognition result based on the obtained decoding result of each decoding module and the corresponding weight.

[0029] In a sixth possible implementation manner of the second aspect, the training set includes a single-language data set and a mixed-language data set, and the single-language data set is less than or equal to ten orders of magnitude of the mixed-language data set.

[0030] In a seventh possible implementation manner of the second aspect, the obtaining unit is further configured to:

[0031] select data with a word frequency less than or equal to an initial word frequency threshold from the single-language corpus data to add to the single-language data set;

[0032] if the single-language data set is still less than the preset multiple of the mixed-language data set after traversing the single-language corpus data, increase the initial word frequency threshold to continue traversing the single-language corpus data, so as to increase the single-language data set.

[0033] To achieve the above object, according to a third aspect of an embodiment of the present application, a storage medium is provided, which can include a stored program, wherein when the program runs, the device where the storage medium is located is controlled to execute the speech recognition method in any one of the first aspect.

[0034] To achieve the above object, according to a fourth aspect of the embodiments of the present application, an electronic device is provided, the device comprising at least one processor, and at least one memory connected with the processor; the processor is configured to invoke program instructions in the memory, and execute the speech recognition method according to any one of the first aspect.

[0035] By the above technical solution, the speech recognition method provided by the embodiments of the present application acquires speech information, the speech information comprising speech in at least two languages. The speech information is recognized based on a preset speech recognition model to determine a speech recognition result, wherein the preset speech recognition model comprises at least one independent target decoding module, and the target decoding module is obtained based on a decoding module trained by a training set comprising an error index identifier. Due to the data-oriented characteristics of the speech recognition system, the more complete the training data and the larger the data volume, the better the performance of the recognition system. In order to ensure the accuracy of the speech recognition model in the case of scarce training data, the error types are summarized based on the recognized errors, and the error index is introduced into the decoding module of the speech recognition model, thereby adding a prediction task of the error index sequence. After training by the training set comprising the error index identifier, the speech recognition accuracy of the speech recognition model can be significantly improved. Especially for the misrecognition problem caused by the different difficulty levels of recognizing the same word in different mixed language corpus sequences. And under the premise of ensuring the training cost, data labeling cost and limited training data, various recognition errors caused by the scarcity of mixed language corpus are solved, thereby improving the recognition performance of the overall speech recognition model.

[0036] The speech recognition device, storage medium and electronic device described above also have corresponding effects due to the adoption of the speech recognition method described above. The above description is only a summary of the technical solutions of the embodiments of the present application. In order to make the technical means of the embodiments of the present application more clear, the embodiments of the present application can be implemented according to the content of the description. And in order to make the above and other objects, features and advantages of the embodiments of the present application more obvious and easy to understand, the specific embodiments of the embodiments of the present application are described below. BRIEF DESCRIPTION OF DRAWINGS

[0037] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments and are not meant to limit the present application. Moreover, the same reference numerals in the attached drawings indicate the same or similar components. In the drawings:

[0038] Figure 1 A schematic flowchart of a speech recognition method provided by the embodiments of the present application is shown;

[0039] Figure 2 Fig. 1 shows a schematic framework diagram of a preset speech recognition model according to an embodiment of the present application;

[0040] Figure 3 Fig. 2 shows a schematic framework diagram of another preset speech recognition model according to an embodiment of the present application;

[0041] Figure 4 Fig. 3 shows a schematic structural block diagram of a speech recognition device according to an embodiment of the present application;

[0042] Figure 5 Fig. 4 shows a schematic structural block diagram of an electronic device for speech recognition according to an embodiment of the present application. DETAILED DESCRIPTION

[0043] Exemplary embodiments of the present application will be described in detail with reference to the drawings. Although exemplary embodiments of the present application are shown in the drawings, it is understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the application to those skilled in the art.

[0044] To solve the problem of poor accuracy of speech recognition in the current speech recognition process, the present application provides a speech recognition method, as shown in Figure 1 The method can include S110 and S120.

[0045] S110: Obtain speech information, the speech information including speech in at least two languages.

[0046] It should be noted that the speech information can be speech information obtained in various scenarios, for example, a segment of speech including at least two languages in a conference exchange. The languages can be English, Japanese, etc. different national languages, or can be Cantonese, Shanghai dialect, etc. dialects with certain differences in different regions of the same country.

[0047] S120: Identify the speech information based on a preset speech recognition model to determine a speech recognition result.

[0048] The preset speech recognition model includes at least one independent target decoding module, which is obtained based on at least a training set including error index identification.

[0049] Exemplarily, the preset speech recognition model can use the encoding module Encoder and the decoding module Decoder structure of a typical neural network model Transformer model. The Encoder adopts a convolutional neural network for time domain downsampling to reduce the computational complexity. The Transformer structure improves the parallel computing capability of the preset speech recognition model and alleviates the information loss problem in the sequential computing process. Based on modeling and training of different tasks, the Decoder after the Encoder can be connected in parallel in multiple modeling modes, and the Transformer model structure is also adopted.

[0050] Exemplarily, the error index identifier can be set based on the identified error type. In the multiple prediction tasks of the preset speech recognition model, at least one independent target decoding module includes the prediction task of the error index identifier, and after a large amount of training, the speech recognition accuracy of the speech recognition model can be significantly improved. It can be understood that the preset speech recognition model also includes at least one basic decoding module, which is respectively used to perform other prediction tasks except the prediction task of the target decoding module. For example, the basic decoding module can include a main task decoding module or an alignment decoding module.

[0051] Exemplarily, the decoding module can be used as the main task of the preset speech recognition model through the Attention algorithm mechanism, and the error index prediction task is assisted to improve the decoding performance of the preset speech recognition model. It should be noted that the algorithm mechanism is not limited here, which can be multiple Attention algorithms under the Attention algorithm mechanism, or other algorithms with similar functions. Moreover, after obtaining the decoding results of each decoding module Decoder, comprehensive analysis can be performed on the decoding results of each decoding module Decoder to determine the final speech recognition result.

[0052] In summary, the speech recognition method provided by the embodiments of the present application comprises the following steps: obtaining speech information, wherein the speech information comprises speech in at least two languages; identifying the speech information based on a preset speech recognition model to determine a speech recognition result, wherein the preset speech recognition model comprises at least one independent target decoding module, and the target decoding module is obtained based on at least a training set comprising error index identifiers. The decoding module has a data-oriented characteristic due to the performance of the speech recognition system, that is, the more complete the training data and the larger the data volume, the better the performance of the recognition system. In the case of scarce training data, in order to ensure the accuracy of speech recognition by the speech recognition model, the error types are summarized based on the recognized errors, and the error index obtained by summarization is introduced into the decoding module of the speech recognition model, so as to add a prediction task of the error index sequence. After training of the training set comprising the error index identifiers, the speech recognition accuracy of the speech recognition model can be significantly improved. In particular, the misrecognition problem caused by different degrees of difficulty in recognizing the same word in different mixed language corpus sequences is solved. Moreover, under the premise of ensuring the training cost, data labeling cost and limited training data, various recognition errors caused by the scarcity of mixed language corpus are solved, and the recognition performance of the overall speech recognition model is improved.

[0053] In some examples, the number of error index identifiers is determined based on the number of language types. It can be understood that, since the error types in speech recognition between two languages are basically certain, the more the number of languages in the mixed language, the more the error types. Therefore, the number of error index identifiers can be determined according to the number of language types. For example, the language types are language 1 and language 2, and the errors in speech recognition include the error of misrecognizing language 1 as language 2 and the error of misrecognizing language 2 as language 1. If language 3 is mixed, the error types will at least increase the error types of misrecognizing language 3 as language 1 or language 2, and the error types of misrecognizing language 1 or language 2 as language 3.

[0054] In some examples, the error types of the error index identifiers can be obtained based on the comparison between the initial recognition result of the initial speech recognition model and the training data. For example, the initial speech recognition model can be trained in the first stage, the recognition errors existing in the initial speech recognition model in the first stage are analyzed, the error index identifier is defined according to the summarized recognition error types, the prediction task of the error index sequence is added, and the recognition training of the speech recognition model in the second stage is performed. Under the premise of ensuring the training cost, data labeling cost and limited training data, various recognition errors caused by the scarcity of mixed language corpus are solved, and the recognition performance of the overall speech recognition model is improved. Of course, the error index identifier can also be obtained by other ways or experience, which is not limited herein.

[0055] In some examples, the error types identified by the above error index may include same-language recognition errors in the target language, cross-language recognition errors in the target language, insertion errors in the target language, and deletion errors in the target language. Taking the possible error types in a Chinese-English mixed language as an example, the above same-language recognition errors may include Chinese being misrecognized as Chinese and English being misrecognized as English. The above cross-language recognition errors in the target language may include Chinese being misrecognized as English and English being misrecognized as Chinese. The above insertion errors in the target language may include Chinese insertion errors and English insertion errors. The above deletion errors in the target language may include English deletion errors and Chinese deletion errors.

[0056] In some examples, the initial speech recognition model can be trained in the first stage. After aligning the decoding result with the original annotation, the recognition errors existing in the initial speech recognition model in the first stage can be analyzed. For example, Chinese being misrecognized as Chinese means that the corresponding position in the original annotation is Chinese, and the corresponding position in the decoding result is also Chinese, but it is different from the original text. For example, if the original annotation is "你好" (Hello) and the decoding result is "您好" (How do you do), then the character "你" has the error of Chinese being misrecognized as Chinese. Chinese being misrecognized as English means that the corresponding position in the original annotation is Chinese, and the corresponding position in the decoding result becomes English. For example, if the original annotation is "你好" (Hello) and the decoding result is "你hi", then the character "好" has the error of Chinese being misrecognized as English. Chinese deletion error means that the corresponding position in the original annotation is Chinese, and there is a Chinese character missing in the corresponding position in the decoding result. For example, if the original annotation is "你好" (Hello) and the decoding result is "你***", where "***" is a placeholder, indicating the missing Chinese character at this position when the original annotation and the decoding result are best aligned. Then, the character "好" has a Chinese deletion error. Chinese insertion error means that there is a corresponding character in the decoding result at the position that is vacant in the original annotation. For example, if the original annotation is "你***好" (Hello) and the decoding result is "你真好" (You are so nice), where "***" is a placeholder, then it can be considered that both "你" and "好" have insertion errors. The English errors can be known in the same way and will not be elaborated here.

[0057] In some examples, still taking the speech recognition model of a Chinese-English mixed language as an example, when modeling, considering that there are approximately 6,000 to 7,000 common Chinese characters and about 20,000 common English words, therefore, considering the generalization of the model and ensuring that the model can handle out-of-vocabulary words (words that do not appear during training) as much as possible, Chinese can be modeled and trained using individual characters (char), and English can be modeled and trained using subword units to jointly form the result vocabulary of the mixed language. This vocabulary corresponds to the decoding output part of the speech recognition model, that is, the output layer of the neural network model.

[0058] In some examples, the language id language recognition identifier of the training audio can be generated based on the determined vocabulary and the text of the training audio. The text of the training audio can be split into units in the vocabulary according to the vocabulary. According to whether the split result belongs to Chinese or English specifically, a language id sequence is generated. For example, if the Chinese language id is set to 0 and the English language id is set to 1, and the original text is "hi早上好,你今天吃breakfast了吗", when single characters (char) are used for modeling and training in Chinese and subword units are used for modeling and training in English, the split result is "hi早上好你今天吃break fast了吗", and the generated language id is [1,0,0,0,0,0,0,0,1,1,1,0,0].

[0059] In some examples, if single characters (char) are used for modeling and training in both Chinese and English, the split result becomes "hi早上好你今天吃bre ak fast了吗", and the generated language id is [1,0,0,0,0,0,0,0,1,1,1,1,1,1,1,1,1,0,0]. In some examples, if single characters (char) are used for modeling and training in Chinese and words (word) are used for modeling and training in English, the split result becomes "hi早上好你今天吃breakfast了吗", and the generated language id is [1,0,0,0,0,0,0,0,1,0,0]. The granularity of the units in the vocabulary can be selected according to the actual situation of the language, such as the number of splits after language splitting, and no limitation is made here.

[0060] In some examples, the above initial speech recognition model may include an encoding module and independent multitask decoding modules that reuse the same encoding module Encoder, which may include: a language id language recognition identifier decoding module, an Attention decoding module, and a CTC (Connectionist Temporal Classification) alignment task decoding module. Among them, the Attention decoding module processes the main task, and the CTC alignment task decoding module and the language id language recognition identifier decoding module process auxiliary tasks to assist in improving the decoding performance of the Attention decoding module.

[0061] In some examples, for the training and optimization of the initial speech recognition model described above, it is noted that the mechanism of the multi-task is through the weighted sum of the multi-task loss function (loss), that is, through the weighted sum of the loss functions of the multiple Decoder decoding modules as the final loss function, and the weights are updated through the gradient backpropagation of the Encoder, so as to continuously optimize the initial speech recognition model described above. The loss function is used to evaluate the degree of difference between the predicted value and the true value of the model, and the better the loss function, the better the performance of the model. Taking the initial speech recognition model with a language id language recognition and identification decoding module, an Attention decoding module, and a CTC alignment task decoding module as an example, the formula of the final loss function Loss1 can be expressed as:

[0062] Loss1=W decoder *Loss decoder +W ctc *Loss ctc +W languageid

[0063] *Loss languageid

[0064] Wherein, Loss decoder is the loss function of the Attention decoding module, W aecoder is the weight of the loss function of the Attention decoding module, Loss ctc is the loss function of the CTC alignment task decoding module, W ctc is the weight of the loss function of the CTC alignment task decoding module, Loss languageid is the loss function of the language id language recognition and identification decoding module, and W language id is the weight of the loss function of the language id language recognition and identification decoding module. For example, the weight values of the various decoding modules can be obtained by multiple experiments on a small data set, and the weight values of the various decoding modules corresponding to the optimal performance.

[0065] In some examples, as shown in Figure 2 , the preset speech recognition model can include an encoding module, and multiple independent multi-task decoding modules that reuse the same encoding module Encoder can include a language id language recognition and identification decoding module, an Attention decoding module, a CTC alignment task decoding module, and an independent error index identification decoding module. Among them, the Attention decoding module processes the main task, and the CTC alignment task decoding module, the language id language recognition and identification decoding module, and the error index identification decoding module process the auxiliary task, which further assists in improving the decoding performance of the Attention decoding module.

[0066] Exemplary, with the error types that may exist in the mixed language of Chinese and English as an example, the error index identification in the identification can use 0 to represent that the text is correct, and use 1-8 to represent the error types that may exist in the mixed language of Chinese and English: Chinese error into Chinese, English error into English, Chinese error into English, English error into Chinese, Chinese insertion error, English insertion error, English deletion error and Chinese deletion error. Based on this, the result sequence of the error index identification decoding module of a certain training audio can be represented as:

[0067] S error id ={s1,s2,...s i …,s n}s i ∈(0~8)

[0068] In some examples, for the training and optimization of the above-mentioned preset speech recognition model, it needs to be noted that the action mechanism of the multi-task is the weighted sum of the multi-task loss function, that is, the weighted sum of the loss functions of the multiple Decoder decoding modules is used as the final loss function, and the Encoder is updated by gradient back propagation to update the weight, so as to continuously optimize the above-mentioned preset speech recognition model. The loss function is used to evaluate the degree of difference between the predicted value and the true value of the model, and the better the loss function, the better the performance of the model. Taking the preset speech recognition model with language id language identification decoding module, Attention decoding module, CTC alignment task decoding module and independent error index identification decoding module as an example, the formula of the final loss function Loss2 can be represented as:

[0069] Loss2=W decoder *Loss decoder +W ctc *Loss ctc +Wl anguage id

[0070] *Loss language id +W error id *Loss error id

[0071] Wherein, Loss decoder is the loss function of the Attention decoding module, and W decoder is the weight of the loss function of the Attention decoding module. Loss ctc is the loss function of the CTC alignment task decoding module, and W ctc is the weight of the loss function of the CTC alignment task decoding module. Loss language id is the loss function of the language id language identification decoding module, and Wlanguage id The weight of the loss function of the language id language identification and decoding module. error id The loss function of the independent error index identification and decoding module is W error id The weight of the loss function of the independent error index identification and decoding module.

[0072] Exemplarily, the preset speech recognition model further includes at least one basic decoding module, and the method can further include:

[0073] The weight of each decoding module is determined based on the training result of the preset speech recognition model based on the preset data training set;

[0074] The decoding result of each decoding module is obtained respectively to determine the speech recognition result, including:

[0075] The decoding result of each decoding module is obtained respectively, and the speech recognition result is determined based on the obtained decoding result of each decoding module and the corresponding weight.

[0076] The preset data training set can be a small training set, and since the error index identification and decoding module is completely independent of other decoding modules at this time, the weight values of the decoding modules can be the weight values of the decoding modules corresponding to the optimal performance through multiple experiments on a small data set. Each decoding module can include the target decoding module and the basic decoding module, and the basic decoding module can include a main task decoding module or an alignment decoding module, which can be used to perform other prediction tasks other than the prediction task of the target decoding module.

[0077] The error index identification itself also has a certain degree of language identification function.

[0078] In some examples, the target decoding module is obtained based on at least a training set including error index identification and language identification. The error index identification can be integrated into the language id decoding module. For example, in a speech recognition model for mixed Chinese and English, 0 and 1 can be used to represent correct Chinese and correct English, and 2-9 can be used to represent possible error types of mixed Chinese and English. Based on this, the result sequence of the combined language identification and error index identification decoding module of a certain training audio can be represented as:

[0079] S error id = {s1, s2,... s i …, s n} s i ∈ (0-9)

[0080] Still taking the mixed Chinese and English as an example, set 0 represents the correct Chinese text, 1 represents the correct English text, 2 represents the Chinese error into Chinese, 3 represents the English error into English, 4 represents the Chinese error into English, 5 represents the English error into Chinese, 6 represents the Chinese insertion error, 7 represents the English insertion error, 8 represents the English deletion error, and 9 represents the Chinese deletion error. Taking the original label "nǐmen hǎo" as an example, "nǐmen hi", then "nǐ" has a Chinese error into Chinese, which is marked as 2. "men" is correct, which is correct Chinese, marked as 0. And "hǎo" has a Chinese error into English, marked as 4. Therefore, the generated S error id can be represented as [2, 0, 4].

[0081] In some examples, as shown in Figure 3 , the above-mentioned preset speech recognition model can include an encoding module, and multiple independent multi-task decoding modules of the same encoding module Encoder can include: a fused language id language recognition identifier and error index identifier decoding module, an Attention decoding module and a CTC alignment task decoding module. Among them, the Attention decoding module processes the main task, and the CTC alignment task decoding module, combined with the language id language recognition identifier and error index identifier decoding module, processes the auxiliary task, further assisting to improve the decoding performance of the Attention decoding module.

[0082] In some examples, taking the preset speech recognition model with the language id language recognition identifier and error index identifier decoding module, the Attention decoding module and the CTC alignment task decoding module as an example, the formula of the final loss function Loss3 can be represented as:

[0083] Loss3=W decoder *Loss decoder +W ctc *Loss ctc +W language id+error id

[0084] *Loss language id+error id

[0085] Among them, Loss decoder is the loss function of the Attention decoding module, W decoder is the weight of the Attention decoding module loss function. Loss ctc is the loss function of the CTC alignment task decoding module, W ctc is the weight of the CTC alignment task decoding module loss function. Loss language id is the loss function of the language id language recognition identifier decoding module, Wlanguage id is a weight of a loss function of a language id language recognition identification decoding module. language id+error id is a loss function of a language id language recognition identification decoding module combined with an error index identification. language id+error id is a weight of a loss function of a language id language recognition identification decoding module combined with an error index identification.

[0086] Exemplarily, since the error index identification is obtained based on the language recognition identification decoding module at this time, the weight of each decoding module can be reused to adjust the optimal result before the adjustment, for example, the weight of each decoding module in the initial speech recognition model can be used, and the weight of each decoding module in the first stage can be fine-tuned on the basis of the weight of each decoding module in the first stage in the second stage combined with the error index identification, so as to save the training time and accelerate the convergence of the model.

[0087] In some examples, the training set includes a single language data set and a mixed language data set, and the single language data set is less than or equal to ten orders of magnitude of the mixed language data set. It should be noted that taking a speech recognition model of a Chinese-English mixed language as an example, the training set can include a pure Chinese data set, a pure English data set, and a Chinese-English mixed data set. The pure Chinese and pure English data sets are relatively easy to obtain in a large amount, but in order to avoid affecting the recognition performance of the Chinese-English mixed data set, the duration of the pure Chinese and pure English data sets should not exceed ten orders of magnitude of the Chinese-English mixed data set on the premise of covering as many common words and Chinese word groups as possible. For example, the duration of the pure Chinese and pure English data sets can be 10 times the duration of the Chinese-English mixed data set. The duration of the pure Chinese and pure English data sets can also be 5 times the duration of the Chinese-English mixed data set. The duration of the pure Chinese and pure English data sets can also be 40 times the duration of the Chinese-English mixed data set.

[0088] In some examples, in order to ensure that the obtained training set audio data covers as many texts as possible, for the acquisition of the training set audio data, the method can further include:

[0089] Selecting data with a word frequency less than or equal to an initial word frequency threshold from the single language corpus data to add to the single language data set;

[0090] If the single language data set is still less than the preset multiple of the mixed language data set after traversing the single language corpus data, increasing the initial word frequency threshold to continue traversing the single language corpus data to increase the single language data set.

[0091] For example, to obtain Chinese training data, a Chinese character table can be obtained by counting all text sequences, and the size is w. If the length of the Chinese-English mixed corpus is h, the upper limit of the time of the added Chinese data is 10h. All Chinese data can be shuffled, and n pieces of data can be randomly extracted, so that the sum of their lengths is greater than the length of the Chinese-English mixed corpus. Count the word frequency of all Chinese characters covered by the n pieces of data, and the word frequency of the characters not covered is 0, and the average value f is solved n , the initial word frequency threshold is T0=f n . From the remaining Chinese data, traverse the data to check if the word frequency of each character exceeds the threshold T. If none of them exceeds the threshold, add it to the Chinese training set and update the word frequency of each character and the length of the selected Chinese data h t . If k t >=10h, stop searching. If there is a word frequency exceeding T, put it back and extract the next piece of data. If all the remaining Chinese data are traversed and no Chinese data meeting the conditions are found, and h t <10h, update T=T0+f n , until the length of the collected Chinese training data meets 10h. The way to obtain English training data is the same, and will not be repeated here. By using the above training set acquisition method, it can be possible to consider avoiding affecting the recognition performance of the Chinese-English mixed data set, and to achieve the effect of covering all common words or Chinese word groups as much as possible to improve the accuracy of the speech recognition model.

[0092] According to some embodiments, as an implementation of the method shown in the above Figure 1 , the embodiments of the present application also provide a speech recognition device for implementing the method shown in the above Figure 1 . The device embodiments correspond to the foregoing method embodiments, and for the sake of readability, the details of the foregoing method embodiments will not be described one by one, but it should be clear that the device in the present embodiment can correspondingly implement all the contents in the foregoing method embodiments. As shown in the above Figure 4 , the device includes an acquisition unit 310, an identification unit 320, and a determination unit 330, wherein

[0093] The acquisition unit 310 can be used to acquire speech information, and the speech information includes speech in at least two languages;

[0094] The identification unit 320 can be used to identify the speech information based on a preset speech recognition model, wherein the preset speech recognition model includes at least one independent target decoding module, and the target decoding module is obtained based on at least one training set including an error index identifier;

[0095] The determining unit 330 can be configured to obtain the decoding result of each decoding module respectively to determine the speech recognition result.

[0096] According to the speech recognition device provided by the embodiment of the present application, the speech information including at least two kinds of speech is obtained, the speech information is recognized based on the preset speech recognition model, the preset speech recognition model includes at least one independent target decoding module, the target decoding module is obtained based on at least one training set including error index identification, the decoding result of each decoding module is obtained respectively to determine the speech recognition result. Because the speech recognition system has the data-oriented characteristics, the more complete the training data is and the larger the data volume is, the better the performance of the recognition system is. In order to ensure the accuracy of the speech recognition model in the case of insufficient training data, the error types are summarized based on the recognized errors, and the error index is introduced into the decoding module of the speech recognition model, so that the prediction task of the error index sequence is added, and the speech recognition accuracy of the speech recognition model is improved significantly through the training of the training set including the error index identification. Especially for the misrecognition problem caused by the different difficulty levels of the same word in different mixed language corpus sequences. And under the premise of ensuring the training cost, the data labeling cost and the limited training data, the various recognition errors caused by the scarcity of mixed language corpus are solved, and the recognition performance of the overall speech recognition model is improved.

[0097] In some examples, the number of error index identifications is determined based on the number of language types.

[0098] In some examples, the error type of the error index identification is obtained by comparing the initial recognition result of the initial speech recognition model with the training data.

[0099] In some examples, the error type of the error index identification includes same language recognition error of the target language, cross language recognition error of the target language, insertion error of the target language and deletion error of the target language.

[0100] In some examples, the target decoding module is obtained based on at least one training set including error index identification and language recognition identification.

[0101] In some examples, the determining unit can be further configured to:

[0102] determine the weight of each decoding module based on the training result of the preset data training set in the preset speech recognition model;

[0103] The decoding result of each decoding module is obtained respectively, and the speech recognition result is determined based on the obtained decoding result of each decoding module and the corresponding weight.

[0104] In some examples, the training set includes a single language data set and a mixed language data set, and the single language data set is less than or equal to ten times the mixed language data set.

[0105] In some examples, the obtaining unit is further configured to:

[0106] In the single language corpus data, data with a word frequency less than or equal to an initial word frequency threshold is selected to join the single language data set.

[0107] If the single language data set is still less than the preset multiple of the mixed language data set after the single language corpus data is traversed, the initial word frequency threshold is increased to continue traversing the single language corpus data, so as to increase the single language data set.

[0108] It should be noted that the processor includes a core, and the core retrieves the corresponding program unit from the memory. The core can be set to one or more, and the accuracy of the speech recognition in the current speech recognition process can be improved by adjusting the core parameters.

[0109] The embodiment of the present application provides a storage medium, which stores a program, and the program is executed by a processor to realize the speech recognition method.

[0110] The embodiment of the present application provides a processor, which is used for running a program, and the program is executed to perform the speech recognition method.

[0111] The embodiment of the present application provides a device 400, as shown in the figure, the device includes at least one processor 410 and at least one memory 420 connected with the processor; wherein the processor 410, the memory 420 communicates with each other; the processor 410 is used for calling the program instruction in the memory 420, to execute the speech recognition method described above. Figure 5

[0112] The device herein can be a server, a PC, a PAD, a mobile phone, etc.

[0113] ​The application further provides a computer program product suitable for executing a program for initializing the following method steps when executed on a process management device: obtaining voice information, the voice information comprising voice in at least two languages; identifying the voice information based on a preset voice recognition model, wherein the preset voice recognition model comprises at least one independent target decoding module, and the target decoding module is obtained based on at least a training set comprising error index identification; and obtaining the decoding result of each decoding module to determine a voice recognition result.

[0114] In some examples, the number of error index identifications is determined based on the number of language categories.

[0115] In some examples, the error type of the error index identification is obtained by comparing the initial recognition result of the initial voice recognition model with the training data.

[0116] In some examples, the error type of the error index identification comprises same language recognition error of the target language, cross language recognition error of the target language, insertion error of the target language, and deletion error of the target language.

[0117] In some examples, the target decoding module is obtained based on at least a training set comprising error index identification and language recognition identification.

[0118] In some examples, the method further comprises:

[0119] training the weight of each decoding module in the preset voice recognition model based on a preset data training set;

[0120] The obtaining of the decoding result of each decoding module to determine the voice recognition result comprises:

[0121] The obtaining of the decoding result of each decoding module comprises: obtaining the decoding result of each decoding module based on the obtained decoding result of each decoding module and the corresponding weight to determine the voice recognition result.

[0122] In some examples, the training set comprises a single language data set and a mixed language data set, and the single language data set is less than or equal to ten orders of magnitude of the mixed language data set.

[0123] In some examples, the method further comprises:

[0124] In the single language corpus data, data with a word frequency less than or equal to an initial word frequency threshold is selected to be added to the single language data set.

[0125] If the single language data set is still less than the preset multiple of the mixed language data set after traversing the single language corpus data, the initialization word frequency threshold is increased to continue traversing the single language corpus data, so as to increase the single language data set.

[0126] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions specified in the flowchart block or blocks. Figure 1 The flowchart and / or block diagram in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to the present application. In this regard, each flowchart block and / or block in the figures can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that the flowchart blocks and / or blocks in the figures can represent a phase of a possible implementation. The flowchart blocks and / or blocks in the figures can also represent a part of a possible implementation. Figure 1 The flowchart and / or block diagram in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to the present application. In this regard, each flowchart block and / or block in the figures can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that the flowchart blocks and / or blocks in the figures can represent a phase of a possible implementation. The flowchart blocks and / or blocks in the figures can also represent a part of a possible implementation.

[0127] In one typical arrangement, the device includes one or more processors (CPU), memory, and a bus. The device can also include an input / output interface, a network interface, and the like.

[0128] The memory can include non-persistent memory and / or volatile memory, e.g., random access memory (RAM) such as static random access memory (SRAM), dynamic random access memory (DRAM), or other random access memories: non-volatile memory, e.g., read only memory (ROM), flash memory, or other non-volatile storage; and / or a combination of different memory types. The memory includes at least one memory chip. The memory is an example of computer readable media.

[0129] Computer readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definition herein, computer readable media does not include transitory media, such as modulated data signals and carrier waves.

[0130] It is also to be noted that the terms "comprising", "including", and any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0131] Those skilled in the art will appreciate that embodiments of the present application can be devised for a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer-readable program code.

[0132] The embodiments of the present application are only illustrative and are not intended to limit the present application. Various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A speech recognition method, characterized in that, include: Acquire voice information, wherein the voice information includes voices in at least two languages; The speech information is recognized based on a preset speech recognition model to determine the speech recognition result. The preset speech recognition model includes at least one independent target decoding module. The target decoding module is a decoding module trained on a training set including error index identifiers. The error types of the error index identifiers include same-language recognition errors, cross-language recognition errors, insertion errors, and deletion errors in the target language. The number of error index identifiers is determined based on the number of language types.

2. The method according to claim 1, characterized in that, The error type identified by the error index is obtained by comparing the initial recognition results of the initial speech recognition model with the training data.

3. The method according to claim 1, characterized in that, The target decoding module is obtained by training on at least a training set including error index identifiers, and includes: The target decoding module is trained on a training set that includes error index identifiers and language identification identifiers.

4. The method according to claim 1, characterized in that, The preset speech recognition model further includes at least one basic decoding module, and the method further includes: The weight of each decoding module is determined based on the training results in the preset speech recognition model using a preset data training set. The decoding module obtains the decoding result of each decoding module, and determines the speech recognition result based on the obtained decoding result of each decoding module and the corresponding weight.

5. The method according to any one of claims 1-4, characterized in that, The training set includes a single-language dataset and a mixed-language dataset, wherein the single-language dataset is less than or equal to ten times the size of the mixed-language dataset.

6. The method according to claim 5, characterized in that, The method further includes: Data with word frequencies less than or equal to the initial word frequency threshold are selected from the single-language corpus data and added to the single-language dataset; If, after traversing the single-language corpus, the single-language dataset is still less than a preset multiple of the mixed-language dataset, then the initial word frequency threshold is increased to continue traversing the single-language corpus, thereby increasing the single-language dataset.

7. A voice recognition device, characterized in that, include: An acquisition unit is used to acquire voice information, the voice information including voices in at least two languages; The recognition unit is used to recognize the speech information based on a preset speech recognition model to determine the speech recognition result. The preset speech recognition model includes at least one independent target decoding module. The target decoding module is a decoding module trained on at least a training set including error index identifiers. The error types of the error index identifiers include same-language recognition errors, cross-language recognition errors, insertion errors, and deletion errors in the target language. The number of error index identifiers is determined based on the number of language types.

8. The apparatus according to claim 7, characterized in that, The error type identified by the error index is obtained by comparing the initial recognition results of the initial speech recognition model with the training data.

9. The apparatus according to claim 7, characterized in that, The target decoding module is trained on a training set that includes error index identifiers and language identification identifiers.

10. The apparatus according to claim 7, characterized in that, The preset speech recognition model further includes at least one basic decoding module, and also includes a determination unit, used for: The weight of each decoding module is determined based on the training results in the preset speech recognition model using a preset data training set. The decoding result of each decoding module is obtained separately, and the speech recognition result is determined based on the obtained decoding result of each decoding module and the corresponding weight.

11. The apparatus according to any one of claims 7-10, characterized in that, The training set includes a single-language dataset and a mixed-language dataset, wherein the single-language dataset is less than or equal to ten times the size of the mixed-language dataset.

12. The apparatus according to claim 11, characterized in that, The acquisition unit is also used for: Data with word frequencies less than or equal to the initial word frequency threshold are selected from the single-language corpus data and added to the single-language dataset; If, after traversing the single-language corpus, the single-language dataset is still less than a preset multiple of the mixed-language dataset, then the initial word frequency threshold is increased to continue traversing the single-language corpus, thereby increasing the single-language dataset.

13. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, it controls the device where the storage medium is located to perform the speech recognition method as described in any one of claims 1 to 6.

14. An electronic device, characterized in that, The device includes at least one processor and at least one memory connected to the processor; the processor is used to call program instructions in the memory to execute the speech recognition method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Efficient self-adaptive speech recognition engine-oriented hot word error correction method and system

    CN118471201A

  • Speech recognition system and method for automatically calibrating data label

    WO2022030805A1