Speech recognition methods, devices, storage media and equipment

By using a multi-granularity speech recognition model, which employs a combined training set granularity of multiple independent decoding parts, the problem of poor accuracy in mixed-language speech recognition models is solved, thereby improving the accuracy and performance of mixed-language speech recognition.

CN115810347BActive Publication Date: 2025-11-14BEIJING YUANLI WEILAI SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111070846.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-13
Publication Date
2025-11-14
Estimated Expiration
2041-09-13

AI Technical Summary

Technical Problem

The recognition performance of mixed-language speech recognition models is worse than that of single-language models, making it difficult to meet the needs of mixed-language communication scenarios, mainly due to the scarcity of training data.

Method used

A multi-granularity speech recognition model is adopted, which includes at least two independent decoding parts. Each decoding part corresponds to a different training set granularity. The speech recognition result is determined by obtaining the decoding result of each decoding part separately.

Benefits of technology

It improves the accuracy and performance of mixed-language speech recognition, alleviates the dilemma of scarce training data, enhances attention to multiple pronunciation combinations, and improves the accuracy of the recognition model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115810347B_ABST
    Figure CN115810347B_ABST
Patent Text Reader

Abstract

This application discloses a speech recognition method, apparatus, storage medium, and device, primarily aimed at addressing the problem of poor accuracy in current speech recognition processes. The method includes: acquiring speech information, which includes speech in at least two languages; recognizing the speech information based on a multi-granularity speech recognition model, wherein the multi-granularity speech recognition model includes at least two independent decoding parts, and the granularity of the training set corresponding to each decoding part is different; and acquiring the decoding result of each decoding part to determine the speech recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and in particular to a speech recognition method, apparatus, storage medium, and device. Background Technology

[0002] As globalization intensifies, multilingual communication is increasingly prevalent in politics, economics, sports, and culture, ranging from official negotiations to everyday conversations. Speech recognition technology, by converting audio signals into more intuitive text records, has experienced rapid development in recent years, spawning numerous applications such as video subtitling and automated meeting minutes, greatly facilitating people's lives. Simultaneously, this has further increased the demands on the performance of multilingual speech recognition systems.

[0003] However, due to the high flexibility of mixed-language communication and the high cost of transcribing mixed-language audio, training data is scarce. Therefore, the recognition performance of mixed-language speech recognition models is often worse than that of single-language speech recognition models, making it difficult to meet the needs of existing mixed-language communication scenarios. Summary of the Invention

[0004] In view of the above problems, this application provides a speech recognition method, device, storage medium and equipment, the main purpose of which is to solve the problem of poor accuracy of speech recognition in the current speech recognition process.

[0005] To address the aforementioned technical problems, in a first aspect, embodiments of this application provide a speech recognition method, which may include:

[0006] Acquire voice information, which includes voice information in at least two languages;

[0007] The above speech information is recognized based on a multi-granularity speech recognition model, wherein the multi-granularity speech recognition model includes at least two independent decoding parts, and the combination of granularity of the training set corresponding to each decoding part is different.

[0008] The decoding results of each of the above decoding parts are obtained separately to determine the speech recognition result.

[0009] In the first possible implementation of the first aspect above, the combination of languages ​​and granularity in each of the above training sets is different.

[0010] In a second possible implementation of the first aspect above, the training granularity includes a single character, a word segmentation unit, and a word, wherein the word segmentation unit is composed of the single character, and the word is composed of the word segmentation unit.

[0011] In a third possible implementation of the first aspect above, obtaining the decoding result of each of the decoding parts to determine the speech recognition result may include:

[0012] Obtain the decoding result for each of the above decoding parts;

[0013] Calculate the confidence score of the decoding result for each of the above decoding parts;

[0014] The decoding result with the highest confidence score is determined as the speech recognition result.

[0015] In the fourth possible implementation of the first aspect above, before determining the decoding result with the highest confidence score as the speech recognition result, it may further include:

[0016] The obtained decoding results are compared. If there are identical decoding results, the posterior probabilities of the identical decoding results are combined to calculate the confidence score of the identical decoding results.

[0017] In the fifth possible implementation of the first aspect above, before determining the decoding result with the highest confidence score as the speech recognition result, the method may further include:

[0018] Calculate the size of the resulting vocabulary associated with each of the above decoding parts;

[0019] Based on the size of each of the above-mentioned result vocabularies obtained through calculation, update the confidence score of the decoding result of the above-mentioned decoding part associated with the above-mentioned result vocabulary.

[0020] In the sixth possible implementation of the first aspect above, the resulting vocabulary of the speech recognition model includes natural mixed Chinese and English text outside the training set and / or text obtained after processing a single language text within the training set.

[0021] In the seventh possible implementation of the first aspect above, the processing method of processing the single-language text in the training set includes replacing the names and / or pronouns in the single-language text with other languages.

[0022] Secondly, embodiments of this application also provide a speech recognition device, which may include: an acquisition unit, a recognition unit, and a determination unit.

[0023] The aforementioned acquisition unit can be used to acquire speech information, which includes speech in at least two languages;

[0024] The aforementioned recognition unit can be used to recognize the aforementioned speech information based on a multi-granularity speech recognition model, wherein the aforementioned multi-granularity speech recognition model includes at least two mutually independent decoding parts, and the granularity of the training set corresponding to each of the aforementioned decoding parts is different;

[0025] The aforementioned determining unit can be used to obtain the decoding result of each of the aforementioned decoding parts to determine the speech recognition result.

[0026] In the first possible implementation of the second aspect above, the combination of languages ​​and granularity in each of the above training sets is different.

[0027] In a second possible implementation of the second aspect above, the training granularity may include a single character, a word segmentation unit, and a word, wherein the word segmentation unit is composed of the single character and the word is composed of the word segmentation unit.

[0028] In a third possible implementation of the second aspect described above, the determining unit can specifically be used for:

[0029] Obtain the decoding result for each of the above decoding parts;

[0030] Calculate the confidence score of the decoding result for each of the above decoding parts;

[0031] The decoding result with the highest confidence score is determined as the speech recognition result.

[0032] In a fourth possible implementation of the second aspect described above, the determining unit may also be used for:

[0033] Compare the obtained decoding results. If there are identical decoding results, combine the posterior probabilities of the identical decoding results to calculate the confidence score of the identical decoding results.

[0034] In a fifth possible implementation of the second aspect described above, the determining unit may also be used for:

[0035] Calculate the size of the resulting vocabulary associated with each of the above decoding parts;

[0036] Based on the size of each of the above-mentioned result vocabularies obtained through calculation, update the confidence score of the decoding result of the above-mentioned decoding part associated with the above-mentioned result vocabulary.

[0037] In the sixth possible implementation of the second aspect above, the resulting vocabulary of the speech recognition model may include natural mixed Chinese and English text outside the training set and / or text obtained after processing a single language text within the training set.

[0038] In the seventh possible implementation of the second aspect above, the processing method for processing the single-language text in the training set may include replacing the names and / or pronouns in the single-language text with other languages.

[0039] To achieve the above objectives, according to a third aspect of the embodiments of this application, a storage medium is provided, which may include a stored program, wherein, when the program is running, the device where the storage medium is located controls the execution of the speech recognition method described in any one of the first aspects.

[0040] To achieve the above objectives, according to a fourth aspect of the embodiments of this application, an electronic device is provided, the device including at least one processor and at least one memory connected to the processor; the processor is configured to call program instructions in the memory to execute the speech recognition method described in any one of the first aspects.

[0041] By employing the above technical solution, the speech recognition method provided in this application acquires speech information, which includes speech in at least two languages. It then recognizes the speech information based on a multi-granularity speech recognition model, wherein the multi-granularity speech recognition model includes at least two independent decoding parts, each of which corresponds to a different combination of granularity in the training set. The decoding result of each decoding part is acquired to determine the speech recognition result. Due to the data-driven nature of speech recognition systems, the more complete and larger the training data, the better the performance of the recognition system. Because the different combinations of granularity in the training set corresponding to the decoding parts allow for multiple different label sequences to be obtained from the same mixed-language audio, the overall usability of the mixed-language training corpus increases several times, alleviating the scarcity of training set data to some extent and improving the accuracy of speech recognition. Furthermore, when training the speech recognition model by superimposing label sequences of different granularities onto the same mixed-language audio, the speech recognition model considers multiple pronunciation combinations, further improving the accuracy of speech recognition. Furthermore, since the multi-granularity speech recognition model includes at least two independent decoding modules, the decoding parts of the models trained from training sets of different granularities are independent of each other. By obtaining the decoding results of each decoding part separately, the decoding result with the highest matching degree to the pronunciation combination of the audio to be decoded can be selected as the final decoding result, further improving the recognition performance and accuracy of the speech recognition model.

[0042] The above-mentioned voice recognition device, storage medium, and electronic device also have corresponding effects due to the adoption of the above-mentioned voice recognition method. The above description is only an overview of the technical solution of the embodiments of this application. In order to better understand the technical means of the embodiments of this application and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of this application more obvious and understandable, the following are specific implementation methods of the embodiments of this application. Attached Figure Description

[0043] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0044] Figure 1 A schematic flowchart of a speech recognition method provided in an embodiment of this application is shown;

[0045] Figure 2 A schematic framework diagram of a multi-granularity speech recognition model provided in an embodiment of this application is shown;

[0046] Figure 3 A schematic structural block diagram of a speech recognition device provided in an embodiment of this application is shown;

[0047] Figure 4 A schematic structural block diagram of an electronic device for speech recognition provided in an embodiment of this application is shown. Detailed Implementation

[0048] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the embodiments of the present application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough in its understanding and will fully convey the scope of the present application to those skilled in the art.

[0049] To address the issue of poor accuracy in current speech recognition processes, embodiments of this application provide a speech recognition method, such as... Figure 1 As shown, the method may include: S110 to S130.

[0050] S110: Obtain voice information, which includes voice information in at least two languages.

[0051] It should be noted that the aforementioned voice information can be obtained in various scenarios. For example, it could be a segment of audio from a conference conversation that includes at least two languages. These languages ​​can be from different countries, such as English and Japanese, or dialects within the same country, such as Cantonese and Shanghainese, which have certain differences between different regions.

[0052] S120: Recognize the above speech information based on a multi-granularity speech recognition model.

[0053] The aforementioned multi-granularity speech recognition model may include at least two independent decoding parts, and the combination of granularity of the training set corresponding to each decoding part is different.

[0054] For example, the aforementioned multi-granularity speech recognition model can utilize the typical Transformer neural network model's Encoder and Decoder structure. The Encoder employs a convolutional neural network for temporal downsampling, reducing computational complexity. The Transformer structure enhances the parallel computation capability of the multi-granularity speech recognition model, mitigating information loss during sequential computation. For instance, since the granularity of the training set can refer to different granularity-based segmentation methods for the same mixed-language speech, different segmentation granularities result in different label sequences for the same mixed-language speech. In other words, labeling speech data sequences based on different granularity segmentation methods yields different label sequences. Therefore, combining different granularity segmentation methods to label speech data sequences can produce different label sequences. Modeling and training based on different granularities allows for the parallel connection of Decoders with various modeling methods after the Encoder, also employing the Transformer model structure.

[0055] S130: Obtain the decoding result of each of the above decoding parts to determine the speech recognition result.

[0056] For example, an Attention algorithm can be used to obtain the highest-probability decoding path at each granularity through decoding by each decoder, serving as the candidate decoding result. It should be noted that the algorithm mechanism here is not limited; it can be various Attention algorithms or other algorithms with similar functionality. Furthermore, after obtaining the decoding results from each decoder, the results can be compared and comprehensively analyzed to determine the final speech recognition result.

[0057] In summary, the speech recognition method provided in this application acquires speech information, which includes speech in at least two languages. It then recognizes the speech information based on a multi-granularity speech recognition model, wherein the multi-granularity speech recognition model includes at least two independent decoding parts, each of which corresponds to a different combination of granularity in the training set. The decoding result of each decoding part is acquired to determine the speech recognition result. Due to the data-driven nature of speech recognition systems, the more complete and larger the training data, the better the performance of the recognition system. Because the different combinations of granularity in the training set corresponding to the decoding parts allow for multiple different label sequences to be obtained from the same mixed-language audio, the overall usability of the mixed-language training corpus increases several times, alleviating the scarcity of training set data to some extent and improving the accuracy of speech recognition. Furthermore, when training the speech recognition model by superimposing label sequences of different granularities onto the same mixed-language audio, the speech recognition model considers multiple pronunciation combinations, further improving the accuracy of speech recognition. Furthermore, since the multi-granularity speech recognition model includes at least two independent decoding parts, the decoding parts of the models trained from training sets of different granularities are independent of each other. By obtaining the decoding results of each decoding part separately, the decoding result with the highest matching degree to the pronunciation combination of the audio to be decoded can be selected as the final decoding result, further improving the recognition performance and accuracy of the speech recognition model.

[0058] In some examples, the combination of language and granularity differs in each of the aforementioned training sets. Taking a speech recognition model for mixed Chinese and English as an example, the recognition result of output Decoder1 can be obtained based on a combination of Chinese and English at the first training granularity; the recognition result of output Decoder2 can be obtained based on a combination of Chinese and English at the second training granularity; the recognition result of output Decoder3 can be obtained based on a combination of Chinese at the first training granularity and English at the second training granularity; and the recognition result of output Decoder4 can be obtained based on a combination of Chinese at the second training granularity and English at the first training granularity. Therefore, because the combination of language and granularity differs in each of the aforementioned training sets, more independent decoding parts with different granularities can be formed through the combination of language and training granularity. Due to the different splitting granularities, different label sequences can be obtained based on the same mixed language speech. That is, labeling speech data sequences based on different granularity splitting methods can yield different label sequences, solving the problem of inaccurate recognition caused by scarce training data.

[0059] In some examples, the selection of the training set does not need to satisfy all possible combinations of languages ​​and granularities. For example, assume that Chinese has two splitting granularities and English has two splitting granularities. Then, for a speech recognition model of mixed Chinese and English languages, all possible combinations of languages ​​and granularities include the combination of Chinese and English at the first training granularity, the combination of Chinese and English at the second training granularity, the combination of Chinese at the first training granularity and English at the second training granularity, and the combination of Chinese at the second training granularity and English at the first training granularity.

[0060] Understandably, some speech recognition models may only include Decoder1, Decoder2, and Decoder3. Specifically, the recognition result of the output part Decoder1 can be obtained based on a combination of Chinese and English at the first training granularity; the recognition result of the output part Decoder2 can be obtained based on a combination of Chinese and English at the second training granularity; and the recognition result of the output part Decoder3 can be obtained based on a combination of Chinese at the first training granularity and English at the second training granularity.

[0061] In some other speech recognition models, the speech recognition model may only include Decoder1 and Decoder2. The recognition result of the output part Decoder1 can be obtained based on a combination of Chinese and English at a first training granularity, while the recognition result of the output part Decoder2 can be obtained based on a combination of Chinese and English at a second training granularity.

[0062] In some speech recognition models, only Decoder2, Decoder3, and Decoder4 may be included. The recognition result of the output part Decoder2 can be obtained based on a combination of Chinese characters at the second training granularity and English characters at the second training granularity; the recognition result of the output part Decoder3 can be obtained based on a combination of Chinese characters at the first training granularity and English characters at the second training granularity; and the recognition result of the output part Decoder4 can be obtained based on a combination of Chinese characters at the second training granularity and English characters at the first training granularity. For example, the above training granularities can include char (single character), subword (subword segmentation unit), and word (word), where the subword segmentation unit is composed of the single character, and the word is composed of the subword segmentation unit. When training a neural network model, a vocabulary and modeling units need to be selected to map the text into an integer sequence that is convenient for neural network operations. Modeling and training can be based on the following granularities: single character granularity, subword segmentation unit granularity, and word granularity. Taking English as an example, if char modeling is used, only the 26 letters and a few symbols such as "," and "." need to be mapped to unique integer indices to achieve the serialization of any English text. These 26 letters and special symbols form the vocabulary. The word modeling method is more direct, mapping all English words to unique integer indices, with all words collectively constituting the vocabulary. However, the number of commonly used words is approximately 20,000 to 30,000. Due to the excessively large vocabulary and the low frequency of some words, the data is relatively sparse, making model training difficult, and this method is generally not used. Taking all factors into consideration, the subword method uses a specific algorithm to split words into affixes. For example, "subword" can be split into the affixes "sub" and "word." Some affixes can be selected to form the vocabulary. This modeling method can effectively control the vocabulary size. Common methods include BPE (byte pair encoder) and unigram unary segmentation, which will not be limited here.

[0063] Taking the Chinese and English mixed speech recognition model as an example, if we use single character (char) and subword segmentation units as the training granularity, we can use the weighted sum of the loss functions of multiple decoders as the final loss function. Then, we can perform gradient backpropagation in the encoder to update the weights, thereby continuously optimizing the model. The loss function is used to evaluate the degree of difference between the model's predicted value and the true value. The better the loss function, the better the model's performance usually is.

[0064] Taking a Chinese and English mixed speech recognition model as an example, if we use single character (char) and subword segmentation units as the training granularity, the final loss function can be expressed as:

[0065] Loss = W cc+ec *Loss cc+ec +W cc+es *Loss cc+es +W cs+ec *Loss cs+ec

[0066] ++W cs+es *Loss cs+

[0067] Among them, Loss cc+ec The loss function W for modeling an Attention Decoder for Chinese and English characters. cc+ec Weight it. Loss cc+es The loss function W for modeling an Attention Decoder for Chinese characters and English subwords. cc+es Weight it. Loss cs+ec The loss function W for modeling an Attention Decoder for Chinese subwords and English characters. cs+ec Weight it. Loss cs+es The loss function W of the Attention Decoder model for Chinese subword + English subword is used. cs+e Assign weights to them. For example, in this scheme, the weights of the four Decoders can be made the same, so the weights of each Decoder are equal and equal to 1 / 4.

[0068] For example, such as Figure 2 As shown, the aforementioned multi-granularity speech recognition model can include one encoding part and four decoding parts. Speech information is encoded in the encoding part, and this speech information can include two languages: Language 1 and Language 2. The granularity of modeling and training for the decoding parts can include single-character granularity and word segmentation unit granularity. The four decoding parts can be divided into: Decoding Part 1, single-character granularity for Language 1 and Language 2; Decoding Part 2, single-character granularity for Language 1 and word segmentation unit granularity for Language 2; Decoding Part 3, word segmentation unit granularity for Language 1 and single-character granularity for Language 2; Decoding Part 4, word segmentation unit granularity for Language 1 and word segmentation unit granularity for Language 2. Through these four decoding parts, four independent decoding results, from Decoding Result 1 to Decoding Result 4, can be obtained.

[0069] For example, since the number of decoding parts can be determined by the number of languages ​​and the training granularity, and the languages ​​include Language 1, Language 2, and Language 3, the granularity of modeling and training the decoding parts can include single character granularity, word segmentation unit granularity, and word granularity. The encoder part can be divided into a maximum of 27 decoding parts: Decoding Part 1 (single character granularity for Language 1, Language 2, and Language 3); Decoding Part 2 (single word segmentation granularity for Language 1, Language 2, and Language 3); and so on. Through these 27 decoding parts, twenty-seven independent decoding results, from Decoding Result 1 to Decoding Result 27, can be obtained. It should be noted that some of these decoding parts can also be selected as the chosen decoding parts to implement the above scheme; this is not limited here.

[0070] According to some embodiments, obtaining the decoding result of each of the decoding parts to determine the speech recognition result may include:

[0071] Obtain the decoding result for each of the above decoding parts;

[0072] Calculate the confidence score of the decoding result for each of the above decoding parts;

[0073] The decoding result with the highest confidence score is determined as the speech recognition result.

[0074] It should be noted that the above multi-granularity speech recognition model can generate multiple independent decoding results for the same audio. The most direct way to determine the speech recognition result is to use the optimal decoding method, that is, to select the decoding result with the highest confidence score from the candidate decoding results as the final output speech recognition result.

[0075] For example, when a decoder performs decoding, to improve decoding performance, it generally uses a beamsearch strategy. This involves constructing a decoding graph through breadth-first search to obtain a suboptimal solution. Setting the beam size of the beam search to n, the final result will consist of n decoding sequences, representing the n most likely decoding results for the current audio. Let's continue with... Figure 2 Taking the multi-granularity quad-decoder model architecture shown as an example, 4n decoding results will be obtained after decoding. Each decoding result corresponds to a confidence score S, which reflects the posterior probability of decoding the text given the audio, and can be expressed as:

[0076] S i =log(P(y) i |x)), i∈1~n

[0077] Where x represents the current audio, y iLet P(y) represent the i-th decoded sequence. i |x) represents the decoded sequence y corresponding to the current audio x. i The probability, S i This is the decoding score, or confidence score, of the result; a higher score corresponds to a higher posterior probability. By sorting the confidence scores S of the decoding results, the decoding result with the highest score can be selected as the final result to confirm the speech recognition result.

[0078] It should be noted that since the beam search results of different decoding parts of the decoder may contain the same decoding results, ignoring this situation may lead to miscalculation of the confidence score parameter of the decoding result, thereby reducing the accuracy of the speech recognition model in speech recognition.

[0079] In some examples, to avoid the aforementioned problems, the process of determining the decoding result with the highest confidence score as the speech recognition result may further include:

[0080] The obtained decoding results are compared. If there are identical decoding results, the posterior probabilities of the identical decoding results are combined to calculate the confidence score of the identical decoding results.

[0081] For example, we can count the number of sequence types appearing in the candidate sequences, i.e., the decoding results output by the decoding part, and finally obtain m sequences. Calculate the average confidence score for each sequence:

[0082]

[0083] Where, N j This means that N is present in all decoding results. j The same decoding result corresponds to the decoding sequence j, S k It is its corresponding confidence score.

[0084] By sorting the confidence scores of the decoding results, the decoding result with the highest score can be selected as the final result to confirm the speech recognition result. Here, considering the possibility that the beam search results of different decoding parts may contain the same decoding result, the posterior probabilities of the above-mentioned identical decoding results are combined to calculate the confidence score of the above-mentioned identical decoding results. This can effectively avoid the problem of miscalculation of the confidence score parameter of the decoding result, and further improve the accuracy of the speech recognition model in speech recognition.

[0085] It should be noted that because the above speech recognition models employ a multi-granularity modeling and training strategy, ultimately using a combination of various modeling units, the vocabulary size corresponding to each decoder differs. For example, vocabulary A has a size of L, and vocabulary B has a size of Q. When the model is not trained, from the perspective of maximum likelihood, given a speech segment, the probability of each unit in the vocabulary given by the model is equal.

[0086] That is, the probability in vocabulary A is The probability of B vocabulary is It is evident that vocabulary size itself affects the output of posterior probabilities, and this effect persists even after the model training has converged.

[0087] Ignoring this situation could lead to miscalculation of the confidence score parameter in the decoding result, thereby reducing the accuracy of the speech recognition model in speech recognition.

[0088] In some examples, to avoid the aforementioned problems, before determining the decoding result with the highest confidence score as the speech recognition result, the method may further include:

[0089] Calculate the size of the resulting vocabulary associated with each of the above decoding parts;

[0090] Based on the size of each of the above-mentioned result vocabularies obtained through calculation, update the confidence score of the decoding result of the above-mentioned decoding part associated with the above-mentioned result vocabulary.

[0091] For example, the sizes V1, V2, V3, and V4 of each result vocabulary can be obtained statistically, and the confidence score S corresponding to each decoding result can be expressed as:

[0092] S i =log(P(y) i |x)), i∈1~n

[0093] Where x represents the current audio, y i Let P(y) represent the i-th decoded sequence. i |x) represents the decoded sequence y corresponding to the current audio x. i The probability, S i This is the decoding score of the result; a higher score corresponds to a higher posterior probability. Typically, the confidence score S of the decoding results is directly sorted, and the decoding result with the highest score is selected as the final result to confirm the speech recognition result. However, considering the vocabulary sizes V1, V2, V3, and V4 of each result, the confidence score of the decoding result can be updated according to the vocabulary size corresponding to the decoder to which the decoding sequence belongs.

[0094]

[0095] Among them, S i ’ V represents the original score of the decoding result i. k This indicates that the decoding result i is the output of Decoder k, and its vocabulary size is V. k Here, by considering the different vocabulary sizes of each decoder, potential miscalculations of the confidence score parameter in the decoding results are avoided, thus improving the accuracy of the speech recognition model. Of course, we can also consider the possibility that beam search results from different decoder parts might contain the same decoding results. Therefore, we can continue to count the types of sequences appearing in the candidate sequences (i.e., the decoding results output by the decoder part), ultimately obtaining m sequences. The average confidence score for each sequence is then calculated:

[0096]

[0097] Where, N j This means that N is present in all decoding results. j The same decoding result corresponds to the decoding sequence j, S k This is the corresponding confidence score. By sorting the confidence scores of the decoding results, the decoding result with the highest score can be selected as the final result to confirm the speech recognition result. Here, since the case that the beam search results of different decoding parts of the decoder may contain the same decoding result is considered, the posterior probabilities of the above-mentioned identical decoding results are combined to calculate the confidence score of the above-mentioned identical decoding results, which further improves the accuracy of the speech recognition model in speech recognition.

[0098] It should be noted that due to the scarcity of mixed-language corpora in the original training set, the types and combinations of mixed-language sequences are limited, failing to reflect the real-world situations of various application scenarios. For example, when using subword modeling, generating a vocabulary solely based on the text content within the training set results in a vocabulary that statistically deviates from the true distribution, leading to poor recognition results from the speech recognition model. However, while annotating mixed-language audio training data is costly, mixed-language text corpora can be readily and abundantly obtained. Therefore, in some examples, the resulting vocabulary of the aforementioned speech recognition model includes natural mixed Chinese-English text outside the training set and / or text obtained by processing single-language text within the training set. The aforementioned natural mixed Chinese-English text outside the training set can be obtained from everyday data in various ways, while the method of processing single-language text within the training set can yield forged mixed-language text. This expands the resulting vocabulary and further improves the accuracy of speech recognition by the trained speech recognition model.

[0099] In some examples, the processing method for single-language text within the training set includes replacing names and / or pronouns in the single-language text with those in other languages. For instance, all English words appearing in the English and Chinese-English text within the set are counted, all English words are filtered out, retaining only English nouns or pronouns. A Chinese-English dictionary is consulted, and replaceable phrases in the Chinese text within the set with the aforementioned English nouns or pronouns, thus obtaining mixed Chinese-English text. Specifically, for example, 1 to 3 words are randomly replaced in each Chinese text. This expands the resulting vocabulary and further improves the accuracy of speech recognition by the trained speech recognition model.

[0100] Furthermore, as a response to the above Figure 1 In addition to the implementation of the method shown, this application also provides a speech recognition device for the above-described speech recognition method. Figure 1 The method shown is implemented accordingly. This device embodiment corresponds to the foregoing method embodiment. For ease of reading, this device embodiment will not repeat the details of the foregoing method embodiment, but it should be clear that the device in this embodiment can implement all the contents of the foregoing method embodiment. Figure 3 As shown, the device includes: an acquisition unit 310, an identification unit 320, and a determination unit 330, wherein,

[0101] The acquisition unit 310 can be used to acquire speech information, which includes speech in at least two languages;

[0102] The recognition unit 320 can be used to recognize the above-mentioned speech information based on a multi-granularity speech recognition model, wherein the multi-granularity speech recognition model includes at least two mutually independent decoding parts, and the combination of granularity of the training set corresponding to each of the above-mentioned decoding parts is different.

[0103] The determining unit 330 can be used to obtain the decoding result of each of the above-mentioned decoding parts to determine the speech recognition result.

[0104] By employing the above technical solution, the speech recognition device provided in this application embodiment acquires speech information, which includes speech in at least two languages. It then recognizes the speech information based on a multi-granularity speech recognition model, wherein the multi-granularity speech recognition model includes at least two independent decoding parts, each of which corresponds to a different combination of granularity in the training set. The decoding result of each decoding part is acquired to determine the speech recognition result. Due to the data-driven nature of speech recognition systems, the more complete and larger the training data, the better the performance of the recognition system. Because the different combinations of granularity in the training set corresponding to the decoding parts allow for multiple different label sequences to be obtained from the same mixed-language audio, the overall usability of the mixed-language training corpus increases several times, alleviating the scarcity of training set data to some extent and improving the accuracy of speech recognition. Furthermore, when training the speech recognition model by superimposing label sequences of different granularities onto the same mixed-language audio, the speech recognition model considers multiple pronunciation combinations, further improving the accuracy of speech recognition. Furthermore, since the multi-granularity speech recognition model includes at least two independent decoding modules, the decoding parts of the models trained from training sets of different granularities are independent of each other. By obtaining the decoding results of each decoding part separately, the decoding result with the highest matching degree to the pronunciation combination of the audio to be decoded can be selected as the final decoding result, further improving the recognition performance and accuracy of the speech recognition model.

[0105] In some examples, the combination of languages ​​and granularity differs for each of the training sets mentioned above.

[0106] In some examples, the training granularity described above may include a single character, a word segmentation unit, and a word, wherein the word segmentation unit is composed of the single character and the word is composed of the word segmentation unit.

[0107] In some examples, the aforementioned determining unit can specifically be used for:

[0108] Obtain the decoding result for each of the above decoding parts;

[0109] Calculate the confidence score of the decoding result for each of the above decoding parts;

[0110] The decoding result with the highest confidence score is determined as the speech recognition result.

[0111] In some examples, the aforementioned determining unit can also be used for:

[0112] Compare the obtained decoding results. If there are identical decoding results, combine the posterior probabilities of the identical decoding results to calculate the confidence score of the identical decoding results.

[0113] In some examples, the aforementioned determining unit can also be used for:

[0114] Calculate the size of the resulting vocabulary associated with each of the above decoding parts;

[0115] Based on the size of each of the above-mentioned result vocabularies obtained through calculation, update the confidence score of the decoding result of the above-mentioned decoding part associated with the above-mentioned result vocabulary.

[0116] In some examples, the resulting vocabulary of the above speech recognition model may include natural mixed Chinese and English text outside the training set and / or text obtained after processing a single language text within the training set.

[0117] In some examples, the processing of monolingual text within the training set described above may include replacing names and / or pronouns in the monolingual text with those in other languages.

[0118] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured; adjusting kernel parameters can address the current problem of poor accuracy in speech recognition.

[0119] This application provides a storage medium on which a program is stored, which, when executed by a processor, implements the speech recognition method.

[0120] This application provides a processor for running a program, wherein the program executes the speech recognition method during runtime.

[0121] This application provides a device 400, such as... Figure 4 As shown, the device includes at least one processor 410 and at least one memory 420 connected to the processor; wherein the processor 410 and the memory 420 communicate with each other; the processor 410 is used to call program instructions in the memory 420 to execute the above-described speech recognition method.

[0122] The devices mentioned in this article can be servers, PCs, tablets, mobile phones, etc.

[0123] This application also provides a computer program product, which, when executed on a process management device, is suitable for executing an initialization program having the following method steps: acquiring speech information, the speech information including speech in at least two languages; recognizing the speech information based on a multi-granularity speech recognition model, wherein the multi-granularity speech recognition model includes at least two mutually independent decoding parts, and the granularity combination of the training set corresponding to each decoding part is different; and acquiring the decoding result of each decoding part to determine the speech recognition result.

[0124] In some examples, the combination of languages ​​and granularity differs for each of the training sets mentioned above.

[0125] In some examples, the training granularity mentioned above includes single characters, word segments, and words, wherein the word segments are composed of the single characters and the words are composed of the word segments.

[0126] In some examples, obtaining the decoding result of each of the decoding parts to determine the speech recognition result may include:

[0127] Obtain the decoding result for each of the above decoding parts;

[0128] Calculate the confidence score of the decoding result for each of the above decoding parts;

[0129] The decoding result with the highest confidence score is determined as the speech recognition result.

[0130] In some examples, before determining the decoding result with the highest confidence score as the speech recognition result, the following may also be included:

[0131] The obtained decoding results are compared. If there are identical decoding results, the posterior probabilities of the identical decoding results are combined to calculate the confidence score of the identical decoding results.

[0132] In some examples, before determining the decoding result with the highest confidence score as the speech recognition result, the above method may further include:

[0133] Calculate the size of the resulting vocabulary associated with each of the above decoding parts;

[0134] Based on the size of each of the above-mentioned result vocabularies obtained through calculation, update the confidence score of the decoding result of the above-mentioned decoding part associated with the above-mentioned result vocabulary.

[0135] In some examples, the resulting vocabulary of the above speech recognition model includes natural mixed Chinese and English text outside the training set and / or text obtained after processing a single language text within the training set.

[0136] In some examples, the processing of monolingual text within the training set described above includes replacing names and / or pronouns in the monolingual text with those in other languages.

[0137] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable process management device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable process management device, generate instructions for implementing the process... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0138] In a typical configuration, the device includes one or more processors (CPUs), memory, and a bus. The device may also include input / output interfaces, network interfaces, etc.

[0139] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM, and memory includes at least one memory chip. Memory is an example of computer-readable media.

[0140] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0141] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0142] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0143] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A speech recognition method, characterized in that, include: Acquire mixed-language speech information, wherein the speech information includes speech from at least two languages; The speech information is recognized based on a multi-granularity speech recognition model, wherein the multi-granularity speech recognition model includes at least two independent decoding parts. Each decoding part corresponds to a different combination of languages ​​and speech information splitting granularities in the training set. The same mixed language speech information is split based on different granularities to obtain different label sequences by splitting and combining the same mixed language speech information. The granularity includes a single character, a word segmentation unit, and a word. The word segmentation unit is composed of the single character, and the word is composed of the word segmentation unit. The decoding results of each decoding part are obtained separately to determine the speech recognition result.

2. The method according to claim 1, characterized in that, The step of obtaining the decoding result of each of the decoding parts to determine the speech recognition result includes: Obtain the decoding result for each of the decoding parts; Calculate the confidence score of the decoding result for each of the decoding parts; The decoding result with the highest confidence score is determined as the speech recognition result.

3. The method according to claim 2, characterized in that, Before determining the decoding result with the highest confidence score as the speech recognition result, the process also includes: The obtained decoding results are compared. If there are identical decoding results, the posterior probabilities of the identical decoding results are combined to calculate the confidence score of the identical decoding results.

4. The method according to claim 2, characterized in that, Before determining the decoding result with the highest confidence score as the speech recognition result, the method further includes: Calculate the size of the resulting vocabulary associated with each of the decoded parts; Based on the size of the resulting vocabulary, update the confidence score of the decoding result of the decoding part associated with the resulting vocabulary.

5. The method according to claim 4, characterized in that, The resulting vocabulary of the speech recognition model includes natural mixed Chinese and English text outside the training set and / or text obtained after processing a single language text within the training set.

6. The method according to claim 5, characterized in that, The processing method for single-language texts within the training set includes replacing names and / or pronouns in the single-language texts with other languages.

7. A voice recognition device, characterized in that, include: An acquisition unit is used to acquire mixed language speech information, wherein the speech information includes speech from at least two languages; The recognition unit is used to recognize the speech information based on a multi-granularity speech recognition model. The multi-granularity speech recognition model includes at least two independent decoding parts. Each decoding part corresponds to a different combination of languages ​​and speech information splitting granularities in the training set. The same mixed language speech information is split based on different granularities to obtain different label sequences by splitting and combining the same mixed language speech information. The granularity includes a single character, a word segmentation unit, and a word. The word segmentation unit is composed of the single character, and the word is composed of the word segmentation unit. A determining unit is used to obtain the decoding result of each of the decoding parts respectively, so as to determine the speech recognition result.

8. The apparatus according to claim 7, characterized in that, The determining unit is specifically used for: Obtain the decoding result for each of the decoding parts; Calculate the confidence score of the decoding result for each of the decoding parts; The decoding result with the highest confidence score is determined as the speech recognition result.

9. The apparatus according to claim 8, characterized in that, The determining unit is further configured to: The obtained decoding results are compared. If there are identical decoding results, the posterior probabilities of the identical decoding results are combined to calculate the confidence score of the identical decoding results.

10. The apparatus according to claim 8, characterized in that, The determining unit is further configured to: Calculate the size of the resulting vocabulary associated with each of the decoded parts; Based on the calculated size of each of the resulting vocabulary lists, the confidence score of the decoding result of the decoding part associated with the resulting vocabulary list is updated.

11. The apparatus according to claim 10, characterized in that, The resulting vocabulary of the speech recognition model includes natural mixed Chinese and English text outside the training set and / or text obtained after processing a single language text within the training set.

12. The apparatus according to claim 11, characterized in that, The processing method for single-language texts within the training set includes replacing names and / or pronouns in the single-language texts with other languages.

13. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, it controls the device where the storage medium is located to perform the speech recognition method as described in any one of claims 1 to 6.

14. An electronic device, characterized in that, The device includes at least one processor and at least one memory connected to the processor; the processor is used to call program instructions in the memory to execute the speech recognition method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Low-resource and multi-language voice recognition model and voice recognition method

    CN110428818A

  • Multi-lingual speech recognition with cross-language context modeling

    US20040088163A1