A speech recognition method, apparatus, device, and storage medium

By extracting the audio features of each language in the target audio and combining language information, the encoder, language recognition and decoding module in the speech recognition model is solved, and the problem of poor mixed speech recognition effect in multilinguals is achieved, achieving higher recognition accuracy.

CN115440217BActive Publication Date: 2025-07-11XIAN XUNFEI SUPER BRAIN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211042586.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-07-11
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

When existing speech recognition technology processes mixed speech in multiple languages, the recognition effect is poor, making it difficult to accurately identify mixed speeches containing multiple languages.

Method used

By extracting the audio features corresponding to each language in the target audio, and combining language information, the speech recognition model is used for speech recognition, including the encoder module, the language recognition module and the decoding module, to reduce the pericarticular interference between different languages.

Benefits of technology

It improves the accuracy of multilingual mixed speech recognition, reduces the interference of similar pronunciation words between different languages on the recognition results, and improves the recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115440217B_ABST
    Figure CN115440217B_ABST
Patent Text Reader

Abstract

The present application discloses a speech recognition method, apparatus, device and storage medium. The method includes: obtaining a target audio to be recognized, where the target audio is a mixed audio including at least two languages; extracting first audio features corresponding to each language in the target audio; and obtaining a speech recognition result of the target audio based on the first audio features corresponding to each language and the language information in the target audio. By the above method, the present application can improve the accuracy of the speech recognition result of multi-language mixed speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition, and in particular, to a speech recognition method, apparatus, device, and storage medium. Background Art

[0002] Currently, there already exist technologies that can directly map speech to text to achieve end-to-end speech recognition. However, current speech recognition technologies have strong recognition capabilities for single languages, but have poor recognition effects for mixed-language speech, especially for mixed-language speech that contains multiple languages in one sentence, which poses a great challenge to multi-language speech recognition.

[0003] In summary, how to improve the accuracy of multi-language mixed speech recognition results is of great significance. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a speech recognition method, apparatus, device, and storage medium that can improve the accuracy of multi-language mixed speech recognition results.

[0005] To solve the above technical problem, one technical solution adopted by this application is: to provide a speech recognition method, which includes: obtaining a target audio to be recognized, where the target audio is a mixed audio including at least two languages; extracting first audio features corresponding to each language in the target audio; and obtaining a speech recognition result of the target audio based on the first audio features corresponding to each language and the language information in the target audio.

[0006] To solve the above technical problem, another technical solution adopted by this application is: to provide a speech recognition apparatus, which includes: an obtaining module for obtaining a target audio to be recognized, where the target audio is a mixed audio including at least two languages; a feature extraction module for extracting first audio features corresponding to each of the languages in the target audio; and a speech recognition module for obtaining a speech recognition result of the target audio based on the first audio features corresponding to the language and the language information in the target audio.

[0007] To solve the above technical problem, another technical solution adopted by this application is: to provide an electronic device, including a memory and a processor coupled to each other, where the memory stores program instructions; and the processor is configured to execute the program instructions stored in the memory to implement the above method.

[0008] To solve the above technical problem, another technical solution adopted by this application is: to provide a computer-readable storage medium, where the computer-readable storage medium is used to store program instructions, and the program instructions can be executed to implement the above method.

[0009] The beneficial effects of the present application are as follows: After obtaining the target audio to be recognized, which includes at least two languages, the present application first extracts the first audio features corresponding to each language in the target audio, and then obtains the speech recognition result of the target audio based on the first audio features corresponding to each language and the language information in the target audio. Since the present application combines the audio features in the target audio and the language information corresponding to the target audio in the speech recognition process, the interference of near-sounding words between different languages on the recognition result can be reduced, and thus the accuracy of the multi-language mixed speech recognition result can be improved. Description of the Drawings

[0010] Figure 1 is a schematic flowchart of an embodiment of the speech recognition method provided by the present application;

[0011] Figure 2 is a schematic framework diagram of the speech recognition model provided by the present application;

[0012] Figure 3 is a schematic flowchart of an embodiment of the speech recognition method provided by the present application;

[0013] Figure 4 is Figure 1 a schematic flowchart of an embodiment of the shown step S13;

[0014] Figure 5 is a schematic framework diagram of an embodiment of the speech recognition device provided by the present application;

[0015] Figure 6 is a schematic structural diagram of an embodiment of the electronic device provided by the present application;

[0016] Figure 7 is a schematic structural diagram of the computer-readable storage medium provided by the present application. Detailed Embodiments

[0017] To make the objectives, technical solutions, and effects of the present application clearer and more definite, the following further describes the present application in detail with reference to the accompanying drawings and by way of examples.

[0018] It should be noted that if there are descriptions involving "first", "second", etc. in the embodiments of the present application, such descriptions of "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first", "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those skilled in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0019] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the speech recognition method provided by this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 1 the process sequence shown. As Figure 1 shown, this embodiment includes:

[0020] S11: Obtain a target audio to be recognized, where the target audio is a mixed audio including at least two languages.

[0021] The method of this embodiment is used to obtain a speech recognition result of the target audio based on the first audio features corresponding to each language in the target audio and the language information of the target audio. Among them, a speech recognition model can be used to comprehensively combine the first audio features corresponding to each language and the language information of the target audio to obtain the speech recognition result of the target audio.

[0022] The target audio is a mixed audio including at least two languages. Among them, the target audio can be a mixed audio of two languages, or a multi-language mixed audio of more than two languages. In addition, the target audio can be an inter-sentence mixed audio of at least two languages (each sentence corresponds to speech that only contains one language, but the languages corresponding to different sentences are different), or an intra-sentence mixed audio of at least two languages. For example, "I feel very happy today" is an intra-sentence mixed audio containing two languages. In the intra-sentence mixed audio of at least two languages, each sentence contains at least two languages. Among them, at least two languages can be a mixed speech of Chinese and foreign languages, or a mixed speech of Mandarin and various dialects. The specific types of languages included in the target audio and the mixing situation need to be determined according to the actual application scenario, and no specific limitation is made here.

[0023] S12: Extract the first audio features corresponding to each language in the target audio.

[0024] In one implementation manner, relevant feature extraction algorithms can be used to extract the first audio features corresponding to each language in the target audio.

[0025] In other implementation manners, a speech recognition model can also be used to extract the first audio features corresponding to each language in the target audio. In a specific implementation manner, as Figure 2As shown, the encoder modules corresponding to each language in the speech recognition model can be used to extract the first audio features corresponding to each language in the target audio. For example, if the target audio contains two languages, language 1 and language 2, the encoder module 1 corresponding to language 1 is used to extract the first audio feature corresponding to language 1 in the target audio, and the encoder module 2 corresponding to language 2 is used to extract the first audio feature corresponding to language 2 in the target audio. Among them, each extracted first audio feature at least includes the audio features of the corresponding language. Optionally, each extracted first audio feature may further include the position information of the audio corresponding to each language relative to the target audio. For example, it includes which frames of the target audio the audio corresponding to each language is.

[0026] Among them, the encoder modules corresponding to each language are trained using the first single-language sample audio of the corresponding language. For example, if encoder module 1 is a Mandarin encoder and encoder module 2 is a Sichuan dialect encoder, the Mandarin encoder is trained using Mandarin audio data (the first single-language sample audio), and the Sichuan dialect encoder is trained using Sichuan dialect audio data (the first single-language sample audio) to obtain the encoder module 1 corresponding to Mandarin and the encoder module 2 corresponding to Sichuan dialect.

[0027] S13: Based on the first audio features corresponding to each language and the language information in the target audio, obtain the speech recognition result of the target audio.

[0028] In an embodiment, before obtaining the speech recognition result of the target audio based on the first audio features corresponding to each language and the language information in the target audio, the language information of the target audio needs to be obtained first. In one implementation, the first audio features corresponding to each language are fused first using the splicing method to obtain the first audio fusion feature, or the first audio features corresponding to each language are fused according to the position information of each extracted first audio feature relative to the target audio to obtain the first audio fusion feature. Then, based on the first fusion feature, the language information of the target audio is obtained. Among them, the obtained language information of the target audio can be the language information corresponding to different audio segments, or the language information corresponding to each audio frame in the target audio.

[0029] In one implementation, as Figure 2 shown, the speech recognition model further includes a language recognition module. Among them, the language information in the target audio is obtained using the language recognition module of the speech recognition model. Specifically, the language recognition module of the speech recognition model is used to process the first audio fusion feature to obtain the language information in the target audio. Among them, the language recognition module is trained using the first sample audio and the trained encoder modules. The first sample audio is a mixed audio including multiple languages (at least two languages), and is labeled with the language information (language label) y corresponding to each audio frame in the first sample audio. i, therefore, the language recognition module trained with the first sample audio can recognize the language information corresponding to each speech frame in the target audio. The training steps of the language recognition module are as follows: First, use the encoder modules corresponding to each language that have been trained to extract the sample audio features corresponding to each language in the first sample audio. Then, fuse the sample audio features and input them into the language recognition module to obtain the language information z corresponding to each audio frame in the first sample audio. i , and then, use the cross-entropy loss function of the language information to adjust the parameters. The cross-entropy loss function of the language information is as follows:

[0030]

[0031] Among them, L represents the loss function, c represents the number of samples of the first sample audio, Tc represents the number of frames of the first sample audio, i represents the i-th frame, and y i represents the true language label, and z i represents the predicted language probability value.

[0032] In one embodiment, as Figure 2 shown, in addition to including the encoder modules corresponding to each language and the language recognition module, the speech recognition model further includes a decoding module. Among them, the steps of obtaining the speech recognition result of the target audio based on the first audio features corresponding to each language and the language information in the target audio can be obtained by using the decoding module of the speech recognition model. The decoding module is trained with the second sample audio, and the second sample audio includes at least one of the second single-language sample audio and the multi-language sample audio. That is to say, the decoding module can be trained with the second single-language sample audio (each single-language sample audio). Among them, the second single-language sample audio and the first single-language sample audio can be the same or different. Of course, the decoding module can also be trained with the multi-language sample audio (that is, each sample audio includes multiple languages), or can be trained with the second single-language sample audio and the multi-language sample audio. Specifically, first train the decoding module with the second single-language sample audio with a large amount of data and perform data tuning, and then perform fine-tuning with the multi-language sample audio. The training steps of the decoding module in the speech recognition model include: First, use the encoder modules corresponding to each language that have been trained to extract the sample audio features corresponding to each language in the second sample audio, and the trained language recognition module to obtain the language information in the second sample audio. Then, based on the sample audio features corresponding to each language and the language information in the second sample audio, obtain the sample speech recognition result of the second sample audio. After that, use the sample speech recognition result to adjust the parameters of the decoding module.

[0033] Of course, in other embodiments, the step of obtaining the speech recognition result of the target audio based on the first audio features corresponding to each language and the language information in the target audio can also be obtained by using a relevant text recognition algorithm.

[0034] It should be noted that in some scenarios, there may be similar-sounding words between different languages, such as the English word "shy" and the Chinese word "晒", the English word "low" and the Chinese word "漏". For the same word, there may be differences in tones between Mandarin Chinese and various dialects due to different pronunciation habits. For example, the first tone in Mandarin is often pronounced as the second tone in Henan dialect, the second tone in Mandarin is often pronounced as the fourth tone in Henan dialect, and the third tone in Mandarin is often pronounced as the first tone in Henan dialect. For example, "河南人" has the second tone in Mandarin (hé nán rén), but is often pronounced as the fourth tone in Henan dialect (hè nàn rèn). It is understandable that if the language information corresponding to the target audio is not combined in the target audio recognition process, such similar-sounding words will interfere with speech recognition and cause recognition errors.

[0035] In this embodiment, after obtaining the target audio to be recognized including at least two languages, the first audio features corresponding to each language in the target audio are first extracted, and then the speech recognition result of the target audio is obtained based on the first audio features corresponding to each language and the language information in the target audio. Since the scheme of this embodiment combines the first audio features in the target audio and the language information corresponding to the target audio in the speech recognition process, the interference of near-sounding words in different languages ​​on the recognition result can be reduced, thereby improving the accuracy of the speech recognition result.

[0036] It should be noted that the language recognition module of the above-mentioned speech recognition model is obtained by training using the first sample audio data, and a large amount of training sample data is required during the model training process, among which the first sample audio data of multi-language mixture is scarce, especially the first sample audio data of similar languages ​​(such as Mandarin and corresponding dialects) mixed in sentences is even more scarce. In order to enhance the first sample audio data used to train the language recognition module to ensure the training effect of the language recognition module, a sufficient number of first sample audio data can be obtained first. In one embodiment, the first sample audio can also be obtained by performing entity translation, insertion, replacement and other operations on the monolingual audio to construct a multi-language mixed voice segment, and then using a speech synthesizer to synthesize the mixed voice segments of each multi-language, and the first sample audio data is marked with the language information corresponding to each audio frame. In other embodiments, the first sample audio is obtained by cutting multiple monolingual audios (wherein the cutting point for cutting is determined based on the target adjacent audio frames belonging to different words of the monolingual audio) and mixing and splicing.

[0037] Specifically, seeFigure 3 , Figure 3 is a schematic flowchart of an embodiment of the speech recognition method provided in this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 3 the process sequence shown. As Figure 3 shown, in this embodiment, the first sample audio data is obtained by cropping and mixing multiple single-language audios, specifically including:

[0038] S31: Obtain multiple single-language audios, and find the target adjacent audio frames belonging to different characters in each single-language audio.

[0039] Specifically, first obtain an appropriate number of multiple single-language audios. Among them, the number of each single language can be, but is not limited to, close to 1:1, or can also be close to 2:1, 3:1, etc. The specific languages and quantities of each single-language audio can be determined according to the actual application scenario. For example, if it corresponds to recognizing the mixed speech of Mandarin and Sichuan dialect, then it corresponds to the audios of separate Mandarin and separate Sichuan dialect. Another example is that if it corresponds to recognizing the mixed speech of Mandarin and Shaanxi dialect, then it corresponds to the audios of separate Mandarin and separate Shaanxi dialect. Here, only examples are used, which do not limit the languages corresponding to the multiple single-language audios, nor the languages included in the first sample audio data.

[0040] In one implementation manner, an existing single-language speech recognition system based on the deep neural network-hidden Markov model architecture (DNN-HMM) can be used to decode each single-language audio frame by frame, and find the target adjacent audio frames belonging to different characters in each single-language audio. That is to say, in this embodiment, each frame in two target adjacent audio frames belongs to different characters.

[0041] S32: For each single-language audio, based on the target adjacent audio frames of the single-language audio, determine at least one cropping point of the single-language audio, and crop the single-language audio according to the at least one cropping point to obtain multiple audio segments of the single-language audio, where at least one audio segment of the single-language audio contains at least two characters.

[0042] In this embodiment, after finding the target adjacent audio frames belonging to different characters in each monolingual audio, it is necessary to determine at least one clipping point of the monolingual audio from the target adjacent audio frames, and clip the monolingual audio according to the at least one clipping point to obtain multiple audio segments of the monolingual audio. Among them, at least one audio segment of the monolingual audio contains at least two characters, that is to say, for the subsequent first sample audio spliced from at least one audio segment of two or more monolingual audios, there may not necessarily be a language switch before and after each character. It can be understood that since there may not necessarily be a language switch before and after each character, the position of the language switch is not fixed. Therefore, while ensuring that the language recognition module learns the language switch information, the interference caused by audio splicing can be suppressed.

[0043] In one implementation, after finding the target adjacent audio frames belonging to different characters in each monolingual audio, at least one clipping point of the monolingual audio can be determined according to a certain probability distribution (such as Poisson distribution).

[0044] S33: Select at least one audio segment of two or more monolingual audios for splicing to obtain the first sample audio.

[0045] In one implementation, after obtaining the respective audio segments corresponding to each monolingual language, at least one audio segment of two or more monolingual audios is randomly selected for splicing to obtain the first sample audio. Among them, since the first sample audio is spliced from at least one randomly selected audio segment of each monolingual language, this splicing method can be applied to each audio sentence or each audio segment, and then each audio segment is clipped into individual audio sentences to obtain the first sample audio. Therefore, the first sample audio obtained by this random splicing may contain both intra-sentence mixed audio and inter-sentence mixed audio. In some embodiments, for better training to recognize the language information corresponding to each frame of audio in intra-sentence mixed audio, each audio sentence is spliced from at least one audio segment of two or more randomly selected monolingual audios to obtain the first sample audio that completely belongs to intra-sentence mixed audio.

[0046] It should be noted that the purpose of obtaining the first sample audio data in this embodiment is only to train a language recognition module capable of performing language recognition, rather than obtaining the text information corresponding to the first sample audio data. Therefore, the obtained first sample audio data can be an incomplete or unsmooth sentence. In addition, in order to train the language recognition module of the speech recognition model, the corresponding language information can be labeled for each audio frame in the first sample audio data. The labeling method can be, but is not limited to, adding a 1-hot language label for each frame of audio, or using a word embedding representation method, etc. It can be understood that a language recognition module trained with the first sample audio data in which each audio frame contains the corresponding language information annotation can recognize the language information corresponding to each frame of audio in the target audio.

[0047] Meanwhile, it should be noted that for the in-sentence mixed speech recognition of multiple languages, especially for the in-sentence mixed speech recognition of languages that are similar (Mandarin and its corresponding dialects), there is obvious crosstalk of near-homophonic words compared with the in-sentence mixed speech recognition of multiple languages. Therefore, the language recognition module trained using the in-sentence mixed audio of multiple languages as the first sample audio data can also obtain good recognition results for the in-sentence mixed speech of multiple languages. Of course, if it is only used for recognizing the in-sentence mixed speech of multiple languages, the in-sentence mixed speech of multiple languages can also be specifically used as the first sample audio data to train a language recognition module that can recognize the in-sentence mixed speech of multiple languages.

[0048] In one embodiment, as Figure 2 shown, the speech recognition model includes an encoder module corresponding to each language, a language recognition module, and a decoding module. After each module of the speech recognition model has been trained and optimized, first use the encoder module of the speech recognition model to extract the first audio features corresponding to each language in the target audio, then splice the first audio features corresponding to each language to obtain the first audio fusion feature. After that, use the language recognition module of the speech recognition model to process the first audio fusion feature to obtain the language information in the target audio. Then, based on the first audio features corresponding to each language and the language information in the target audio, obtain the speech recognition result of the target audio, where the obtained language information in the target audio includes the language information corresponding to each audio frame in the target audio.

[0049] Specifically, please refer to Figure 4 , Figure 4 is Figure 1 a schematic flowchart of one embodiment of step S13 shown. It should be noted that if there are substantially the same results, this embodiment is not limited to the Figure 4 shown process sequence. As Figure 4 shown, in this embodiment, based on the first audio features corresponding to each language and the language information in the target audio, obtaining the speech recognition result of the target audio specifically includes:

[0050] S41: Based on the first audio features corresponding to each language and the language information corresponding to each audio frame, obtain the language audio features corresponding to multiple coding stages, and obtain the decoded text of each coding stage based on the language audio features corresponding to each coding stage.

[0051] In this embodiment, multiple encoding stages are executed sequentially. The decoded texts of different encoding stages represent the text information of different audio parts in the target audio. That is to say, the text information contained in the target audio is not obtained by decoding the decoding module at one time, but the decoding module sequentially decodes the language audio features corresponding to each encoding stage. Among them, the language audio features corresponding to each encoding stage are the input features of the decoding module, and the input features are obtained based on the first audio features corresponding to each language and the language information corresponding to each audio frame. Among them, the language audio features corresponding to each encoding stage include the second audio features corresponding to each language in the encoding stage.

[0052] In one implementation manner, after obtaining the second audio features corresponding to each language in the current encoding stage, the currently executed encoding stage is used as the current encoding stage for multi-stage decoding. Specifically, the second audio features corresponding to each language in the current encoding stage can be fused by splicing to obtain the second fused audio feature corresponding to the current encoding stage. Then, the second fused audio feature corresponding to the current encoding stage is decoded to obtain the decoded text of the current encoding stage.

[0053] Among them, in one implementation manner, obtaining the language audio features corresponding to multiple encoding stages based on the first audio features corresponding to each language and the language information corresponding to each audio frame includes: using the currently executed encoding stage as the current encoding stage. Among them, for each language included in the target audio, based on the first audio feature corresponding to the language, the language information of each audio frame corresponding to the language, and the decoded text of the reference encoding stage, the second audio features corresponding to each language in the current encoding stage are obtained, where the reference encoding stage is the encoding stage executed before the current encoding stage.

[0054] In a specific implementation manner, the decoding module is a decoding module based on the Attention mechanism. Among them, the Attention mechanism is used to use the language information to guide the weight allocation of the second audio features corresponding to multiple languages at different time points in the sequence to achieve decoding of the corresponding language for each audio frame. Specifically, referring to formulas (2) to (4), taking the target audio including language p and language x as an example, first use the encoding stage to be decoded as the current encoding stage. Then, for each audio frame corresponding to each language, based on the first audio feature corresponding to the audio frame, the language information corresponding to the audio frame, and the decoded text of the reference encoding stage, the attention weights of the audio frame are obtained using formulas (2) to (3). Then, based on the first audio features corresponding to each audio frame corresponding to the language and the attention weights, the second audio features corresponding to each language in the current encoding stage are obtained using (4), that is, the second audio features c p,u and c x,u .

[0055]

[0056]

[0057]

[0058] In formula (2), p represents language p, u-1 is the decoded text in the reference encoding stage, u is the text to be decoded in the current encoding stage, and z LID,i is the language information corresponding to the i-th frame in the target audio output by the language recognition module. represents the feature vector of language p corresponding to the i-th frame. represents the feature vector corresponding to the decoded text u-1 in the reference encoding stage, W p 、W LID 、V, and b are parameters determined through training, and e p,u,i is the initial attention weight of the i-th frame for the text u to be decoded in the current encoding stage, with a value in [-1, 1].

[0059] In formula (3), a p,u,i represents the attention weight after normalizing the initial weight e p,u,i with a value in [0, 1].

[0060] In formula (4), T represents the number of audio frames of the target audio, and c p,u represents the concatenated vector obtained by multiplying the feature vectors of each audio frame of language p by the corresponding normalized weights.

[0061] Similarly, the feature vector obtained by multiplying the feature vectors of each audio frame of language x by the corresponding normalized weights can be calculated using formulas (2) to (4), and the feature vector c x,u . Among them, c p,u and c x,u respectively represent the second audio features corresponding to language p and language x. By fusing c p,u and c x,u the second fused audio feature in the current encoding stage is obtained, and then the decoding module decodes the second fused audio feature corresponding to the current encoding stage to obtain the decoded text in the current encoding stage.

[0062] It should be noted that in some scenarios, due to individual differences, different languages are not completely different, and there will be similar situations. For example, the speech of some people in Henan is similar to Mandarin, but the speech of some people in Henan is quite different from Mandarin. If the language information is strictly used to distinguish Mandarin and Henan dialect in the target audio, errors are likely to occur. Therefore, in the embodiment of the present method, the decoding module does not strictly use the language information corresponding to each speech frame for decoding, but through the Attention mechanism, uses the language information corresponding to each speech frame to calculate the attention weight of the language information in each speech frame to the text in the current decoding stage, so that the decoding module decodes by combining the attention weight of the language information in each speech frame to the current text to be decoded, thereby improving the accuracy of the decoding result of the decoding module.

[0063] S42: Obtain the speech recognition result of the target audio based on the decoded texts of each encoding stage.

[0064] In this embodiment, after obtaining the decoded texts of each encoding stage in step S41, the decoded texts of each encoding stage can be integrated to obtain the speech recognition result of the target audio. Among them, the speech recognition result of the target audio is the text corresponding to the target audio.

[0065] Please refer to Figure 5 , Figure 5 is a schematic framework diagram of an embodiment of the speech recognition device provided by the present application. In this embodiment, the speech recognition device 50 includes an acquisition module 51, a feature extraction module 52, and a speech recognition module 53. The acquisition module 51 is used to acquire the target audio to be recognized, and the target audio is a mixed audio including at least two languages; the feature extraction module 52 is used to extract the first audio features corresponding to each language in the target audio; the speech recognition module 53 is used to obtain the speech recognition result of the target audio based on the first audio features corresponding to each language and the language information in the target audio.

[0066] In some embodiments, before the speech recognition module 53 obtains the speech recognition result of the target audio based on the first audio features corresponding to each language and the language information in the target audio, it further includes: fusing the first audio features corresponding to each language to obtain a first audio fusion feature; obtaining the language information in the target audio based on the first audio fusion feature.

[0067] In some embodiments, obtaining the language information corresponding to each audio frame based on the first audio fusion feature includes: using the language recognition module of the speech recognition model to process the first audio fusion feature to obtain the language information in the target audio; wherein, the language recognition module is trained using the first sample audio, and the first sample audio is a mixed audio including multiple languages and is labeled with the language information corresponding to each audio frame in the first sample audio.

[0068] In some embodiments, the steps for the acquisition module 51 to acquire the first sample audio include: acquiring a plurality of monolingual audios, and finding target adjacent audio frames belonging to different characters in each monolingual audio; for each monolingual audio, based on the target adjacent audio frames of the monolingual audio, determining at least one cutting point of the monolingual audio, and cutting the monolingual audio according to the at least one cutting point to obtain a plurality of audio segments of the monolingual audio, wherein at least one audio segment of the monolingual audio contains at least two characters; selecting at least one audio segment of two or more monolingual audios for splicing to obtain the first sample audio.

[0069] In some embodiments, the steps for the feature extraction module 52 to extract the first audio features corresponding to each language in the target audio are executed by the encoder modules corresponding to each language of the speech recognition model; the steps for obtaining the speech recognition result of the target audio based on the first audio features corresponding to each language and the language information in the target audio are obtained by the decoding module of the speech recognition model.

[0070] In some embodiments, the encoder modules corresponding to each language are trained using the first monolingual sample audio of the corresponding language, and the decoding module is trained using the second sample audio, and the second sample audio includes at least one of the second monolingual sample audio and the multilingual sample audio; and / or, the language information in the target audio is obtained using the language recognition module of the speech recognition model, and the training steps of the decoding module include: using the trained encoder module to extract the sample audio features corresponding to each language in the second sample audio, and using the trained language recognition module to obtain the language information in the second sample audio, wherein the language recognition module is trained using the first sample audio and the trained encoder module; obtaining the sample speech recognition result of the second sample audio based on the sample audio features corresponding to each language and the language information in the second sample audio; adjusting the parameters of the decoding module using the sample speech recognition result.

[0071] In some embodiments, the language information in the target audio acquired by the acquisition module 51 includes the language information corresponding to each audio frame in the target audio; obtaining the speech recognition result of the target audio based on the first audio features corresponding to each language and the language information in the target audio includes: obtaining the language audio features corresponding to each coding stage based on the first audio features corresponding to each language and the language information corresponding to each audio frame, and obtaining the decoded text of each coding stage based on the language audio features corresponding to each coding stage, wherein a plurality of coding stages are executed in sequence, and the decoded texts of different coding stages represent the text information of different audio parts in the target audio, and the language audio features corresponding to the coding stage include the second audio features corresponding to each language in the coding stage; obtaining the speech recognition result of the target audio based on the decoded text of each coding stage.

[0072] In some embodiments, the speech recognition module 53 obtains language audio features corresponding to multiple encoding stages based on the first audio features corresponding to each language and the language information corresponding to each audio frame, including: taking the currently executed encoding stage as the current encoding stage; for each language included in the target audio, obtaining the second audio feature corresponding to the language in the current decoding stage based on the first audio feature corresponding to the language, the language information of each audio frame corresponding to the language, and the decoded text of the reference encoding stage, where the reference encoding stage is the encoding stage executed before the current encoding stage; and / or obtaining the decoded text of each encoding stage based on the language audio features corresponding to each encoding stage, including: taking the currently executed encoding stage as the current encoding stage; fusing the second audio features corresponding to each language in the current encoding stage to obtain the second fused audio feature corresponding to the current encoding stage; and decoding the second fused audio feature corresponding to the current encoding stage to obtain the decoded text of the current encoding stage.

[0073] In some embodiments, the feature extraction module 52 obtains the second audio feature corresponding to the language in the current encoding stage based on the first audio feature corresponding to the language, the language information of each audio frame corresponding to the language, and the decoded text of the reference encoding stage, including: for each audio frame corresponding to the language, obtaining the attention weight of the audio frame based on the first audio feature corresponding to the audio frame, the language information of the audio frame, and the decoded text of the reference encoding stage; and obtaining the second audio feature corresponding to the language in the current encoding stage based on the first audio features corresponding to each audio frame of the language and the attention weight.

[0074] Please refer to Figure 6 , Figure 6 FIG. is a schematic structural diagram of an embodiment of an electronic device provided by the present application. In this embodiment, the electronic device 60 includes a processor 61 and a memory 62.

[0075] The processor 61 may also be referred to as a CPU (Central Processing Unit). The processor 61 may be an integrated circuit chip with signal processing capabilities. The processor 61 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor, or the processor 61 may also be any conventional processor 61, etc.

[0076] The memory 62 in the electronic device 60 is used to store program instructions required for the operation of the processor 61.

[0077] The processor 61 is used to execute program instructions to implement the methods provided by any of the above embodiments and any non-conflicting combinations.

[0078] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of the computer-readable storage medium provided by this application. The computer-readable storage medium 70 in the embodiments of this application stores program instructions 71, and when the program instructions 71 are executed, they implement the methods provided by any of the above embodiments and any non-conflicting combinations. Among them, the program instructions 71 can form a program file and be stored in the above computer-readable storage medium 70 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) can execute all or part of the steps of the methods in various embodiments of this application. And the aforementioned computer-readable storage medium 70 includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.

[0079] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0080] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The similarities or similarities between them can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0081] In several embodiments provided by this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0082] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0083] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0084] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0085] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A speech recognition method, characterized in that, The method includes: Obtaining a target audio to be recognized, where the target audio is a mixed audio including at least two languages; Extracting first audio features corresponding to each of the languages in the target audio; Based on the first audio features corresponding to each of the languages and the language information in the target audio, obtaining a speech recognition result of the target audio; the language information is obtained by processing the first audio fusion feature using a language recognition module of a speech recognition model after fusing the first audio features corresponding to each of the languages to obtain a first audio fusion feature; Wherein, the language recognition module is trained using a first sample audio; the obtaining step of the first sample audio includes: obtaining a plurality of single-language audios, and finding target adjacent audio frames belonging to different characters in each of the single-language audios; for each of the single-language audios, determining at least one cutting point of the single-language audio based on the target adjacent audio frames of the single-language audio, and cutting the single-language audio according to the at least one cutting point to obtain a plurality of audio segments of the single-language audio, wherein at least one audio segment of the single-language audio contains at least two characters; selecting at least one audio segment of two or more of the single-language audios for splicing to obtain the first sample audio.

2. The method according to claim 1, characterized in that, The first sample audio is labeled with the language information corresponding to each audio frame in the first sample audio.

3. The method according to claim 1, wherein The step of extracting the first audio features corresponding to each of the languages in the target audio is executed by an encoder module corresponding to each of the languages of the speech recognition model; the step of obtaining the speech recognition result of the target audio based on the first audio features corresponding to each of the languages and the language information in the target audio is obtained by using a decoding module of the speech recognition model.

4. The method according to claim 3, wherein Each encoder module corresponding to the languages is trained using a first single-language sample audio corresponding to the language, and the decoding module is trained using a second sample audio, where the second sample audio includes at least one of a second single-language sample audio and a multi-language sample audio; And / or, the language information in the target audio is obtained by using the language recognition module of the speech recognition model, and the training step of the decoding module includes: Using the trained encoder module to extract sample audio features corresponding to each of the languages in the second sample audio, and using the trained language recognition module to obtain the language information in the second sample audio, where the language recognition module is trained using the first sample audio and the trained encoder module; Based on the sample audio features corresponding to each of the languages and the language information in the second sample audio, obtaining a sample speech recognition result of the second sample audio; Using the sample speech recognition result to adjust the parameters of the decoding module.

5. The method according to claim 1, characterized in that, The language information in the target audio includes the language information corresponding to each audio frame in the target audio; the obtaining of the speech recognition result of the target audio based on the first audio features corresponding to each of the languages and the language information in the target audio includes: Based on the first audio features corresponding to each of the languages and the language information corresponding to each of the audio frames, obtain the language audio features corresponding to multiple encoding stages, and obtain the decoded text for each of the encoding stages based on the language audio features corresponding to each of the encoding stages. Among them, the multiple encoding stages are executed sequentially, and the decoded texts of different encoding stages represent the text information of different audio parts in the target audio. The language audio features corresponding to the encoding stage include the second audio features corresponding to each of the languages in the encoding stage; Based on the decoded texts of each of the encoding stages, obtain the speech recognition result of the target audio.

6. The method according to claim 5, wherein The obtaining the language audio features corresponding to multiple encoding stages based on the first audio features corresponding to each of the languages and the language information corresponding to each of the audio frames includes: Take the currently executed encoding stage as the current encoding stage; For each language included in the target audio, based on the first audio features corresponding to the language, the language information of each of the audio frames corresponding to the language, and the decoded text of the reference encoding stage, obtain the second audio features corresponding to the language in the current encoding stage; where the reference encoding stage is the encoding stage executed before the current encoding stage; And / or, the obtaining the decoded text for each of the encoding stages based on the language audio features corresponding to each of the encoding stages includes: Take the currently executed encoding stage as the current encoding stage; Fuse the second audio features corresponding to each of the languages in the current encoding stage to obtain the second fused audio feature corresponding to the current encoding stage; Decode the second fused audio feature corresponding to the current encoding stage to obtain the decoded text of the current encoding stage.

7. The method according to claim 6, characterized in that, The obtaining the second audio features corresponding to the language in the current encoding stage based on the first audio features corresponding to the language, the language information of each of the audio frames corresponding to the language, and the decoded text of the reference encoding stage includes: For each of the audio frames corresponding to the language, based on the first audio features corresponding to the audio frame, the language information of the audio frame, and the decoded text of the reference encoding stage, obtain the attention weight of the audio frame; Based on the first audio features corresponding to each of the audio frames corresponding to the language and the attention weights, obtain the second audio features corresponding to the language in the current encoding stage.

8. A voice recognition device, characterized in that, The apparatus includes: An acquisition module, configured to acquire a target audio to be recognized, where the target audio is a mixed audio including at least two languages; A feature extraction module, configured to extract the first audio features corresponding to each of the languages in the target audio; A speech recognition module, configured to obtain the speech recognition result of the target audio based on the first audio features corresponding to each of the languages and the language information in the target audio; the language information is obtained by processing the first audio fusion feature using the language recognition module of the speech recognition model after fusing the first audio features corresponding to each of the languages; Among them, the language identification module is trained using the first sample audio; the obtaining steps of the first sample audio include: obtaining a plurality of single-language audios, and finding target adjacent audio frames belonging to different characters in each of the single-language audios; for each of the single-language audios, based on the target adjacent audio frames of the single-language audio, determining at least one cut point of the single-language audio, and cutting the single-language audio according to the at least one cut point to obtain a plurality of audio segments of the single-language audio, wherein at least one audio segment of the single-language audio contains at least two characters; selecting at least one of the audio segments of two or more of the single-language audios for splicing to obtain the first sample audio.

9. An electronic device, characterized in that, Comprising a memory and a processor coupled to each other, The memory stores program instructions; The processor is configured to execute the program instructions stored in the memory to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program instructions, and the program instructions can be executed to implement the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Cross-language speech recognition method and device

    CN110349564A

  • Speech recognition method and device and computer readable storage medium

    CN114283786A