A Hybrid Bilingual Speech Recognition Method and System

By adopting the Encoder-Decoder framework and BPE word segmentation method in hybrid bilingual speech recognition, combining the Attention mechanism and data augmentation technology, problems such as high system complexity and language recognition dependence in the existing technology are solved, and a more efficient and robust hybrid bilingual speech recognition effect is achieved.

CN114267333BActive Publication Date: 2025-06-03NAT COMP NETWORK & INFORMATION SECURITY MANAGEMENT CENT GUANGDONG BRANCH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111509949.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-06-03
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

The existing hybrid bilingual speech recognition schemes have the problems of high overall complexity of the system, large calculation volume, recognition performance depends on language recognition results, speech length and quality affect the recognition effect, and non-subject language information is submerged.

Method used

The multilingual joint modeling method based on the Encoder-Decoder framework is adopted to construct a multilingual unified modeling unit through BPE word segmentation, without distinguishing languages, and using the Attention mechanism and data augmentation technology to achieve end-to-end hybrid bilingual speech recognition.

Benefits of technology

Reliance on language recognition is avoided, system complexity and calculation amount is reduced, recognition performance is improved, and recognition effect is significantly improved in the processing of non-subject language information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114267333B_ABST
    Figure CN114267333B_ABST
Patent Text Reader

Abstract

The present invention discloses a hybrid bilingual speech recognition method and system. The method includes the following steps: a data processing step, including: performing BPE shared dictionary making, data augmentation, and feature extraction operations on a certain amount of target bilingual audio data and target bilingual text corpus to provide effective data input for the training of the backend network; an Encoder-Decoder training step, including: training a speech recognizer using a Transformer structure on the effective data obtained in the data processing step. The present invention relates to the technical field of bilingual hybrid continuous speech recognition. According to the input monolingual speech data, bilingual mixed speech data, or bilingual mixed-up speech data of the target language, the content information of the speech is automatically transcribed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for hybrid bilingual speech recognition, belonging to the technical field of speech recognition. Background Art

[0002] With the development and breakthrough of deep learning technology, especially end-to-end technology, in the field of speech recognition, continuous speech recognition technology has been widely applied in fields such as education, medical care, smart cities, entertainment, and military, and the actual effects of its applications in various fields have been generally recognized by the industry. However, in actual business operations, the speech data received at the front end is not all in a single language, and it may be mixed with two or more languages. For example, in the Guangdong region, the speech data received at the front end may be a mixture of Mandarin, Cantonese, or English. Using only a single-language engine for continuous speech recognition of such multi-language mixed speech cannot achieve the requirements of actual use. Therefore, the research on multi-language mixed continuous speech recognition technology has become a hot topic in the field of speech recognition.

[0003] Hybrid bilingual speech recognition is the most common and important one in the scenario of hybrid multi-language recognition, and it is also one of the important challenges faced by current speech recognition technology. The hybrid bilingual speech data stream usually includes two situations. One is that the target speech stream contains both bilingual data of language A and language B and their proportions are quite equal; the other is that the speech stream of target language A contains a small number of words and phrases of language B. For these two types of bilingual mixed speech streams, there are currently two mainstream solutions. One is the method of first classifying the language of the speech and then performing speech recognition. The language is segmented through a short-time language recognition model, and then the audio after language classification is sent to the corresponding speech recognition system for recognition; the other is to add the words of language B to the speech recognition system of target language A, and the pronunciation of the words of language B is simulated using the phoneme system of language A.

[0004] The current hybrid bilingual speech recognition solutions have achieved certain effects in bilingual speech recognition, but there are still some problems. For the speech stream recognition solution for bilingual mixing, the system needs to integrate language classification and multiple speech recognition modules. The large number of models increases the computational complexity of the system, and the overall complexity of the system is relatively high; moreover, the speech recognition module depends on the language classification result, and incorrect language classification will lead to completely incorrect recognition results. In actual scenarios, both the speech length and quality will affect the language classification effect, and thus affect the overall recognition effect of the system. For the speech stream recognition solution for bilingual mixing, on the one hand, using the phoneme system of language A cannot well represent the words of language B, resulting in poor recognition effects; on the other hand, since the speech recognition system of language A is mainly trained through the speech of language A, and the proportion of language B speech is relatively small, the speech recognition system tends to recognize the words of language A, resulting in the correct words of language B not being correctly recognized.

[0005] Existing technical solutions for mixed bilingual speech recognition are mainly divided into speech recognition solutions based on mixed bilingual speech streams and speech recognition solutions based on mixed bilingual jumbled speech streams according to the mixing method of bilingual speech streams.

[0006] For mixed bilingual speech streams, existing technical solutions usually construct a mixed language speech recognition system by combining language identification and speech recognition. A language identifier is constructed at the front end of the system to first determine the language category of the input data, and then the audio features are sent to the speech recognizer of the corresponding language to obtain the speech recognition result. This method strongly depends on the effect of the language identifier. To overcome the errors introduced by language identification, a hybrid multi-language speech recognition scheme using multiple monolingual speech recognizers in parallel is proposed. Each speech recognizer in the system is independent and trained separately. During recognition, the audio features are sent to all the speech recognizers in the hybrid system, and the one with the highest likelihood probability among all the speech recognition results is selected as the final recognition result, which has better robustness and scalability.

[0007] For mixed bilingual jumbled speech streams, there are mainly two existing technical solutions: One solution is to automatically segment the jumbled speech into small segments containing complete word boundaries according to word segments. The language identifier determines the language category of each small segment, and then the adjacent segments of the same language are re-integrated and sent to the corresponding recognizer for recognition. Finally, the recognition results are spliced to obtain the final result. Another solution is to add words of non-main languages to the speech recognition system of the main language. The pronunciation of words of non-main languages is simulated using the phoneme system of the main language, and then the mixed bilingual speech recognition system is trained with the monolingual speech recognition scheme of the main language.

[0008] Combining the two solutions of mixed bilingual and mixed bilingual speech recognition can solve continuous mixed bilingual speech recognition without stopping when switching languages. With the development of neural networks, there have currently emerged some good-performing language identification and speech recognition methods. Selecting appropriate language identification and speech recognition methods according to different data and application scenarios to construct a reasonable continuous mixed bilingual speech recognition scheme can achieve the purpose of recognizing languages and speech, and has achieved considerable results to a certain extent. Summary of the Invention

[0009] The above scheme can solve the problem of mixed bilingual speech recognition to a certain extent, but there are still some shortcomings. The mixed language bilingual recognition system is constructed by cascading language recognition and multiple monolingual speech recognizers. The overall complexity of this type of system is high and the amount of calculation is large; and the overall recognition performance of the system depends on the results of language recognition. Errors in language recognition will directly lead to complete errors in the speech recognition results. In actual scenarios, the length and quality of speech will affect the language recognition effect, and then affect the recognition effect of the system. Although the use of multiple speech recognizers in parallel can solve the errors caused by language recognition, the amount of calculation of the system is large. The parallel recognition system based on language segmentation also has the problem of large overall system calculation. Another method is to use the phoneme system of the main language to simulate another mixed language, and then combine the main language with the unified modeling scheme. If the pronunciation characteristics of the two mixed languages ​​are quite different, the phoneme system of the main language cannot well represent the words of the other mixed language, resulting in poor recognition effect; at the same time, the main language data accounts for a higher proportion of the data in the mixed language speech recognition model training data, and the non-main language data accounts for a lower proportion. If forced division is used, it will be very unfavorable for the non-main language, causing the speech recognition model to be more inclined to recognize the words of the main language, resulting in the inability to effectively recognize the words of the non-main language.

[0010] In general, the existing hybrid bilingual speech recognition solutions have the following shortcomings: First, the overall system recognition effect depends on the accuracy of language recognition due to joint language recognition; second, the imbalance of mixed language data; and third, the scarcity of mixed bilingual data. This solution aims to address some of the problems existing in the existing hybrid bilingual speech recognition solutions and proposes an end-to-end continuous speech recognition unified modeling solution based on BPE (byte-pair encoding) word segmentation. The hybrid bilingual unified modeling unit is constructed using BPE word segmentation, and the end-to-end hybrid bilingual speech recognition model is directly trained, which gets rid of the dependence on language recognition and avoids the OOV (Out Of Vocabulary) problem of whole word modeling and the loss of semantic information and difficulty in training convergence caused by character modeling.

[0011] The object of the present invention is to overcome the technical defects existing in the prior art, solve the above technical problems, and propose a hybrid bilingual speech recognition method and system. A multi-language joint modeling method based on the Encoder-Decoder framework is adopted, and a multi-language unified modeling unit is constructed based on BPE word segmentation. Without distinguishing languages, the output words are automatically determined through data driving and the context where the words are located. Based on data driving, the network can predict the text content based on context information. The multi-language data joint training method is adopted, without separately training multiple models, and at the same time, the reuse of multi-language training data can be realized. The model can learn most of the feature extraction capabilities from the speech data of languages with rich corpora and apply them to low-resource languages to improve the recognition effect. To avoid the problem that the information of non-main languages in mixed languages is submerged by a large amount of main language data, resulting in poor recognition effect of non-main languages, this solution adds an Attention mechanism for mixed non-main languages at the Encoder end, so that the Decoder can predict the input speech based on both context information and the pronunciation information of non-main languages learned by the Attention module during the decoding process. To solve the problem of scarce mixed bilingual speech data, this solution also adopts speech augmentation methods such as speech synthesis, adding noise, and speed perturbation.

[0012] The present invention specifically adopts the following technical solutions: A hybrid bilingual speech recognition method, comprising the following steps:

[0013] A data processing step, including: performing BPE shared dictionary making, data augmentation, and feature extraction operations on a certain amount of target bilingual audio data and target bilingual text corpora to provide effective data input for the training of the backend network;

[0014] An Encoder-Decoder training step, including: training a speech recognizer using a Transformer structure on the effective data obtained in the data processing step.

[0015] As a preferred embodiment, the BPE shared dictionary making specifically includes: constructing a shared word segmentation dictionary for the target bilingual by using the BPE byte pair encoding algorithm on the target bilingual text corpus.

[0016] As a preferred embodiment, the data augmentation specifically includes: performing entity translation, insertion, and replacement operations on the target bilingual text corpus to construct a bilingual mixed text corpus, and then using a synthesizer to synthesize the bilingual mixed text corpus with the target bilingual audio data to generate mixed bilingual speech data; performing noise addition and speed change operations on the mixed bilingual speech data for data augmentation to generate augmented data.

[0017] As a preferred embodiment, the feature extraction includes: combining the original target bilingual audio data, the target bilingual text corpus, the synthesized mixed bilingual speech data, and the augmented data after adding noise and changing speed as training data to extract features, generating the input audio feature sequence X = {x 1 , x 2 , …, x T}.

[0018] As a preferred embodiment, the training steps of the Encoder-Decoder specifically include:

[0019] Encoding step of the Encoder: Through the Self-Attention self-attention mechanism operation, the LayerNorm mean and variance operation, and the FeedForward positive feedback operation, convert the input audio feature sequence X = {x 1 , x 2 , …, x T} into a high-level feature representation sequence

[0020] Decoding step of the Decoder: Based on the previous output y u-1 , the previous hidden state and the previous context vector c u-1 , calculate the current hidden state Then calculate the output symbol through the Softmax logistic regression layer Based on the probability distribution of the previous predicted labels and the input feature sequence

[0021] Both the Attention attention mechanism layers in the Encoder and the Decoder have Residual residual connections, directly passing the information of the previous layer to the next layer; the Self-Attention self-attention mechanism layers in the Encoder and the Decoder both calculate the attention weights of all vectors in the previous layer, output the context vector, and thus establish the alignment relationship between the input sequence and the output sequence; different from the Encoder, the Decoder also includes an Encoder-Decoder Attention encoding-decoding attention mechanism layer, whose input information includes the output of the previous layer and the output result of the Encoder decoder.

[0022] The present invention also proposes a mixed bilingual speech recognition system, including:

[0023] The data processing module specifically performs: performing BPE shared dictionary creation, data augmentation, and feature extraction operations on a certain amount of target bilingual audio data and target bilingual text corpora to provide effective data input for the training of the backend network;

[0024] The Encoder-Decoder training module specifically performs: training a speech recognizer on the effective data obtained from the data processing step using a Transformer structure.

[0025] As a preferred embodiment, the BPE shared dictionary creation specifically includes: constructing a shared word segmentation dictionary for the target bilingual using the BPE byte pair encoding algorithm on the target bilingual text corpus.

[0026] As a preferred embodiment, the data augmentation specifically includes: performing entity translation, insertion, and replacement operations on the target bilingual text corpus to construct a bilingual mixed text corpus, and then using a synthesizer to synthesize the bilingual mixed text corpus with the target bilingual audio data to generate mixed bilingual speech data; performing noise addition and speed change operations on the mixed bilingual speech data for data augmentation to generate augmented data.

[0027] As a preferred embodiment, the feature extraction includes: combining the original target bilingual audio data and target bilingual text corpus, the synthesized mixed bilingual speech data, and the augmented data after noise addition and speed change as training data to extract features, generating an input audio feature sequence X = {x 1 ,x 2 ,…,x T}.

[0028] As a preferred embodiment, the Encoder-Decoder training module specifically includes:

[0029] The Encoder encoder encoding step: converting the input audio feature sequence X = {x 1 ,x 2 ,…,x T} into a high-level feature representation sequence through Self-Attention self-attention mechanism operations, LayerNorm mean and variance operations, and FeedForward positive feedback operations

[0030] The Decoder decoder decoding step: based on the previous output y u-1 , the previous hidden state , and the previous context vector c u-1 , calculating the current hidden state and then calculating the output symbol through a Softmax logistic regression layer Based on the probability distributions of previous predicted tags and input feature sequences

[0031] Both the Attention mechanism layers in the Encoder and the Decoder have Residual connections, directly passing the information of the previous layer to the next layer; the Self-Attention mechanism layers in the Encoder and the Decoder both calculate the attention weights of all vectors in the previous layer, output context vectors, and thus establish the alignment relationship between the input sequence and the output sequence; different from the Encoder, the Decoder also includes an Encoder-Decoder Attention mechanism layer, and its input information includes the output of the previous layer and the output result of the Encoder.

[0032] The beneficial effects achieved by the present invention: The present invention uses the BPE tokenization method to construct a unified monolingual model for mixed languages, avoiding the OOV problem of whole-word modeling and the loss of semantic information and the problem of difficult training convergence caused by character modeling; adopting an end-to-end unified modeling scheme for mixed bilinguals, enabling speech recognition of mixed bilinguals without constructing a language identification module, and making training more convenient and fast; using various methods such as entity translation, audio splicing, speech synthesis, speed change, and noise addition to augment the training data of the mixed bilingual model and improve the generalization ability of the model. Description of the Drawings

[0033] Figure 1 is a flowchart of a method for mixed bilingual speech recognition according to the present invention. Detailed Embodiments

[0034] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.

[0035] Embodiment 1: As Figure 1 shown, the present invention proposes a method for mixed bilingual speech recognition, including the following steps:

[0036] The data processing step includes: performing BPE shared dictionary making, data augmentation, and feature extraction operations on a certain amount of target bilingual audio data and target bilingual text corpus to provide effective data input for the training of the backend network;

[0037] The Encoder-Decoder training step includes: training a speech recognizer using the Transformer structure on the effective data obtained in the data processing step.

[0038] As a preferred embodiment, the production of the BPE shared dictionary specifically includes: constructing a shared word segmentation dictionary for the target bilingual by using the BPE byte pair encoding algorithm for the target bilingual text corpus.

[0039] As a preferred embodiment, the data augmentation specifically includes: constructing a bilingual mixed text corpus by performing entity translation, insertion, and replacement operations on the target bilingual text corpus, and then using a synthesizer to synthesize the bilingual mixed text corpus with the target bilingual audio data into mixed bilingual speech data; performing noise addition and speed change operations on the mixed bilingual speech data for data augmentation to generate augmented data.

[0040] As a preferred embodiment, the feature extraction includes: combining the original target bilingual audio data, the target bilingual text corpus, the synthesized mixed bilingual speech data, and the augmented data after noise addition and speed change as training data to extract features, generating an input audio feature sequence X = {x 1 , x 2 , …, x T}.

[0041] As a preferred embodiment, the training steps of the Encoder-Decoder specifically include:

[0042] The encoding step of the Encoder encoder: converting the input audio feature sequence X = {x 1 , x 2 , …, x T} into a high-level feature representation sequence through Self-Attention self-attention mechanism operation, LayerNorm mean and variance operation, and FeedForward positive feedback operation

[0043] The decoding step of the Decoder decoder: calculating the current hidden state based on the previous output y u-1 , the previous hidden state and the previous context vector c u-1 , and then calculating the output symbol through the Softmax logistic regression layer Based on the probability distribution of the previous predicted labels and the input feature sequence

[0044] ​Both the Attention mechanism layers in the Encoder and the Decoder have Residual connections, which directly pass the information of the previous layer to the next layer; the Self-Attention mechanism layers in the Encoder and the Decoder both calculate the attention weights of all vectors in the previous layer, output context vectors, and then establish the alignment relationship between the input sequence and the output sequence; different from the Encoder, the Decoder also includes an Encoder-Decoder Attention mechanism layer, and its input information includes the output of the previous layer and the output result of the Encoder.

[0045] Embodiment 2: The present invention also proposes a hybrid bilingual speech recognition system, including:

[0046] A data processing module, which specifically executes: performing BPE shared dictionary making, data augmentation, and feature extraction operations on a certain amount of target bilingual audio data and target bilingual text corpus to provide effective data input for the training of the backend network;

[0047] An Encoder-Decoder training module, which specifically executes: training a speech recognizer on the effective data obtained in the data processing step using a Transformer structure.

[0048] As a preferred embodiment, the BPE shared dictionary making specifically includes: constructing a shared word segmentation dictionary for the target bilingual by using the BPE byte pair encoding algorithm on the target bilingual text corpus.

[0049] As a preferred embodiment, the data augmentation specifically includes: performing entity translation, insertion, and replacement operations on the target bilingual text corpus to construct a bilingual mixed text corpus, and then using a synthesizer to synthesize the bilingual mixed text corpus with the target bilingual audio data to generate mixed bilingual speech data; performing noise addition and speed change operations on the mixed bilingual speech data for data augmentation to generate augmented data.

[0050] As a preferred embodiment, the feature extraction includes: combining the original target bilingual audio data and target bilingual text corpus, the synthesized mixed bilingual speech data, and the augmented data after noise addition and speed change together as training data to extract features, and generating an input audio feature sequence X = {x 1 , x 2 , …, x T}.

[0051] As a preferred embodiment, the Encoder-Decoder training module specifically includes:

[0052] Encoder encoding steps: Through Self-Attention operation, LayerNorm operation for calculating mean and variance, and FeedForward operation, the input audio feature sequence X = {x 1 , x 2 , …, x T} is converted into a high-level feature representation sequence

[0053] Decoder decoding steps: Based on the previous output y u-1 , the previous hidden state , and the previous context vector c u-1 , calculate the current hidden state Then calculate the output symbol through the Softmax logistic regression layer Based on the probability distributions of the previous predicted labels and the input feature sequence

[0054] Both the Attention mechanism layers in the Encoder and the Decoder have Residual connections, directly passing the information of the previous layer to the next layer; the Self-Attention mechanism layers in the Encoder and the Decoder both calculate the attention weights of all vectors in the previous layer, output the context vector, and thus establish the alignment relationship between the input sequence and the output sequence; different from the Encoder, the Decoder also includes an Encoder-Decoder Attention mechanism layer, whose input information includes the output of the previous layer and the output result of the Encoder.

[0055] The Encoder-Decoder training module mainly adopts the encoding and decoding structure of Transformer, and uses the highly parallel Self-Attention mechanism and MultiHead-Attention mechanism to implement the essential encoding, decoding, and alignment modules in speech recognition. Since there is no order limit for the alignment relationship of the Attention model, to avoid excessive randomness of the corresponding relationship, the CTC method is introduced in the Decoder part, and the CTC loss function is used to solve the problem that it is difficult to correspond one by one between the input sequence and the output sequence, further improving the recognition effect of the Encoder-Decoder model.

[0056] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0057] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks. These computer program instructions can also be stored in a computer-readable memory capable of guiding the computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks. These computer program instructions can also be loaded onto the computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. A hybrid bilingual speech recognition method, characterized in that, it includes the following steps: A data processing step, including: performing BPE shared dictionary making, data augmentation, and feature extraction operations on a certain amount of target bilingual audio data and target bilingual text corpus to provide effective data input for the training of the backend network; An Encoder-Decoder training step, including: training a speech recognizer using a Transformer structure on the effective data obtained in the data processing step; The Encoder-Decoder training step specifically includes: Encoder Encoding Steps: Through Self-Attention mechanism operation, LayerNorm mean and variance operation, and FeedForward positive feedback operation, the input audio feature sequence X = {x 1 , x 2 , …, x T} is converted into a high-level feature representation sequence Decoder decoding steps: Based on the previous output y u-1 , the previous hidden state and the previous context vector c u-1 , calculate the current hidden state Then calculate the output symbol through the Softmax logistic regression layer Based on the probability distribution of the previous predicted label and the input feature sequence Both the Attention attention mechanism layer in the Encoder encoder and the Decoder decoder have Residual residual connections, directly passing the information of the previous layer to the next layer; the Self-Attention self-attention mechanism layer in the Encoder encoder and the Decoder decoder both calculates the attention weights of all vectors in the previous layer and outputs a context vector, thereby establishing an alignment relationship between the input sequence and the output sequence; different from the Encoder encoder, the Decoder decoder also includes an Encoder-Decoder Attention encoding-decoding attention mechanism layer, and its input information includes the output of the previous layer and the output result of the Encoder decoder.

2. The hybrid bilingual speech recognition method according to claim 1, characterized in that, The BPE shared dictionary making specifically includes: constructing a shared word segmentation dictionary for the target bilingual by using the BPE byte pair encoding algorithm on the target bilingual text corpus.

3. The hybrid bilingual speech recognition method according to claim 1, characterized in that, The data augmentation specifically includes: performing entity translation, insertion, and replacement operations on the target bilingual text corpus to construct a bilingual mixed text corpus, and then using a synthesizer to synthesize the bilingual mixed text corpus with the target bilingual audio data to generate mixed bilingual speech data; performing noise addition and speed change operations on the mixed bilingual speech data for data augmentation to generate augmented data.

4. The hybrid bilingual speech recognition method according to claim 1, characterized in that, Merge the original target bilingual audio data, the target bilingual text corpus, the synthesized hybrid bilingual speech data, and the augmented data after adding noise and changing speed as training data to extract features, and generate the input audio feature sequence X = {x 1 , x 2 , …, x T}.

5. A hybrid bilingual speech recognition system, characterized in that, it includes: A data processing module, specifically performing: performing BPE shared dictionary making, data augmentation, and feature extraction operations on a certain amount of target bilingual audio data and target bilingual text corpus to provide effective data input for the training of the backend network; An Encoder-Decoder training module, specifically performing: training a speech recognizer using a Transformer structure on the effective data obtained by the data processing module; The Encoder-Decoder training module specifically includes: Encoder Encoding Steps: Through Self-Attention mechanism operations, LayerNorm mean and variance operations, and FeedForward positive feedback operations, the input audio feature sequence X = {x 1 , x 2 , …, x T} is converted into a high-level feature representation sequence Decoder decoding steps: Based on the previous output y u-1 , the previous hidden state and the previous context vector c u-1 , calculate the current hidden state Then calculate the output symbol through the Softmax logistic regression layer Based on the probability distribution of the previous predicted labels and the input feature sequence Both the Attention mechanism layers in the Encoder and the Decoder have Residual connections, which directly pass the information of the previous layer to the next layer; the Self-Attention mechanism layers in the Encoder and the Decoder both calculate the attention weights of all vectors in the previous layer, output context vectors, and thus establish the alignment relationship between the input sequence and the output sequence; different from the Encoder, the Decoder also includes an Encoder-Decoder Attention mechanism layer, and its input information includes the output of the previous layer and the output result of the Encoder.

6. The hybrid bilingual speech recognition system according to claim 5, wherein, the production of the BPE shared dictionary specifically includes: constructing a shared word segmentation dictionary for the target bilingual by using the BPE byte pair encoding algorithm for the target bilingual text corpus.

7. The hybrid bilingual speech recognition system according to claim 5, wherein, the data augmentation specifically includes: constructing a bilingual mixed text corpus by performing entity translation, insertion, and replacement operations on the target bilingual text corpus, and then using a synthesizer to synthesize the bilingual mixed text corpus and the target bilingual audio data into mixed bilingual speech data; performing noise addition and speed change operations on the mixed bilingual speech data for data augmentation to generate augmented data.

8. The hybrid bilingual speech recognition system according to claim 5, wherein, Merge the original target bilingual audio data, the target bilingual text corpus, the synthesized mixed bilingual speech data, and the augmented data after adding noise and changing speed together as training data to extract features, and generate the input audio feature sequence X = {x 1 , x 2 , …, x T}.

Citation Information

Patent Citations

  • Chinese and English hybrid speech recognition model training method and device

    CN111816169A

  • Tibetan-Chinese translation method and device

    CN112084794A