Mixed-Language Speech Recognition Method, Apparatus, System, and Storage Medium

By recognizing speech information to be recognized and entering the trained transcription model for processing, the problem of low accuracy of speech recognition in multilingual mixed scenarios is solved, and efficient recognition and translating of mixed Mongolian and Chinese voices is achieved, improving user experience.

CN115394287BActive Publication Date: 2025-06-13IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210892864.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-27
Publication Date
2025-06-13
Estimated Expiration
2042-07-27

AI Technical Summary

Technical Problem

Existing speech recognition technology is difficult to accurately recognize in multilingual mixed scenarios, especially in mixed Mongolian and Chinese, where multilingual mixed speech cannot be effectively recognized, resulting in low accuracy of output results.

Method used

A mixed language speech recognition method is adopted to identify speech information, determine its language information by determining the target language, and when determining the target language, the speech information is input into the trained transcription model for the transcription processing. The transcribed model adopts an encoder-decoder framework, and improves the recognition efficiency and accuracy of the model through random mask processing and multi-task training.

Benefits of technology

It improves the accuracy of speech recognition in multilingual voice mixed scenarios, can effectively recognize and transliterate mixed voices in Mongolian and Chinese, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115394287B_ABST
    Figure CN115394287B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, system and storage medium for mixed-language speech recognition. Among them, the method for mixed-language speech recognition includes the following steps: obtaining speech information to be recognized; performing language recognition on the speech information to be recognized to determine the language information of the speech information to be recognized; when the language information includes a target language, inputting the speech information to be recognized into a trained transcription model to convert the speech information to be recognized into text information, where the target language includes a first language and a second language, and the text information includes mixed-language text information corresponding to the first language and the second language. Through the method of the present application, the accuracy of the obtained text information is higher, and it can output a recognition result of mixed multi-language speech, improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method, device, system and storage medium for mixed-language speech recognition. Background Art

[0002] With the development and breakthrough of deep learning technology, especially in the field of speech recognition, speech recognition technology has been widely applied in fields such as entertainment, education, smart cities, healthcare, and military, and the actual effects of its applications in various fields have been recognized by the industry. However, in actual speech recognition, the speech data obtained at the front end is not completely in a single language. Sometimes it may be mixed with two or more languages, such as the mixture of Mongolian and Chinese. Currently, speech recognition technology usually models a single language. In a complex scenario of multi-language mixture, the input speech stream is preprocessed at the front end and segmented into clauses, then the language confidence of the clauses is judged, and then the speech recognition model corresponding to the language with the highest confidence is used to output the recognition result. Finally, the recognition results of the clauses are spliced as the final result of the whole sentence. However, in a complex scenario of language mixture, the accuracy of the output result is not high, and it is even impossible to recognize the speech of multi-language mixture.

[0003] Therefore, improvements are needed to solve at least one of the above problems. Summary of the Invention

[0004] In view of the above problems, the present application provides a method for mixed-language speech recognition, and the method includes the following steps:

[0005] Obtain the speech information to be recognized;

[0006] Perform language recognition on the speech information to be recognized to determine the language information of the speech information to be recognized;

[0007] When the language information includes a target language, input the speech information to be recognized into a trained transcription model to convert the speech information to be recognized into text information. The target language includes a first language and / or a second language, and the text information includes the mixed-language text information corresponding to the first language and the second language. Among them,

[0008] The training process of the transcription model includes:

[0009] During the training process, perform random masking on the extracted acoustic features. Among them, the random masking process includes: randomly masking a predetermined number of time-domain features in the spectrogram corresponding to the acoustic features, and / or randomly masking a predetermined number of frequency-domain features in the spectrogram corresponding to the acoustic features.

[0010] In one embodiment, the transcription model is a speech recognition model based on an encoder-decoder framework. Among them, the encoder of the transcription model to be trained includes a feature extraction module, a convolutional network module, multiple first Transformer network structures, a feed-forward neural network layer, a deconvolution network module, a fully connected layer, and a normalization network module connected in sequence. The decoder of the transcription model to be trained includes a conversion module, a convolutional network module, multiple second Transformer network structures, a feed-forward neural network layer, and a normalization network module connected in sequence. Among them, the method pre-trains to obtain the trained transcription model through the following steps:

[0011] Obtain training data, where the training data includes speech information and text labels corresponding to the speech information;

[0012] Extract the acoustic features of the speech information in the current period of the input training data set through the feature extraction module, and perform the random masking process on the extracted acoustic features;

[0013] Extract the speech coding features of a fixed dimension in the acoustic features through the convolutional network module, the multiple first Transformer network structures, and the feed-forward neural network layer of the encoder;

[0014] Calculate the loss of the speech coding features of the fixed dimension based on the CTC loss function to obtain the first loss;

[0015] Use the deconvolution network module to upsample the time dimension corresponding to the speech coding features of the fixed dimension to be consistent with the time dimension of the input speech information in the current period, and use the fully connected layer and the normalization network module to process the output of the deconvolution network module to obtain the predicted phoneme labels;

[0016] Calculate the second loss of the phoneme sequence of the predicted phoneme labels relative to the phoneme sequence of the true label by using the cross-entropy loss function;

[0017] Obtain the text label corresponding to the speech information before the current period in the training data;

[0018] Input the text label corresponding to the speech information before the current period into the conversion module to be converted into a character embedding vector;

[0019] Input the character embedding vector into the convolutional network model of the decoder to extract the abstract text representation information;

[0020] Input the abstract text representation information into the multiple second Transformer network structures of the decoder to extract the high-dimensional abstract text representation information;

[0021] Feature - weighted fusion of the fixed - dimensional speech coding features output by the feed - forward neural network layer of the encoder and the high - dimensional abstract text representation information is performed through an attention mechanism to obtain fused features;

[0022] The fused features are input into the feed - forward neural network layer and the normalization network module for processing to obtain a predicted text sequence;

[0023] The cross - entropy loss function is used to calculate the third loss at the character level of the predicted text sequence;

[0024] The sentence - level loss function is used to calculate the fourth loss of the predicted text sequence;

[0025] The first loss, the second loss, the third loss, and the fourth loss are weighted and summed to obtain an overall loss;

[0026] The model parameters in the transcribing model to be trained are adjusted using the overall loss to obtain the trained transcribing model.

[0027] In one embodiment, the obtaining of training data includes:

[0028] Obtain the true text labels corresponding to the speech information in the training dataset;

[0029] Perform random text feature perturbation on the true text labels to obtain text labels corresponding to the speech information, where the random text feature perturbation includes: randomly selected positions of the true text labels are replaced with characters or phonemes of non - true labels at a predetermined ratio.

[0030] In one embodiment, the first language is a minority language, and the training data of the transcribing model includes the synthetic speech of the first language and the corresponding text information, the text information corresponding to the original speech of the target language, the spliced speech of the target language and the corresponding text information, and the augmented speech of the target language. The synthetic speech is obtained by synthesizing the phoneme sequence corresponding to the historical text of the first language and the voiceprint information of the historical speech of the first language. The spliced speech is obtained by splicing two randomly selected speeches in the training data. The augmented speech is obtained by adding background noise to the original speech.

[0031] In one embodiment, the language identification of the obtained speech information to be recognized to determine the language information of the speech information to be recognized includes:

[0032] Use the trained language recognition model to perform language recognition on the speech information to be recognized, so as to predict the scores of the target languages in the speech information to be recognized. Among them, the scores of the target languages include the first score of the first language and the second score of the second language;

[0033] Compare the first score with the first threshold, and compare the second score with the second threshold. When the first score is less than the first threshold and the second score is less than the second threshold, it is determined that the language information includes the target language.

[0034] In one embodiment, the transcription model is a speech recognition model based on an encoder-decoder framework. The speech information to be recognized includes speech segments in multiple time periods. Inputting the speech information to be recognized into the trained transcription model to convert the speech information to be recognized into text information includes:

[0035] Perform encoding and decoding processing on the speech segment in each time period to predict the predicted text label corresponding to the speech segment in each time period;

[0036] Merge the predicted text labels corresponding to the speech segments in all time periods in chronological order to obtain the predicted text label corresponding to the speech information to be recognized;

[0037] Obtain the text information corresponding to the speech information to be recognized according to the predicted text label corresponding to the speech information to be recognized.

[0038] In one embodiment, obtaining the speech information to be recognized includes:

[0039] Obtain the original speech information;

[0040] Segment the original speech information through voice activity endpoint detection and filter out the invalid speech in the original speech information to obtain the speech information to be recognized.

[0041] On the other hand, the present application also provides a hybrid language speech recognition device, and the device includes:

[0042] An acquisition module for acquiring speech information to be recognized;

[0043] A language recognition module for performing language recognition on the speech information to be recognized to determine the language information of the speech information to be recognized;

[0044] A transcription module, which is configured to input the to-be-recognized speech information into a trained transcription model when the language information includes a target language, so as to convert the to-be-recognized speech information into text information. The target language includes a first language and a second language, and the text information includes mixed language text information corresponding to the first language and the second language. Among them,

[0045] The training process of the transcription model includes:

[0046] During the training process, random masking processing is performed on the extracted acoustic features. Among them, the random masking processing includes: randomly occluding a predetermined number of time-domain features in the spectrogram corresponding to the acoustic features, and / or randomly occluding a predetermined number of frequency-domain features in the spectrogram corresponding to the acoustic features.

[0047] On the other hand, the present application also provides a mixed language speech recognition system, which includes a memory and a processor. A computer program is stored on the memory and run by the processor. When the computer program is run by the processor, the processor is made to execute the foregoing mixed language speech recognition method.

[0048] On the other hand, the present application provides a storage medium, on which a computer program is stored. When the computer program runs, it executes the foregoing mixed language speech recognition method.

[0049] To solve at least one of the foregoing technical problems, the present application provides a mixed language speech recognition method, device, system and storage medium. Through the mixed language speech recognition method of the present application, the to-be-recognized speech information is first subjected to language recognition to determine the language information of the to-be-recognized speech information, so as to screen the data of the speech information. When it is screened that the speech information includes a target language, the to-be-recognized speech information is input into a trained transcription model to convert the to-be-recognized speech information into text information, so that the data input into the transcription model more meets the model requirements, improves the recognition efficiency of the model, and further makes the obtained text information more accurate, can output a recognition result of mixed multi-language speech, and improves the user experience. Description of the Drawings

[0050] The following drawings of the present application are used as a part of the present application to understand the present application. The drawings show the embodiments of the present application and their descriptions, and are used to explain the device and principle of the present application. In the drawings,

[0051] Figure 1 A schematic flowchart showing a mixed language speech recognition method according to an embodiment of the present application is shown.

[0052] Figure 2Another schematic flowchart showing the mixed-language speech recognition method according to an embodiment of the present application.

[0053] Figure 3 A schematic flowchart showing the simulation of Mongolian monolingual corpus synthesis according to an embodiment of the present application.

[0054] Figure 4 A schematic block diagram showing a transcription model according to an embodiment of the present application.

[0055] Figure 5 A schematic diagram showing a random speech feature mask according to an embodiment of the present application.

[0056] Figure 6 A schematic diagram showing a random text feature perturbation according to an embodiment of the present application.

[0057] Figure 7 A schematic block diagram showing a mixed-language speech recognition device according to an embodiment of the present application.

[0058] Figure 8 A schematic block diagram showing a mixed-language speech recognition system according to an embodiment of the present application. Detailed implementation manners

[0059] In order to make the objectives, technical solutions, and advantages of the present application more apparent, exemplary embodiments according to the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein. Based on the embodiments of the present application described herein, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.

[0060] Based on at least one of the foregoing technical problems, as Figure 1 shown, the present application provides a mixed-language speech recognition method 100, which includes the following steps:

[0061] Step S110, obtaining the speech information to be recognized;

[0062] Step S120, performing language recognition on the speech information to be recognized to determine the language information of the speech information to be recognized;

[0063] Step S130, when the language information includes a target language, inputting the speech information to be recognized into a trained transcription model to convert the speech information to be recognized into text information, where the target language includes a first language and a second language, and the text information includes mixed-language text information corresponding to the first language and the second language.

[0064] Through the mixed-language speech recognition method of the present application, first perform language recognition on the speech information to be recognized to determine the language information of the speech information to be recognized, so as to screen the data of the speech information. When it is screened that the speech information includes the target language, input the speech information to be recognized into the trained transcription model to convert the speech information to be recognized into text information, so that the data input into the transcription model is more in line with the model requirements, improve the recognition efficiency of the model, and further make the accuracy of the obtained text information higher, and can output the recognition result of mixed-language speech (for example, the speech including multiple languages in a whole sentence and a clause), that is, the mixed-language text information, and improve the user experience.

[0065] The above-mentioned speech information to be recognized may refer to speech data including at least one language. In the embodiments of the present application, the speech information to be recognized may include speech data of multiple languages. For example, it may simultaneously include mixed speech of a first language and a second language. Among them, the first language may be one or more different minority languages such as Mongolian, French, Japanese, etc., and the second language may be Chinese. Or the first language may be a local dialect and the second language is Mandarin Chinese. Or the first language may be other ethnic languages and the second language is Chinese. Among them, in the present application, a minority language may refer to other languages except English and Chinese.

[0066] In the present application, the case of mixing Mongolian and Chinese is mainly taken as an example, but this is not intended to be restrictive. In addition to being applicable to the case of mixing Mongolian and Chinese, the present application can also be applicable to the case of mixing other languages and Chinese, or the case of mixing any two or more languages.

[0067] In step S110, the speech information to be recognized may be the obtained original speech information. The original speech information may be long speech stream or short speech stream. The long speech stream may refer to speech with a duration greater than or equal to a preset duration, while the short speech stream may refer to speech with a duration below the preset time. The preset duration may be reasonably set according to prior experience and will not be specifically limited here.

[0068] In some embodiments, the obtaining of the speech information to be recognized includes: obtaining the original speech information. Optionally, the original speech information may be a long speech stream; segment the original speech information through voice activity detection (VAD) and filter the invalid speech in the original speech information to obtain the speech information to be recognized. Through segmentation and filtering, the speech included in the speech information to be recognized is basically valid speech, thus avoiding the interference of invalid speech on the subsequent speech recognition effect, and further improving the accuracy of the speech recognition output result.

[0069] Among them, VAD can be used to separate the voice signal and non-voice signal (i.e., invalid voice, such as background noise like music and reverberation) in the original voice information. VAD can be displayed by any suitable method well-known to those skilled in the art. For example, by frame segmentation, judging the energy of a frame, the zero-crossing rate and other simple methods to determine whether it is a voice segment (which can also be called valid voice); 2. By detecting whether there is a fundamental period in a frame to determine whether it is a voice segment (which can also be called valid voice); 3. Training a model by the method of Deep Neural Networks (DNN) to classify whether it is a voice frame, and using DNN for voice classification, and then separating the voice segment (which can be called valid voice) and non-voice segment (i.e., invalid voice).

[0070] Whether to use VAD for segmentation and filtering can be reasonably selected according to the actual application scenario. For example, when most of the application scenarios involve phrase-stream voice (such as WeChat voice), VAD may not be used for segmentation and filtering, while when most of the application scenarios involve long-stream voice, VAD can be used for segmentation and filtering.

[0071] Alternatively, in some embodiments, it can be determined whether to use VAD for segmentation and filtering according to the duration of the original voice information. For example, when the duration is greater than or equal to a preset duration, VAD is applied, while when the duration is less than the preset duration, VAD is not applied. By setting it so flexibly, the amount of data processing can be reduced while ensuring the subsequent voice recognition effect.

[0072] Furthermore, in step S120, the language information of the voice information to be recognized can be determined based on any suitable method well-known to those skilled in the art. For example, it can be recognized and obtained based on a trained language recognition model. Or, in some embodiments, the language recognition of the obtained voice information to be recognized to determine the language information of the voice information to be recognized includes: performing language recognition on the voice information to be recognized by a trained language recognition model to predict the scores of the target languages in the voice information to be recognized, where the scores of the target languages include the first score of the first language and the second score of the second language; comparing the first score with the first threshold, and comparing the second score with the second threshold. When the first score is less than the first threshold and the second score is less than the second threshold, it is determined that the language information includes the target language. Among them, the first threshold and the second threshold can be reasonably set according to actual needs, for example, they can be 60 points, 70 points, 80 points, or 90 points, etc.

[0073] Language recognition can be used to determine the language information of the speech information to be recognized, so as to screen the data of the speech information. When it is screened that the speech information includes the target language, the speech information to be recognized is input into the trained transcription model to convert the speech information to be recognized into text information, so that the data input into the transcription model is more in line with the model requirements, improving the recognition efficiency of the model, and further making the accuracy of the obtained text information higher. It can output the recognition results of multi-language speech mixing, enhancing the user experience.

[0074] In some embodiments, the trained language recognition model is used to perform language recognition on the speech information to be recognized to predict the score of the target language in the speech information to be recognized, which may include: the trained language recognition model can be used to extract the language representation features of each speech segment in the speech information to be recognized, and compare the similarity between the language representation features and the language representation features of the target language to obtain the similarity scores of the languages of each speech segment and the target language (i.e., the scores of the target language, such as the first score and the second score).

[0075] The language representation features of each speech segment can be determined according to the acoustic features of the speech information to be recognized. For example, the bottleneck features of the speech information to be recognized can be extracted as its acoustic features, and then the acoustic features are mapped through a series of orthogonal projection spaces to obtain low-dimensional acoustic features as the language representation features of the speech information to be recognized, and then the language representation features of each speech segment are extracted from the language representation features of the speech information to be recognized. Or it can also be determined based on other suitable methods for the language representation features of each speech segment.

[0076] Optionally, the speech information to be recognized includes multiple speech segments, and the language information of each speech segment can be recognized in sequence to determine the language information of each speech segment, and the subsequent recognition can be the speech recognition of each speech segment in sequence. Some details of the specific recognition process will be described later.

[0077] In this application, the speech segments can be divided in any suitable way, and speech segments of any duration. For example, each speech frame in the speech information to be recognized can be used as a speech segment.

[0078] Through the above language recognition process, language information can be recognized, so as to determine whether the speech information to be recognized includes the target language. When the speech information to be recognized includes the target language, for example, some speech segments of the speech information to be recognized include the first language and the second language, some speech segments may also include the first language, and some speech segments may also include the second language, then these speech segments including the target language can be input into the trained transcription model for recognition, so as to convert the speech information to be recognized into text information. Further, in step S130, when the language information includes the target language, the speech information to be recognized is input into the trained transcription model to convert the speech information to be recognized into text information. The target language includes the first language and the second language, and the text information includes the text information corresponding to the first language and the text information corresponding to the second language. Through the trained transcription model of the present application, the recognition of speech in mixed languages and the writing of text can be realized, and the recognition effect is good and the accuracy is high.

[0079] In some embodiments, as Figure 4 shown, the transcription model of the present application is a speech recognition model based on the encoder-decoder framework. The speech information to be recognized includes speech segments in multiple time periods. The inputting the speech information to be recognized into the trained transcription model to convert the speech information to be recognized into text information includes: performing encoding and decoding processing on the speech segments in each time period to predict the predicted text labels corresponding to the speech segments in each time period; merging the predicted text labels corresponding to the speech segments in all time periods in chronological order to obtain the predicted text labels corresponding to the speech information to be recognized; and obtaining the text information corresponding to the speech information to be recognized according to the predicted text labels corresponding to the speech information to be recognized.

[0080] Among them, the duration of each time period can be reasonably set according to actual needs. For example, it can be 1s, 2s, 3s, 5s, 10s, etc.

[0081] The speech recognition model based on the encoder-decoder framework can be a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory network (LSTM), a gated recurrent unit (GRU), a neural network structure embedded with Attention, and so on.

[0082] In some embodiments, encoding and decoding processing is performed on the speech segments of each time period to predict the predicted text labels corresponding to the speech segments of each time period, including: encoding the speech segment of the current time period through an encoder (i.e., the encoder part in the transcription model) to obtain speech encoding features of a fixed dimension; decoding the speech encoding features of the fixed dimension and the historical predicted text labels obtained between the current time periods through a decoder (i.e., the decoder part in the transcription model) to predict the predicted text labels corresponding to the speech segments of the current time period. By traversing the speech segments of each time period, the speech segments corresponding to the target language are sequentially recognized to predict text labels, and the text labels can correspond to character sequences or phoneme sequences.

[0083] In some embodiments, the decoder decodes the speech encoding features of the fixed dimension and the historical predicted text labels obtained between the current time periods to predict the predicted text labels corresponding to the speech segments of the current time period, including: the decoder extracts high-dimensional abstract text representation information from the historical predicted text labels that have been predicted and obtained before the current time period. For example, if the current time period is the t-th time period, the historical predicted text labels can be the predicted text labels recognized by the trained transcription model at the (t - 1)-th time period and before; the decoder performs attention mechanism-based fusion processing on the speech encoding features of the fixed dimension and the high-dimensional abstract text representation information to obtain a fusion feature; the decoder processes the fusion feature to predict the predicted text labels corresponding to the speech segments of the current time period.

[0084] The method for extracting the speech encoding features of a fixed dimension can be any suitable method. For example, encoding the speech segment of the current time period through an encoder to obtain speech encoding features of a fixed dimension includes: extracting the acoustic features of the speech segment of the current time period, such as FilterBank features, Mel cepstral coefficient features (MFCC), or other suitable features. Any suitable method can be used to extract the acoustic features, and no specific limitation is made here; inputting the acoustic features into a convolutional network module to extract the abstract speech representation information in the acoustic features; inputting the abstract speech representation information into multiple Transformer network structures of the encoder to extract high-dimensional abstract speech representation information; inputting the high-dimensional abstract speech representation information into a feed-forward neural network layer for processing to output the speech encoding features of the fixed dimension.

[0085] Among them, the number of Transformer network structures of the encoder can be reasonably set according to actual needs. Each Transformer network structure can be implemented based on the Transformer block well-known to those skilled in the art, and specific limitations are not made here. In some embodiments, the encoding part adopts a Transformer structure. The Transformer adopts a structure with Self-Attention as the basic unit, and the Transformer can more effectively learn the context relationship of the input, so as to give richer and more accurate speech coding features. For example, the Transformer block used in the encoder part can be composed of a predetermined number (such as 6) of identical composite layers. Each composite layer is composed of a multi-head self-attention mechanism and a fully connected position feed-forward network. Except for the first composite layer, other composite layers use the output of the previous layer as the input. In the composite layer, after each single layer, it will be processed by a network similar to the residual structure and a normalization layer.

[0086] In some embodiments, the decoder part can extract high-dimensional abstract text representation information through any suitable method. For example, the decoder extracts high-dimensional abstract text representation information from the historical predicted text labels that have been predicted before the current period, including: converting the historical predicted text labels into character embedding vectors (i.e., text embedding vectors); inputting the character embedding vectors into a convolutional network model to extract abstract text representation information; inputting the abstract text representation information into multiple Transformer network structures of the decoder to extract the high-dimensional abstract text representation information.

[0087] In some embodiments, a Transformer-based decoder can be used to generate text content. Optionally, the decoder can also have N composite layers. The difference is that each composite layer can have four single layers: an ordinary Multi-head Self-Attention, a Multi-head Self-Attention for processing the relationship between topic words (referred to as topic multi-head self-attention), a Multi-head Self-Attention for processing knowledge graph information (referred to as knowledge multi-head self-attention), and a fully connected feed-forward neural network. A normalization layer is also used for processing between each single layer. Among them, for the topic multi-head self-attention mechanism, the hidden state vector of the topic word is used for weight calculation. In the knowledge multi-head self-attention mechanism, the hidden state vector of the triple is used for weight calculation. Among them, N is any positive integer set in advance, and can be 6 or 12 or other suitable values.

[0088] By predicting the predicted text labels of each speech segment and merging the predicted text labels of all speech segments in chronological order, the predicted text labels of the speech information to be recognized are obtained. Furthermore, the text information corresponding to the speech information to be recognized can be generated through the predicted text labels.

[0089] After the text information is obtained, it can also be displayed on a display so that the user can obtain the text information.

[0090] In a specific example, such as Figure 3 shown, taking the speech recognition of mixed Mongolian and Chinese as an example. First, the original speech data can be segmented by VAD and the invalid speech can be filtered to obtain the speech information to be recognized. The speech information to be recognized includes multiple speech segments, and these speech segments are valid speech. Then, the trained language recognition model is used to recognize the language of the speech segment in the current time period to predict the score of the target language in the speech segment of the current time period. Among them, the score of the target language includes the first score of the first language and the second score of the second language; the first score is compared with the first threshold, and the second score is compared with the second threshold. When the first score is less than the first threshold and the second score is less than the second threshold, it is determined that the language information includes the target language. When the language information of the speech segment includes the target language, such as Mongolian and Chinese, the speech segment in the current time period is recognized through the trained transcription model (such as the Mongolian-Chinese transcription model) to convert the speech segment into text information (that is, the text result), and the text information is output. Through the transcription model of the present application, it is possible to output the mixed speech result (such as the Mongolian-Chinese mixed text result) for the speech containing mixed languages (such as mixed Mongolian and Chinese) in the whole sentence and clause.

[0091] Below, the training process of the transcription model will be introduced.

[0092] Taking the mixed speech of Mongolian and Chinese as an example, Mongolian, as the main ethnic language of the Mongolian ethnic group in China, has a long history. With the increasingly close exchanges between the Mongolian and Han ethnic groups in society, politics, economy, etc., and the rapid development of the Internet and information industry in recent years, the phenomenon of mixed languages has emerged in daily communication. Among them, due to the fact that the traditional Mongolian script cannot describe some newly constructed words, the situation of mixed Mongolian and Chinese speech is increasing day by day. However, the related technologies mainly model for a single language. In the complex scenario of multi-language mixing, the input speech stream is preprocessed at the front end and segmented into clauses, then the language confidence of the clauses is judged, and then the recognition result is output through the speech recognition model corresponding to the language with the highest confidence. Finally, the recognition results of the clauses are spliced as the final whole sentence result. However, the above-mentioned related technologies have the following defects:

[0093] (1) The entire sentence of speech is pre - processed into clauses, and then the speech recognition model in the corresponding language of the clause is used to output the recognition result. Therefore, the situation of mixed Mongolian and Chinese in the clause still cannot be solved, which is very likely to lead to the inability to obtain accurate text results or even the inability to obtain recognition results.

[0094] (2) In actual scenarios, due to the relatively difficult collection of multi - language mixed corpora and the lack of a unified linguistic norm, it is difficult to construct training corpora. For example, low - resource Mongolian - Chinese mixed corpora cannot meet the training needs of complex speech recognition models, resulting in the model's performance not reaching the practical threshold.

[0095] The training process can include processes such as data construction, model construction, and model training.

[0096] First, in the data construction part, to solve the problems such as the relatively difficult collection of multi - language mixed corpora and the lack of a unified linguistic norm, as Figure 3 shown, first synthesize and simulate a single - language corpus of, for example, Mongolian to obtain synthetic speech of the target language. The synthetic speech is obtained by synthesizing the phoneme sequence corresponding to the historical text of the target language and the voiceprint information of the historical speech of the target language. For example, it is obtained through synthesis by a text - to - speech (TTS) model. The TTS model can also be GlowTTS or other suitable models. The specific synthesis process is not specifically limited here.

[0097] Since the text data of small languages such as Mongolian is easier to collect than speech data, after converting it into a phoneme sequence, the corresponding speech data is generated through a synthesizer. In order to improve the timbre diversity of the synthetic speech and make it more conform to real Mongolian speech, the voiceprint information of the historical speech of the target language, such as the voiceprint characteristics corresponding to real Mongolian speech, is additionally added, and a voiceprint characteristic is randomly selected for speech synthesis during synthesis.

[0098] Through the speech synthesis scheme, it is possible to conveniently expand speech data using more easily collected text data, thus alleviating the problem that the real corpus of some target languages, such as Mongolian, is scarce and cannot construct complex model training data. However, the timbre in the synthesized Mongolian speech is less (still based on the voiceprint characteristics of a small number of speakers for speech synthesis). Therefore, although the speech data is sufficient, the performance of the recognition model obtained by ordinary methods is still not ideal.

[0099] Furthermore, in order to improve the performance of the recognition model, during the data construction of this application, multi-language random splicing and data augmentation were also carried out. On the one hand, the ED model (i.e., the Encoder-Decoder model) usually has a higher recognition rate for short speech and a lower recognition rate for long speech. This is mainly because after the training data is segmented by VAD and then labeled, most of the training set consists of short speech, resulting in weak generalization ability for long speech. On the other hand, it is more difficult to collect mixed language corpora than single-language corpora. Therefore, during the data construction of this application, multi-language random splicing was performed. Two voices were randomly selected from the training set for splicing to obtain spliced voices. At the same time, the corresponding text labels were also spliced in sequence to form a parallel data, which can increase the duration of the voice. Then, it was mixed with the original training data set to expand the corpus and increase data diversity, especially increasing the distribution diversity of long and short language materials.

[0100] In addition, background noises such as music and reverberation were added to the original voice to further augment the training data.

[0101] Therefore, the training data of the transcribing model includes the synthetic voice of the target language and the corresponding text information, the text information corresponding to the original voice of the target language, the spliced voice of the target language and the corresponding text information, and the augmented voice of the target language. Among them, the synthetic voice is obtained by synthesizing the phoneme sequence corresponding to the historical text of the target language and the voiceprint information of the historical voice of the target language. The spliced voice is obtained by splicing two randomly selected voices from the training data. The augmented voice is obtained by adding background noise to the original voice.

[0102] Next, in the model construction part

[0103] Such as Figure 4As shown, the transcribing model is a speech recognition model based on the encoder-decoder framework, which mainly includes two branches, the Encoder branch (i.e., the encoder part) and the Decoder branch (i.e., the decoder part). The encoder of the transcribing model to be trained includes a feature extraction module, a convolutional network module, multiple first Transformer network structures, a feed-forward neural network layer, a deconvolution network module, a fully connected layer, and a normalization network module connected in sequence. The decoder of the transcribing model to be trained includes a conversion module, a convolutional network module, multiple second Transformer network structures, a feed-forward neural network layer, and a normalization network module connected in sequence. Optionally, the normalization network module includes, but is not limited to, a softmax module. There are differences in the structures between the transcribing model to be trained and the trained transcribing model. Among them, the deconvolution network module, the fully connected layer, and the normalization network module of the encoder are used to assist in training. In the trained transcribing model, there are no deconvolution network module, fully connected layer, and normalization network module in the encoder. Among them, the feature extraction module of the Encoder branch can be used to extract FilterBank features from audio data (such as speech information) as input, extract abstract speech representation information through the convolutional network module, then extract high-dimensional abstract speech representation information through N Transformer Blocks, and finally obtain fixed-dimensional speech coding features through a feed-forward module (such as a feed-forward neural network (FFN) layer).

[0104] On the one hand, the Decoder branch is used to convert the text sequence corresponding to the speech into a character embedding vector (Embedding) as input, extract abstract text representation information through the convolutional network module model, and then extract high-dimensional abstract text representation information through M Transformer Blocks; on the other hand, it takes the output of the Encoder branch as input and performs feature weighted fusion with the high-dimensional abstract text features through the attention mechanism, and finally outputs the predicted text sequence through the feed-forward neural network module and softmax.

[0105] Next, in the model training part, the following steps will be carried out:

[0106] 1) Random speech feature masking

[0107] During the training process, random masking processing is performed on the extracted acoustic features. Among them, the random masking processing includes: randomly occluding a predetermined number of time-domain features in the spectrogram corresponding to the acoustic features, and / or randomly occluding a predetermined number of frequency-domain features in the spectrogram corresponding to the acoustic features;

[0108] The speech recognition model based on the encoder-decoder framework will have a certain overfitting problem. The model can predict known data well, but predict unknown data poorly. In the model training stage of this application, time-domain and frequency-domain random speech feature masks are designed. Without introducing additional data, by directly enhancing the spectrogram (i.e., the spectrogram of the extracted acoustic features), the overfitting problem is solved, thereby improving the speech recognition accuracy. The time-domain speech feature mask refers to randomly masking some time-domain features in the spectrogram (such as the T-MASK marked positions shown in Figure 5 ), and the frequency-domain speech feature mask refers to randomly masking some frequency-domain features in the spectrogram (such as the F-MASK marked positions shown in Figure 5 ). This strategy can help the model be more robust in the face of the loss of some frequency signals and the absence of signals in some time periods.

[0109] 2) Random text feature perturbation

[0110] In the speech recognition model training stage, when decoding and predicting the text label at time t, the true text labels at time t-1 and before will be introduced as input information. However, in the application stage, when decoding and predicting the text label at time t, only the model-predicted text labels at time t-1 and before can be relied on as input.

[0111] To alleviate the deviation between the model training and model inference links regarding the input information of historical text labels, this application designs two random text feature perturbation strategies, one for the text character level granularity and the other for the finer text phoneme level granularity. This strategy is only applied in the model training stage. During training, random text feature perturbation is performed on the true text labels corresponding to the text information in the training dataset. Among them, the random text feature perturbation includes replacing randomly selected random positions of the true text labels with characters or phonemes of non-true labels at a predetermined ratio (i.e., characters or phonemes different from the true text labels), thereby increasing the robustness of the model in the inference stage when there are problems with some predicted text labels at historical moments for predicting the current text label. The random perturbation of text at both the character and phoneme levels also greatly enhances the anti-interference performance of the current moment decoding in the model inference stage against incorrect predicted labels at historical moments.

[0112] Since the transcription model of this application adopts an end-to-end framework, in the training stage, the Encoder branch takes speech as input to extract FilterBank features. In this process, the aforementioned random feature masking is adopted to increase the diversity of input data. Then, shallow representations are extracted through a convolutional network, and then high-dimensional speech representations are further obtained through N transformer blocks. Finally, the output of the Encoder branch is obtained through a feed-forward layer (i.e., a feed-forward neural network layer); the input of the Decoder branch includes two parts. One is the true text label corresponding to the speech in the training data (such as an artificially annotated text label), which is then converted into text Embedding and a convolutional network to obtain shallow text representations, and then deep text representations (i.e., high-dimensional abstract text representation information) are obtained through N transformer blocks. The other input is the output of the Encoder branch. After aligning and fusing it with the deep text information at the historical moment, the text output at the current moment is predicted through a feed-forward layer and softmax. After obtaining the complete text output at all moments in sequence, the internal parameters of the model are iteratively adjusted by comparing with the known artificially annotated results and error feedback.

[0113] To improve the performance of the speech recognition model, this application proposes a multi-task (Multi Task) training model, including the Connectionist Temporal Classification (CTC) loss function (i.e., Encoder CTC Loss), the encoder phone Cross Entropy Loss (CE) loss function (Encoder Phone CE Loss) for the Encoder branch, the decoder character CE loss function (i.e., DecoderChar CE Loss) for the Decoder branch, and the decoder Sequence Discriminative training (SDT) loss function (i.e., SDT Loss). Among them, the SDT loss function is also a loss function for minimizing the word error rate training.

[0114] In some embodiments, obtaining the trained transcription model through the following steps in advance includes the following: obtaining training data, where the training data includes speech information and text labels corresponding to the speech information. Optionally, the process of obtaining training data includes: obtaining the true text labels corresponding to the speech information in the training dataset, and performing random text feature perturbation on the true text labels corresponding to the text information in the training dataset. The random text feature perturbation includes replacing random positions of randomly selected true text labels with characters or phonemes of non-true labels at a predetermined ratio. Extracting the acoustic features of the speech information in the current time period of the input training dataset through the feature extraction module, and performing the aforementioned random masking process on the extracted acoustic features; extracting the speech coding features of a fixed dimension in the acoustic features through the convolutional network module of the encoder, the multiple first Transformer network structures, and the feedforward neural network layer; calculating the loss based on the CTC loss function for the speech coding features of the fixed dimension to obtain the first loss; using the deconvolution network module to upsample the time dimension corresponding to the speech coding features of the fixed dimension to be consistent with the time dimension of the input speech information in the current time period, and using the fully connected layer and the normalization network module to process the output of the deconvolution network module to obtain the predicted phoneme labels; calculating the second loss of the phoneme sequence of the predicted phoneme labels relative to the phoneme sequence of the true labels by using the cross-entropy loss function; obtaining the text labels corresponding to the speech information before the current time period in the training data; inputting the text labels corresponding to the speech information before the current time period into the conversion module to be converted into character embedding vectors; inputting the character embedding vectors into the convolutional network model of the decoder to extract abstract text representation information; inputting the abstract text representation information into the multiple second Transformer network structures of the decoder to extract the high-dimensional abstract text representation information; performing feature weighted fusion on the speech coding features of the fixed dimension output by the feedforward neural network layer of the encoder and the high-dimensional abstract text representation information through the attention mechanism to obtain the fusion features; inputting the fusion features into the feedforward neural network layer and the normalization network module for processing to obtain the predicted text sequence; calculating the third loss at the character level of the predicted text sequence by using the cross-entropy loss function; calculating the fourth loss of the predicted text sequence by using the sentence-level loss function; performing weighted summation on the first loss, the second loss, the third loss, and the fourth loss to obtain the overall loss; adjusting the model parameters in the transcription model to be trained by using the overall loss to obtain the trained transcription model, and obtaining the trained transcription model through an iterative update process that makes the overall loss smaller and smaller until the transcription model to be trained converges.

[0115] Among them, the Encoder CTC Loss is mainly used to improve the performance of the Mongolian-Chinese recognition model in noisy audio and long audio. By introducing the CTC loss function, it helps the Encoder input features to better complete feature alignment, thereby obtaining high-purity information coding spikes to improve the decoding effect. The loss of the CTC loss function is used to characterize the difference between the predicted label sequence and the true label sequence containing blanks.

[0116] The specific formula is as follows, where y * represents the true label sequence containing blanks, and x represents the input.

[0117]

[0118] The introduction of the Encoder Phone CE Loss strengthens the prediction ability of the Encoder output features for phoneme information. The transposed convolution module is used to upsample the Encoder output in the time dimension to be consistent with the original input, and then the fully connected layer is connected to the softmax and the Encoder Phone CE loss function to obtain the second loss, that is, the phoneme label prediction loss. The specific formula is as follows, where y n represents the phoneme sequence corresponding to the true label, W yn represents the weight, represents the loss corresponding to the phoneme sequence, X n represents the input, and C represents the category.

[0119]

[0120]

[0121] The Decoder Char CE loss obtains the character-level loss (i.e., the third loss) of the text sequence predicted by the decoder through the fully connected layer connected to the softmax. The specific formula is basically the same, and the true phoneme sequence label needs to be replaced with the true word sequence label.

[0122] The Decoder SDT loss is a sentence-level loss function, aiming to balance the insertion and deletion errors introduced by other objective functions. The specific formula is as follows, where is the average word error rate of the N-best sequences (i.e., candidate sequences), represents the probability that the output for the given input x is the i-th candidate sequence, and W(yi, y * ) represents the word error rate of yi, and x represents the input.

[0123]

[0124] The above loss functions are weighted and summed to obtain the overall loss function for network training, as shown in the following formula:

[0125]

[0126] Among them, the values of the weights W1, W2, W3, and W4 can be reasonably set according to actual needs. For example, all four can be made 0.25, or alternatively, W2 and W3 can be made greater than W1 and W4 respectively.

[0127] Through the four loss functions designed in multi-task training, and the performance complementarity between different loss functions, the overall performance of the transcription model is guaranteed.

[0128] It should be noted that the Connectionist Temporal Classification (CTC) loss function refers to a loss function based on time series annotation. The corresponding methods for constructing the CTC loss function, cross-entropy loss function, and SDT loss function in the existing related technologies are also applicable to this application.

[0129] Based on the above description, through the mixed-language speech recognition method of this application, first perform language recognition on the to-be-recognized speech information to determine the language information of the to-be-recognized speech information, so as to screen the data of the speech information. When it is screened that the speech information includes the target language, input the to-be-recognized speech information into the trained transcription model to convert the to-be-recognized speech information into text information, thereby making the data input into the transcription model more in line with the model requirements, improving the recognition efficiency of the model, and further making the accuracy of the obtained text information higher, being able to output the recognition result of multi-language speech mixture, and enhancing the user experience.

[0130] Next, it will be combined with Figure 7 Describe a mixed-language speech recognition device 700 provided according to another aspect of this application, which can be used to execute the mixed-language speech recognition method according to the embodiments of this application described above.

[0131] As Figure 7As shown in the figure, the mixed-language speech recognition device 700 may include: an acquisition module 710, configured to acquire speech information to be recognized; a language recognition module 720, configured to perform language recognition on the speech information to be recognized to determine the language information of the speech information to be recognized; a transcription module 730, configured to, when the language information includes a target language, input the speech information to be recognized into a trained transcription model to convert the speech information to be recognized into text information, where the target language includes a first language and / or a second language, and the text information includes mixed-language text information corresponding to the first language and the second language. Details of each module of the device may refer to the relevant descriptions of the foregoing method, and will not be described one by one here.

[0132] Next, a mixed-language speech recognition system 800 provided according to another aspect of the present application will be described in conjunction with Figure 8 It can be used to execute the mixed-language speech recognition method according to the embodiments of the present application described above.

[0133] The mixed-language speech recognition device in the foregoing embodiment can be used in the mixed-language speech recognition system 800, and the mixed-language speech recognition system 800 can be, for example, various terminal devices, such as mobile phones, computers, tablet computers, etc.

[0134] As Figure 8 shown, the mixed-language speech recognition system 800 may include a memory 810 and a processor 820. The memory 810 stores a computer program run by the processor 820. When the computer program is run by the processor 820, the processor 820 executes the mixed-language speech recognition method 100 according to the embodiments of the present application described above. Those skilled in the art can understand the specific operations of the mixed-language speech recognition method 100 according to the embodiments of the present application in combination with the foregoing content. For the sake of brevity, specific details will not be described here.

[0135] The processor 820 may be any processing system well known in the art. For example, a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor, a microcontroller, a field programmable gate array (FPGA), etc. The present invention is not limited thereto.

[0136] Among them, the memory 810 is used to store network parameters of one or more neural networks. Exemplarily, the memory 810 may be a RAM, a ROM, an EEPROM, a flash memory or other storage technologies, a CD-ROM, a digital versatile disc (DVD) or other optical disc storage devices, a magnetic tape cassette, a magnetic tape, a magnetic disk storage device or other magnetic storage systems, or any other medium that can be used to store desired information and can be accessed by the processor 820.

[0137] The mixed-language speech recognition system 800 further includes a display (not shown), which can be used to display various visual information, such as the text information obtained by transcription, etc.

[0138] The mixed-language speech recognition system 800 may further include a communication interface (not shown), and information interaction between hardware such as a processor, a communication interface, and a memory can be achieved through a communication bus.

[0139] In addition, according to the embodiments of the present application, a storage medium is further provided. Program instructions are stored on the storage medium, and when the program instructions are run by a computer or a processor, they are used to execute the corresponding steps of the mixed-language speech recognition method 100 of the embodiments of the present application. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.

[0140] Although example embodiments have been described herein with reference to the drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present application thereto. Those of ordinary skill in the art can make various changes and modifications therein without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as claimed by the appended claims.

[0141] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0142] In several embodiments provided by the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0143] In the specification provided herein, numerous specific details are set forth. However, it will be understood that embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0144] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the description of the exemplary embodiments of the present application, various features of the present application are sometimes grouped together in a single embodiment, figure, or description thereof. However, the methods of the present application should not be construed as reflecting an intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected by the corresponding claims, the inventive point lies in that a corresponding technical problem can be solved by features less than all the features of a single disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.

[0145] Those skilled in the art will appreciate that, except where features are mutually exclusive, any combination can be used of all the features disclosed in this specification (including the accompanying claims, abstract and drawings), as well as of all the processes or units of any method or system so disclosed. Each feature disclosed in this specification (including the accompanying claims, abstract and drawings), unless explicitly stated otherwise, may be replaced by alternative features serving the same, equivalent or similar purpose.

[0146] In addition, those skilled in the art will appreciate that, although some embodiments described herein include certain features included in other embodiments but not others, combinations of features of different embodiments are meant to be within the scope of the present application and form different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.

[0147] It should be noted that the above embodiments illustrate rather than limit the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names.

Claims

1. A method for mixed-language speech recognition, characterized in that, the recognition method includes the following steps: Obtain the speech information to be recognized; Perform language recognition on the speech information to be recognized to determine the language information of the speech information to be recognized; When the language information includes the target language, input the speech information to be recognized into the trained transcription model to convert the speech information to be recognized into text information, where the target language includes the first language and the second language, and the text information includes the mixed-language text information corresponding to the first language and the second language. Among them, the training process of the transcription model includes: During the training process, perform random masking on the extracted acoustic features, where the random masking process includes: randomly masking a predetermined number of time-domain features in the spectrogram corresponding to the acoustic features, and / or randomly masking a predetermined number of frequency-domain features in the spectrogram corresponding to the acoustic features, the first language is a minority language, and the training data of the transcription model includes the synthesized speech of the first language and the corresponding text labels, the text labels corresponding to the original speech of the target language, the spliced speech of the target language and the corresponding text labels, and the augmented speech of the target language. Among them, the synthesized speech is obtained by synthesizing the phoneme sequence corresponding to the historical text of the first language and the voiceprint information of the historical speech of the first language, the spliced speech is obtained by splicing two randomly selected speeches in the training data, and the augmented speech is obtained by adding background noise to the original speech.

2. The recognition method according to claim 1, characterized in that, the transcription model is a speech recognition model based on the encoder-decoder framework. Among them, the encoder of the transcription model to be trained includes a feature extraction module, a convolutional network module, multiple first Transformer network structures, a feed-forward neural network layer, a deconvolutional network module, a fully connected layer, and a normalization network module connected in sequence. The decoder of the transcription model to be trained includes a conversion module, a convolutional network module, multiple second Transformer network structures, a feed-forward neural network layer, and a normalization network module connected in sequence. Among them, the method pre-trains the trained transcription model through the following steps: Obtain training data, where the training data includes speech information and the text labels corresponding to the speech information; Extract the acoustic features of the speech information in the current period in the input training data set through the feature extraction module, and perform the random masking process on the extracted acoustic features; Extract the speech coding features of a fixed dimension in the acoustic features through the convolutional network module, the multiple first Transformer network structures, and the feed-forward neural network layer of the encoder; Calculate the loss of the speech coding features of the fixed dimension based on the CTC loss function to obtain the first loss; Use a deconvolution network module to upsample the time dimension corresponding to the speech coding features of the fixed dimension to be consistent with the time dimension of the input speech information of the current period, and use a fully connected layer and a normalization network module to process the output of the deconvolution network module to obtain predicted phoneme labels; Use the cross-entropy loss function to calculate the second loss of the phoneme sequence of the predicted phoneme labels relative to the phoneme sequence of the true labels; Obtain the text label corresponding to the speech information before the current period in the training data; Input the text label corresponding to the speech information before the current period into the conversion module to be converted into a character embedding vector; Input the character embedding vector into the convolutional network model of the decoder to extract abstract text representation information; Input the abstract text representation information into multiple second Transformer network structures of the decoder to extract high-dimensional abstract text representation information; Perform feature weighted fusion on the fixed-dimension speech coding features output by the feed-forward neural network layer of the encoder and the high-dimensional abstract text representation information through an attention mechanism to obtain fusion features; Input the fusion features into the feed-forward neural network layer and the normalization network module for processing to obtain a predicted text sequence; Use the cross-entropy loss function to calculate the third loss at the character level of the predicted text sequence; Use the sentence-level loss function to calculate the fourth loss of the predicted text sequence; Perform weighted summation on the first loss, the second loss, the third loss, and the fourth loss to obtain an overall loss; Use the overall loss to adjust the model parameters in the transcribing model to be trained to obtain the trained transcribing model.

3. The recognition method according to claim 2, characterized in that, The obtaining of the training data includes: Obtain the true text label corresponding to the speech information in the training data set; Perform random text feature perturbation on the true text label to obtain the text label corresponding to the speech information, where the random text feature perturbation includes: replacing a random position of a randomly selected true text label with characters or phonemes of non-true labels at a predetermined ratio.

4. The recognition method according to claim 1, characterized in that, Performing language recognition on the obtained speech information to be recognized to determine the language information of the speech information to be recognized includes: Performing language recognition on the speech information to be recognized through a trained language recognition model to predict the scores of the target languages in the speech information to be recognized, where the scores of the target languages include the first score of the first language and the second score of the second language; Compare the first score with a first threshold, and compare the second score with a second threshold. When the first score is less than the first threshold and the second score is less than the second threshold, it is determined that the language information includes the target language.

5. The recognition method according to claim 1, characterized in that, The transcription model is a speech recognition model based on an encoder-decoder framework. The speech information to be recognized includes speech segments in multiple time periods. Inputting the speech information to be recognized into the trained transcription model to convert the speech information to be recognized into text information includes: Performing encoding and decoding processing on the speech segment in each time period to predict a predicted text label corresponding to the speech segment in each time period; Merging the predicted text labels corresponding to the speech segments in all time periods in chronological order to obtain a predicted text label corresponding to the speech information to be recognized; Obtaining the text information corresponding to the speech information to be recognized according to the predicted text label corresponding to the speech information to be recognized.

6. The recognition method according to claim 1, wherein, Obtaining the speech information to be recognized includes: Obtaining the original speech information; Segmenting the original speech information through voice activity endpoint detection and filtering out invalid speech in the original speech information to obtain the speech information to be recognized.

7. A mixed-language speech recognition device, wherein, The device includes: An acquisition module for acquiring speech information to be recognized; A language recognition module for performing language recognition on the speech information to be recognized to determine the language information of the speech information to be recognized; A transcription module for, when the language information includes a target language, inputting the speech information to be recognized into the trained transcription model to convert the speech information to be recognized into text information. The target language includes a first language and a second language, and the text information includes mixed-language text information corresponding to the first language and the second language. Among them, The training process of the transcription model includes: During the training process, performing random masking processing on the extracted acoustic features, where the random masking processing includes: randomly occluding a predetermined number of time-domain features in the spectrogram corresponding to the acoustic features, and / or randomly occluding a predetermined number of frequency-domain features in the spectrogram corresponding to the acoustic features; The first language is a minority language. The training data of the transcription model includes the synthetic speech of the first language and the corresponding text labels, the text labels corresponding to the original speech of the target language, the spliced speech of the target language and the corresponding text labels, and the augmented speech of the target language. Among them, the synthetic speech is obtained by synthesizing the phoneme sequence corresponding to the historical text of the first language and the voiceprint information of the historical speech of the first language, the spliced speech is obtained by splicing two randomly selected speeches in the training data, and the augmented speech is obtained by adding background noise to the original speech.

8. A mixed-language speech recognition system, wherein, The system includes a memory and a processor. A computer program is stored on the memory and run by the processor. When the computer program is run by the processor, the processor executes the mixed-language speech recognition method according to any one of claims 1-6.

9. A storage medium, wherein, A computer program is stored on the storage medium, and when the computer program runs, it executes the mixed-language speech recognition method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Speech recognition method and device and related equipment

    CN113870840A

  • Hybrid speech recognition method and device, electronic equipment and storage medium

    CN114694637A