Speech Recognition Method, Apparatus, Electronic Device and Storage Medium
By decoding language recognition in parallel decoding of the recognized speech, the inefficiency and resource occupation problems caused by user language selection and translation steps in the prior art are solved, and high accuracy and low cost speech recognition are achieved.
Patent Information
- Application Number
- CN202111550980.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-12-17
AI Technical Summary
Existing voice recognition technology requires users to actively choose languages and need to translate after recognition, resulting in low interaction efficiency, reduced accuracy and increased resource usage.
By using language recognition to recognize speech, obtain language features, and decode them in parallel under the pronunciation language and preset language, share the encoding features, and directly output bilingual recognition text to avoid translation steps.
Improves the accuracy of preset language recognition text, shortens response time, and reduces deployment and maintenance costs.
Smart Images

Figure CN114171002B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a voice recognition method, apparatus, electronic device, and storage medium. Background Art
[0002] As one of the important interfaces for human-computer interaction, voice recognition technology brings a more convenient experience to users and reduces the interaction threshold between humans and machines. However, the complexity of language and the differences in accents still lead to a decrease in the accuracy of voice recognition, affecting the actual user experience.
[0003] In view of the above problems, currently, separate voice recognition systems are provided for various languages. However, users need to actively cooperate to select the voice recognition system corresponding to the language. Especially in the case where users unconsciously mix multiple languages during the interaction process, the independent voice recognition systems for each language cannot accurately recognize this situation. And even if users cooperate and can recognize the correct text in the corresponding language, people who do not understand the language still cannot directly understand the meaning of the voice and need to be translated through a translation system, which will greatly reduce the interaction efficiency. Summary of the Invention
[0004] The present invention provides a voice recognition method, apparatus, electronic device, and storage medium to solve the problem in the prior art that language needs to be manually selected before voice recognition, and the recognized text needs to be translated again, which affects the interaction efficiency.
[0005] The present invention provides a voice recognition method, including:
[0006] Performing language recognition on the voice to be recognized to obtain the language feature of the voice to be recognized;
[0007] Based on the language feature, performing voice decoding on the encoded feature of the voice to be recognized to obtain the recognition texts of the voice to be recognized in the voice language and a preset language respectively, where the voice language is the language indicated by the language feature.
[0008] According to the voice recognition method provided by the present invention, the performing voice decoding on the encoded feature of the voice to be recognized based on the language feature to obtain the recognition texts of the voice to be recognized in the voice language and a preset language respectively includes:
[0009] Based on the language feature, performing voice decoding on the encoded feature in the voice language to obtain the decoded feature and recognition text of the voice to be recognized in the voice language;
[0010] Based on the decoded feature and the language feature, or based on the language feature, performing voice decoding on the encoded feature in the preset language to obtain the recognition text of the voice to be recognized in the preset language.
[0011] A speech recognition method provided by the present invention, wherein performing language recognition on the speech to be recognized to obtain the language feature of the speech to be recognized includes:
[0012] Performing acoustic feature extraction on the speech to be recognized to obtain the acoustic feature of the speech to be recognized;
[0013] Based on the acoustic feature, performing language recognition on the speech to be recognized to obtain the language feature of the speech to be recognized;
[0014] Based on the acoustic feature, performing speech recognition coding on the speech to be recognized to obtain the coding feature of the speech to be recognized.
[0015] A speech recognition method provided by the present invention, wherein based on the acoustic feature, performing speech recognition coding on the speech to be recognized to obtain the coding feature of the speech to be recognized includes:
[0016] Based on the acoustic feature and the language feature, performing speech recognition coding on the speech to be recognized to obtain the coding feature of the speech to be recognized.
[0017] A speech recognition method provided by the present invention, wherein based on the acoustic feature, performing language recognition on the speech to be recognized to obtain the language feature of the speech to be recognized includes:
[0018] Inputting the acoustic feature into a language recognition model, and the language recognition model extracts the language feature based on the acoustic feature, and performs language recognition based on the extracted language feature to obtain the language of the speech to be recognized output by the language recognition model;
[0019] The language recognition model is trained based on the acoustic features of the sample speeches in each sample language.
[0020] A speech recognition method provided by the present invention, wherein based on the language feature, performing speech decoding on the coding feature of the speech to be recognized to obtain the recognition texts of the speech to be recognized in the speech language and the preset language respectively includes:
[0021] Inputting the acoustic feature and the language feature of the speech to be recognized into a speech recognition model, and the speech recognition model performs speech recognition coding on the speech to be recognized based on the acoustic feature and the language feature, or based on the acoustic feature, and performs speech decoding based on the language feature and the coding feature obtained by coding to obtain the recognition texts of the speech to be recognized output by the speech recognition model in the speech language and the preset language respectively;
[0022] The speech recognition model is trained based on the difference between the sample text and the predicted text of the sample speech in the speech language and the preset language, and the predicted text is determined by the speech recognition model during training based on the acoustic features and language features of the sample speech.
[0023] According to a speech recognition method provided by the present invention, the loss function of the speech recognition model is obtained by weighted summation of the speech language loss and the preset language loss;
[0024] The speech language loss is determined based on the difference between the sample text and the predicted text of the sample speech in the speech language, and the preset language loss is determined based on the difference between the sample text and the predicted text of the sample speech in the preset language.
[0025] The present invention also provides a speech recognition device, including:
[0026] A language recognition unit for performing language recognition on the speech to be recognized to obtain the language features of the speech to be recognized;
[0027] A speech decoding unit for performing speech decoding on the encoded features of the speech to be recognized based on the language features to obtain the recognized texts of the speech to be recognized in the speech language and the preset language respectively, where the speech language is the language indicated by the language features.
[0028] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the steps of the speech recognition method as described in any one of the above are implemented.
[0029] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the speech recognition method as described in any one of the above are implemented.
[0030] The speech recognition method, device, electronic device, and storage medium provided by the present invention perform speech decoding on the encoded features of the speech to be recognized in the speech language and the preset language based on the language features, thereby obtaining a bilingual recognition output in the speech language and the preset language. The speech decoding of the speech language and the preset language is parallel, and there is no need to perform translation based on the recognized text of the speech language, effectively improving the accuracy of the recognized text in the preset language and shortening the response time of speech recognition. The speech decoding of the speech language and the preset speech share the encoded features of the speech to be recognized, that is, they have a unified modeling method, making the deployment more flexible, and thus effectively reducing the deployment and maintenance costs. Description of the Drawings
[0031] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly describe the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0032] Figure 1 is a schematic flowchart of the speech recognition method provided by the present invention;
[0033] Figure 2 is a schematic flowchart of step 120 in the speech recognition method provided by the present invention;
[0034] Figure 3 is a schematic flowchart of step 110 in the speech recognition method provided by the present invention;
[0035] Figure 4 is a schematic flowchart of the speech recognition method provided by the present invention;
[0036] Figure 5 is a schematic structural diagram of the speech recognition device provided by the present invention;
[0037] Figure 6 is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0038] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0039] In the specific process for multi - language speech recognition in the related art, generally, speech data is first received, and then the speech data is input into the speech recognition system corresponding to the language type specified by the user for speech recognition. If the recognized text does not belong to the language that the user can understand, it is further necessary to translate the recognized text into a language that the user can understand. For example, if the speech to be recognized is Cantonese, and the user has selected Cantonese as the language before inputting the speech, the text obtained by speech recognition is the orthographic text of Cantonese. However, if the user himself / herself does not recognize the orthographic text of Cantonese, he / she still cannot directly read the text obtained by speech recognition, and the system still needs to translate the orthographic text of Cantonese into Mandarin text. The above - mentioned solution has the following problems:
[0040] 1) The cascading of speech recognition and text translation will cause the errors and omissions in speech recognition to accumulate further in text translation, thus affecting the text translation effect. For example, if the user specifies the language type as Cantonese, the character accuracy of the Cantonese speech recognition model is 90%, and the accuracy of translating Cantonese into Mandarin in text translation is also 90%, then the final effect of recognizing Mandarin text from Cantonese audio is 90% * 90% = 81%.
[0041] 2) The cascading of speech recognition and text translation makes the overall response time = speech recognition response time + translation response time. This combination undoubtedly increases the waiting time of users and greatly reduces the user experience.
[0042] 3) The cascading of speech recognition and text translation will occupy more hardware resources, resulting in a significant increase in both deployment and later maintenance costs.
[0043] In view of the above problems, an embodiment of the present invention provides a speech recognition method. Figure 1 It is a schematic flowchart of the speech recognition method provided by the present invention, as Figure 1 shown. The method includes:
[0044] Step 110, perform language recognition on the speech to be recognized to obtain the language feature of the speech to be recognized.
[0045] Specifically, the speech to be recognized is the speech data used for speech recognition. The speech to be recognized can be obtained through a sound pickup device. Here, the sound pickup device can be a smart phone, a tablet computer, or also a smart appliance such as a speaker, a TV, and an air conditioner, etc. After the sound pickup device picks up the speech to be recognized through a microphone array, it can also amplify and denoise the speech to be recognized. In addition, the speech to be recognized can be a speech segment formed after the sound pickup ends, or a speech stream during the real-time sound pickup process. The embodiment of the present invention does not make specific limitations on this.
[0046] In order to eliminate the operation of the user specifying the language of the speech to be recognized by himself, thus avoiding the cumbersome problem of user operation and the confusion caused by the user forgetting the operation or misselecting for subsequent speech recognition, the embodiment of the present invention performs language recognition on the speech to be recognized, thereby obtaining a language feature that can reflect the language information of the speech to be recognized itself. The language feature here can be a coding representation of the language type corresponding to the speech to be recognized, or a vector representation of the feature reflecting the performance of the speech to be recognized in the corresponding language type. The embodiment of the present invention does not make specific limitations on this.
[0047] The language of the speech to be recognized can be identified by using the acoustic features of the speech to be recognized. For example, the language corresponding to the acoustic features can be determined through a pre-trained language recognition model. At this time, the language features of the speech to be recognized can be the output features of any layer in the language recognition model during the process of recognizing the language based on the acoustic features.
[0048] Step 120: Based on the language features, perform speech decoding on the encoded features of the speech to be recognized to obtain the recognition texts of the speech to be recognized in the speech language and the preset language respectively, where the speech language is the language indicated by the language features.
[0049] Considering that the language of the speech to be recognized and the language that the user can understand may not be the same language, if translation is carried out after speech recognition, problems will arise in terms of accuracy, response timeliness, and maintenance cost. In view of this situation, the embodiments of the present invention propose to directly apply the encoded features of the speech to be recognized to perform parallel speech decoding in the speech language and the preset language, so as to obtain the recognition texts of the speech to be recognized in the speech language and the preset language respectively.
[0050] The process of speech recognition can generally be decomposed into an encoding process and a decoding process. The encoding process is used to perform speech encoding on the speech to be recognized to obtain the encoded features of the speech to be recognized, and the decoding process is used to decode the encoded features of the speech to be recognized to obtain the recognition text of the speech to be recognized. Here, the encoding process and the decoding process can be implemented through a Transformer in the form of an encoder + decoder or other types of models. The encoded features of the speech to be recognized obtained thereby are used to reflect the semantics of the speech to be recognized, that is, to reflect the speech content of the speech to be recognized.
[0051] Specifically, during the process of speech decoding, the encoded features obtained by encoding can be further divided into two branches for parallel decoding, that is, one branch applies the encoded features to perform speech decoding in the speech language, and the other branch applies the encoded features to perform speech decoding in the preset language. Here, the speech language is the language indicated by the language features, that is, the language used by the speech to be recognized itself, and the preset language is a preset language, generally the language that the user can understand or needs to obtain. The two branches share the encoded features obtained in the speech recognition encoding stage, thereby reducing the computational resource consumption of the bilingual recognition output.
[0052] In the decoding processes of the above two branches, the language features and encoding features can be combined for speech decoding. The language features applied to the speech decoding process in the speech language can guide the speech decoding to execute in a direction more conforming to the expression form of the speech language, while the language features applied to the speech decoding process in the preset language can assist the speech decoding to migrate from the expression form of the speech language to the expression form of the preset language, so as to achieve a more reliable and accurate bilingual recognition output of the speech language and the preset language.
[0053] It should be noted that the above speech language and preset language can be languages within the same country or region. For example, speech recognition can support various dialects, such as Sichuan dialect, Cantonese, Minnan dialect, Uyghur, Shanghai dialect, etc. At this time, Sichuan dialect, Cantonese, Minnan dialect, Uyghur, Shanghai dialect, etc. can all be regarded as possible speech languages, and Mandarin can be used as the preset language. The above speech language and preset language can also cover languages within the same country and region as well as different countries and regions. For example, speech recognition can support multiple languages, such as Mandarin, Japanese, Russian, English, etc. The above languages are all regarded as possible speech languages, and Mandarin or English or another language specified by the user can be used as the preset language, thus ensuring the flexibility of speech recognition.
[0054] The method provided by the embodiments of the present invention performs speech decoding on the encoding features of the speech to be recognized in the speech language and the preset language based on the language features, thereby obtaining a bilingual recognition output in the speech language and the preset language. The speech decoding of the speech language and the preset language is parallel, and there is no need to perform translation based on the text recognized in the speech language, effectively improving the accuracy of the text recognized in the preset language and shortening the response time of speech recognition. The speech decoding of the speech language and the preset speech share the encoding features of the speech to be recognized, that is, they have a unified modeling method, making the deployment more flexible, and thus effectively reducing the deployment and maintenance costs.
[0055] Based on the above embodiments, Figure 2 is a schematic flowchart of step 120 in the speech recognition method provided by the present invention, as Figure 2 shown, step 120 includes:
[0056] Step 121, based on the language features, perform speech decoding on the encoding features in the speech language to obtain the decoding features and recognition text of the speech to be recognized in the speech language;
[0057] Step 122, based on the decoding features and the language features, or based on the language features, perform speech decoding on the encoding features in the preset language to obtain the recognition text of the speech to be recognized in the preset language.
[0058] Specifically, in the bilingual decoding process of sharing the encoded features of the speech to be recognized, the speech decoding for the speech language and the speech decoding for the preset language can be separately described:
[0059] The speech decoding for the speech language can be implemented based on the language feature and the encoded feature, where the language feature can guide the speech decoding for the encoded feature to be performed in a direction more fitting the expression form of the speech language. Specifically during decoding, the language feature and the encoded feature can be first fused, and then the fused feature is applied to the speech decoding process of the speech language. During this process, a decoding feature is generated, and based on the decoding feature, the final decoded text, i.e., the recognized text in the speech language, is output. The decoding feature here can include the encoded representation of the output characters at each decoding moment, or include the semantic representation of the output characters at each decoding moment, that is, the decoding feature can reflect the decoding result of the speech decoding of the speech language.
[0060] The speech decoding for the preset language can also be implemented based on the language feature and the encoded feature, where the language feature can assist the speech decoding in migrating from the expression form of the speech language to the expression form of the preset language. Specifically during decoding, the language feature and the encoded feature can be first fused, and then the fused feature is applied to the speech decoding process of the preset language, and the recognized text in the preset language is obtained through speech decoding. At this time, the speech decoding of the speech language and the preset language are independent of each other and not related.
[0061] In addition, for the speech decoding of the preset language, on the basis of the language feature and the encoded feature, the decoding feature obtained from the speech decoding of the speech language can be further combined. The decoding feature for the speech language can provide a reference for the speech decoding of the preset language, so that the speech decoding of the preset language can not only learn the decoding result in the speech language, but also better migrate from the expression form of the speech language to the expression form of the preset language with the assistance of the language feature. Specifically during decoding, the decoding feature, the language feature, and the encoded feature can be first fused, and then the fused feature is applied to the speech decoding process of the preset language, and the recognized text in the preset language is obtained through speech decoding. At this time, the speech decoding of the preset language refers to the result of the speech decoding of the speech language, and compared with the completely independent decoding method, the reliability of the obtained recognized text is further improved.
[0062] It should be noted that the above fusion methods of the language feature and the encoded feature, or the fusion method of the decoding feature, the language feature, and the encoded feature, can be directly splicing each feature, or adding or weighted summing each feature, etc. The embodiments of the present invention do not make specific limitations on this.
[0063] Based on the above embodiments, Figure 3It is a schematic flowchart of step 110 in the speech recognition method provided by the present invention. As Figure 3 shown, step 110 includes:
[0064] Step 111, extracting acoustic features from the speech to be recognized to obtain the acoustic features of the speech to be recognized.
[0065] Here, the extraction of acoustic features can be achieved by frame-dividing and windowing the speech to be recognized, followed by pre-emphasizing each frame. On this basis, the acoustic features of each frame or each combination of multiple frames are extracted through the fast Fourier transform (FFT). The acoustic features here can be Mel Frequency Cepstrum Coefficient (MFCC) features or Perceptual Linear Predictive (PLP) features, etc.
[0066] Step 112, performing language identification on the speech to be recognized based on the acoustic features to obtain the language features of the speech to be recognized.
[0067] Step 113, performing speech recognition coding on the speech to be recognized based on the acoustic features to obtain the coding features of the speech to be recognized.
[0068] Specifically, the acoustic features of the speech to be recognized obtained in step 111 are applied to the language identification of the speech to be recognized on the one hand and to the speech recognition coding of the speech to be recognized on the other hand. The language identification can be performed before the speech recognition coding or in parallel with the speech recognition coding, that is, step 112 and step 113 can be executed synchronously or successively.
[0069] Among them, for language identification based on acoustic features, specifically, the acoustic features of the speech to be recognized can be input into a pre-trained language identification model, and the intermediate features generated by the language identification model during the process of mapping the acoustic features to the language or the features representing the final result are used as the language features. The language identification model here can be trained based on the acoustic features of sample speeches in various languages.
[0070] For speech recognition coding based on acoustic features, specifically, the acoustic features of the speech to be recognized can be input into the encoder of a pre-trained speech recognition model, and the encoder performs speech recognition coding on the acoustic features to obtain coding features that can reflect the speech content of the speech to be recognized. The speech recognition model here can be trained based on sample speeches in various languages and their corresponding sample recognition texts.
[0071] In the method provided by the embodiment of the present invention, the acquisition of the language feature and the coding feature of the speech to be recognized share the acoustic feature of the speech to be recognized, that is, only one model for extracting acoustic features needs to be deployed, which can provide the features required for input for both the language recognition and speech recognition tasks, helping to reduce the deployment and maintenance costs of speech recognition and improve the response efficiency of speech recognition.
[0072] Based on any of the above embodiments, step 113 includes:
[0073] Based on the acoustic feature and the language feature, perform speech recognition coding on the speech to be recognized to obtain the coding feature of the speech to be recognized.
[0074] Specifically, when performing speech recognition coding, the language feature can be further combined on the basis of the acoustic feature. The language feature can provide a reference for speech recognition coding, so that in the process of speech recognition coding, the characteristics of the acoustic feature in expressing speech content under the speech language can be considered more, thereby improving the reliability and accuracy of the coding feature in representing speech content. Specifically, when coding, the acoustic feature and the language feature can be fused, and then the fused feature is applied to the coding process of speech recognition to obtain the coding feature of the speech to be recognized.
[0075] It should be noted that the above fusion method of the language feature and the acoustic feature can be directly splicing each feature, or adding or weighted summing each feature, etc. The embodiment of the present invention does not make specific limitations on this.
[0076] In the method provided by the embodiment of the present invention, the language feature is referred to during speech recognition coding, which can improve the reliability and accuracy of the coding feature in representing speech content, thereby further improving the reliability and accuracy of speech recognition.
[0077] Based on any of the above embodiments, step 112 includes:
[0078] Input the acoustic feature into the language recognition model. The language recognition model extracts the language feature based on the acoustic feature and performs language recognition based on the extracted language feature to obtain the language of the speech to be recognized output by the language recognition model;
[0079] The language recognition model is trained based on the acoustic features of the sample speeches in each sample language.
[0080] Specifically, the acquisition of language features is achieved through a language recognition model. The trained language recognition model can discriminate the language of the input acoustic features and output the language category to which the speech to be recognized belongs. In this process, the language recognition model needs to further abstract the language features that can reflect the language of the speech to be recognized from the input acoustic features, so as to determine the language of the speech to be recognized based on the language features. Here, the language features of the speech to be recognized can be the output features of any layer in the multi-level language recognition model.
[0081] Before performing step 112, it is necessary to complete the training of the language recognition model first. The specific training method includes the following steps:
[0082] First, collect sample data. The sample data here needs to include the acoustic features of the sample speech of each sample language, which can also be understood as the sample data includes the acoustic features of the sample speech and the language category of the sample speech. The sample language here is the language that needs to be supported in speech recognition. The sample data should include the sample speech of at least two sample languages. The specific number and category of the sample languages included in the sample data can be determined according to the requirements of speech recognition.
[0083] Next, based on the constructed initial model, apply the sample data for model training to obtain the language recognition model. The initial model here can be in the form of a Deep Residual Neural Network (DRNN) or a Deep Convolutional Neural Networks (DCNN), etc. The embodiments of the present invention do not make specific limitations on this. In addition, when training the initial model, the training method can be error Back Propagation (BP).
[0084] Based on any of the above embodiments, step 120 includes:
[0085] Input the acoustic features and language features of the speech to be recognized into the speech recognition model. The speech recognition model performs speech recognition encoding on the speech to be recognized based on the acoustic features and the language features, or based on the acoustic features, and performs speech decoding based on the language features and the encoded features obtained by encoding, to obtain the recognition texts of the speech to be recognized output by the speech recognition model in the speech language and the preset language respectively;
[0086] The speech recognition model is trained based on the difference between the sample text and the predicted text of the sample speech in the speech language and the preset language. The predicted text is determined by the speech recognition model in training based on the acoustic features and language features of the sample speech.
[0087] Specifically, the speech recognition process of bilingual recognition output can be implemented through a speech recognition model. The speech recognition model here includes the functions of encoding and decoding. In the encoding stage, the encoder in the speech recognition model can perform speech recognition encoding on the speech to be recognized based on the input acoustic features and language features, or only apply the input acoustic features to perform speech recognition encoding on the speech to be recognized, thereby obtaining encoding features that can express the speech content. In the decoding stage, the decoder in the speech recognition model can perform speech decoding under the two branches of the speech language and the preset language respectively based on the input language features and the encoding features output by the encoder, so as to realize the dual output of the recognized text under the speech language and the preset language.
[0088] Before performing step 120, it is necessary to complete the training of the speech recognition model first. The training method specifically includes the following steps:
[0089] First, collect sample data. The sample data here needs to include the acoustic features, language features of the sample speech of each sample language, and the sample text of the sample speech under the sample language and the preset language respectively. The sample language here is the language that needs to be supported in speech recognition. The sample data should include the sample speeches of at least two sample languages. The specific number and categories of the sample languages included in the sample data can be determined according to the requirements of speech recognition.
[0090] Next, based on the constructed initial model, apply the sample data for model training to obtain the speech recognition model. The initial model here can be improved based on a general encoder-decoder, such as a transformer. For example, two decoders can be connected after the encoder as the two decoding branches for decoding the speech language and the preset language respectively. Specifically, in the model training process, the speech recognition model during training, that is, the initial model, will decode respectively under the speech language and the preset language based on the input acoustic features and language features of the sample speech, so as to output the predicted text under the speech language and the preset language. Subsequently, by comparing the differences between the sample text and the predicted text of the sample speech under the speech language and the preset language, the loss value of the current initial model can be determined, and then the parameters of the initial model can be updated iteratively until the training is completed to obtain the speech recognition model. Here, in the two decoding layers of the initial model, the softmax can be applied to optimize the output result. When training the initial model, its training method can be the Stochastic Gradient Descent (SGD), which can not only accelerate the convergence speed but also prevent the model from falling into local optima.
[0091] Furthermore, during the training process of the initial model, the acoustic features and language features of the sample speech can be processed and fused first to obtain the fused features of the sample speech. The fused features here have stronger discriminability compared to the acoustic features. On this basis, the fused features can be further processed and then input into the initial model to obtain the predicted text in the speech language and the preset language.
[0092] Based on any of the above embodiments, the loss function of the speech recognition model is obtained by weighted summation of the speech language loss and the preset language loss;
[0093] The speech language loss is determined based on the difference between the sample text and the predicted text of the sample speech in the speech language, and the preset language loss is determined based on the difference between the sample text and the predicted text of the sample speech in the preset language.
[0094] Specifically, considering that the speech recognition model is a dual-output model, whether it is speech recognition in the speech language or speech recognition in the preset language, it is an embodiment of the performance of the speech recognition model. Therefore, during the training of the speech recognition model, its loss function can also be divided into two parts, namely the speech language loss and the preset language loss, for measurement, and the weighted summation of the speech language loss and the preset language loss can be used as the final loss value for updating and iterating the model parameters, so as to achieve multi-objective joint training.
[0095] Among them, the speech language loss is used to represent the loss of speech language recognition, and can be specifically determined based on the difference between the sample text and the predicted text of the sample speech in the speech language. Here, the greater the difference between the sample text and the predicted text in the speech language, the greater the speech language loss, and the smaller the difference between the sample text and the predicted text in the speech language, the smaller the speech language loss.
[0096] The preset language loss is used to represent the loss of preset language recognition, and can be specifically determined based on the difference between the sample text and the predicted text of the sample speech in the preset language. Here, the greater the difference between the sample text and the predicted text in the preset language, the greater the preset language loss, and the smaller the difference between the sample text and the predicted text in the preset language, the smaller the preset language loss.
[0097] Thus, the loss function L of the speech recognition model can be expressed by the following formula:
[0098] L = aL1+(1 - a)L2
[0099] In the formula, L1 is the speech language loss, L2 is the preset language loss, a is the weight, and 0 ≤ a ≤ 1.
[0100] Based on any of the above embodiments, Figure 4It is a schematic flow chart of the speech recognition method provided by the present invention. As Figure 4 shown, speech recognition can be achieved through the following steps:
[0101] First, determine the speech to be recognized, and use the speech to be recognized as input for acoustic feature extraction, thereby obtaining the acoustic features of the speech to be recognized. The acoustic feature extraction here can be implemented by a feature extraction module. Specifically, when performing acoustic feature extraction, in order to improve the discriminability of the acoustic features, the extracted spectral features can be transformed. For example, each frame of speech data and multiple frames of speech data before and after each frame of speech data are used as the input of the feature extraction model, and the output of the feature extraction model is used as the transformed acoustic features.
[0102] After obtaining the acoustic features of the speech to be recognized, input the acoustic features into a pre-trained language recognition model, so as to obtain the language features obtained by performing feature extraction on the acoustic features during the language recognition process.
[0103] Subsequently, apply the acoustic features and language features of the speech to be recognized to speech recognition. Figure 4 The dashed box in is the speech recognition model. The speech recognition model includes an encoding layer and two parallel decoding layers, namely the speech language decoding layer and the preset language decoding layer. Among them, the encoding layer encodes using the input acoustic features and language features to obtain encoded features. The speech language decoding layer decodes the speech under the speech language using the input language features and the encoded features obtained by encoding, thereby obtaining the recognition text under the speech language and the decoding features that can reflect the decoding result of the speech decoding under the speech language. The preset language decoding layer decodes the speech under the preset language using the input language features and the encoded features obtained by encoding, as well as the decoding features obtained by decoding the speech language, thereby obtaining the recognition text under the preset language.
[0104] For example, the speech to be recognized is a Cantonese speech. The language features that can represent Cantonese are obtained through language recognition, and the language features and the acoustic features of the speech to be recognized are input into the speech recognition model, thereby obtaining the correct Chinese characters text in Cantonese and the Mandarin text.
[0105] Based on any of the above embodiments, Figure 5 It is a schematic structural diagram of the speech recognition device provided by the present invention. As Figure 5 shown, the device includes:
[0106] A language recognition unit 510, configured to perform language recognition on the speech to be recognized to obtain the language features of the speech to be recognized;
[0107] A voice decoding unit 520, configured to perform voice decoding on the encoded features of the to-be-recognized voice based on the language feature, so as to obtain the recognized texts of the to-be-recognized voice in the voice language and the preset language respectively, where the voice language is the language indicated by the language feature.
[0108] The device provided by the embodiment of the present invention performs voice decoding on the encoded features of the to-be-recognized voice in the voice language and the preset language based on the language feature, thereby obtaining a bilingual recognition output in the voice language and the preset language. The voice decoding of the voice language and the preset language is parallel, without the need for translation based on the recognized text of the voice language, effectively improving the accuracy of the recognized text in the preset language and shortening the response time of voice recognition. The voice decoding of the voice language and the preset voice share the encoded features of the to-be-recognized voice, that is, they have a unified modeling method, making the deployment more flexible, and thus effectively reducing the deployment and maintenance costs.
[0109] Based on any of the above embodiments, the voice decoding unit is configured to:
[0110] Perform voice decoding on the encoded features in the voice language based on the language feature, so as to obtain the decoded features and the recognized text of the to-be-recognized voice in the voice language;
[0111] Perform voice decoding on the encoded features in the preset language based on the decoded features and the language feature, or based on the language feature, so as to obtain the recognized text of the to-be-recognized voice in the preset language.
[0112] Based on any of the above embodiments, the language recognition unit is configured to:
[0113] Extract acoustic features from the to-be-recognized voice to obtain the acoustic features of the to-be-recognized voice;
[0114] Perform language recognition on the to-be-recognized voice based on the acoustic features to obtain the language feature of the to-be-recognized voice;
[0115] Perform voice recognition encoding on the to-be-recognized voice based on the acoustic features to obtain the encoded features of the to-be-recognized voice.
[0116] Based on any of the above embodiments, the language recognition unit is configured to:
[0117] Perform voice recognition encoding on the to-be-recognized voice based on the acoustic features and the language feature to obtain the encoded features of the to-be-recognized voice.
[0118] Based on any of the above embodiments, the language recognition unit is configured to:
[0119] Input the acoustic features into a language recognition model. The language recognition model extracts language features based on the acoustic features and performs language recognition based on the extracted language features to obtain the language of the speech to be recognized output by the language recognition model.
[0120] The language recognition model is trained based on the acoustic features of sample speeches in various sample languages.
[0121] Based on any of the above embodiments, the speech decoding unit is used for:
[0122] Input the acoustic features and language features of the speech to be recognized into a speech recognition model. The speech recognition model performs speech recognition encoding on the speech to be recognized based on the acoustic features and the language features, or based on the acoustic features, and performs speech decoding based on the language features and the encoding features obtained by encoding to obtain the recognition texts of the speech to be recognized output by the speech recognition model in the speech language and the preset language respectively.
[0123] The speech recognition model is trained based on the difference between the sample text and the predicted text of the sample speech in the speech language and the preset language. The predicted text is determined by the speech recognition model during training based on the acoustic features and language features of the sample speech.
[0124] Based on any of the above embodiments, the loss function of the speech recognition model is obtained by weighted summation of the speech language loss and the preset language loss.
[0125] The speech language loss is determined based on the difference between the sample text and the predicted text of the sample speech in the speech language, and the preset language loss is determined based on the difference between the sample text and the predicted text of the sample speech in the preset language.
[0126] Figure 6 An example of a schematic physical structure diagram of an electronic device is shown as Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute a speech recognition method, and the method includes:
[0127] Perform language recognition on the speech to be recognized to obtain the language features of the speech to be recognized.
[0128] Based on the language feature, perform speech decoding on the encoded feature of the speech to be recognized, and obtain the recognition texts of the speech to be recognized in the speech language and the preset language respectively, where the speech language is the language indicated by the language feature.
[0129] In addition, when the logic instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0130] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the speech recognition method provided by the above-mentioned various methods. The method includes:
[0131] Perform language recognition on the speech to be recognized to obtain the language feature of the speech to be recognized;
[0132] Based on the language feature, perform speech decoding on the encoded feature of the speech to be recognized, and obtain the recognition texts of the speech to be recognized in the speech language and the preset language respectively, where the speech language is the language indicated by the language feature.
[0133] On yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the speech recognition method provided by the above-mentioned various methods. The method includes:
[0134] Perform language recognition on the speech to be recognized to obtain the language feature of the speech to be recognized;
[0135] Based on the language feature, perform speech decoding on the encoded feature of the speech to be recognized, and obtain the recognition texts of the speech to be recognized in the speech language and the preset language respectively, where the speech language is the language indicated by the language feature.
[0136] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0137] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A speech recognition method, characterized in that, including: performing language identification on the speech to be recognized to obtain the language feature of the speech to be recognized; based on the language feature, performing parallel speech decoding on the encoded feature of the speech to be recognized to obtain the recognized texts of the speech to be recognized in the speech language and the preset language respectively, where the speech language is the language indicated by the language feature; the performing parallel speech decoding on the encoded feature of the speech to be recognized includes: applying the encoded feature of the speech to be recognized, dividing it into two branches for parallel decoding, one branch performing speech decoding in the speech language, and the other branch performing speech decoding in the preset language, where the preset language is a preset language, and the speech decoding of the speech language and the preset language share the encoded feature of the speech to be recognized.
2. The speech recognition method according to claim 1, wherein the based on the language feature, performing speech decoding on the encoded feature of the speech to be recognized to obtain the recognized texts of the speech to be recognized in the speech language and the preset language respectively includes: based on the language feature, performing speech decoding on the encoded feature in the speech language to obtain the decoded feature and the recognized text of the speech to be recognized in the speech language; based on the decoded feature and the language feature, or, based on the language feature, performing speech decoding on the encoded feature in the preset language to obtain the recognized text of the speech to be recognized in the preset language.
3. The voice recognition method according to claim 1, wherein the performing language identification on the speech to be recognized to obtain the language feature of the speech to be recognized includes: extracting acoustic features of the speech to be recognized to obtain the acoustic features of the speech to be recognized; based on the acoustic features, performing language identification on the speech to be recognized to obtain the language feature of the speech to be recognized; based on the acoustic features, performing speech recognition encoding on the speech to be recognized to obtain the encoded feature of the speech to be recognized.
4. The speech recognition method according to claim 3, wherein the based on the acoustic features, performing speech recognition encoding on the speech to be recognized to obtain the encoded feature of the speech to be recognized includes: based on the acoustic features and the language feature, performing speech recognition encoding on the speech to be recognized to obtain the encoded feature of the speech to be recognized.
5. The voice recognition method according to claim 3, characterized in that the based on the acoustic features, performing language identification on the speech to be recognized to obtain the language feature of the speech to be recognized includes: inputting the acoustic features into a language identification model, and the language identification model extracts language features based on the acoustic features and performs language identification based on the extracted language features to obtain the language of the speech to be recognized output by the language identification model; the language identification model is trained based on the acoustic features of the sample speeches in each sample language.
6. The voice recognition method according to any one of claims 1 to 5, characterized in that the based on the language feature, performing speech decoding on the encoded feature of the speech to be recognized to obtain the recognized texts of the speech to be recognized in the speech language and the preset language respectively includes: Input the acoustic features and language features of the speech to be recognized into a speech recognition model. The speech recognition model performs speech recognition encoding on the speech to be recognized based on the acoustic features and the language features, or based on the acoustic features, and performs speech decoding based on the language features and the encoded features obtained by the encoding to obtain the recognition texts of the speech to be recognized output by the speech recognition model in the speech language and a preset language respectively; The speech recognition model is trained based on the differences between the sample texts and the predicted texts of the sample speech in the speech language and the preset language. The predicted text is determined by the speech recognition model during training based on the acoustic features and language features of the sample speech.
7. The voice recognition method according to claim 6, wherein The loss function of the speech recognition model is obtained by weighted summation of the speech language loss and the preset language loss; The speech language loss is determined based on the difference between the sample text and the predicted text of the sample speech in the speech language, and the preset language loss is determined based on the difference between the sample text and the predicted text of the sample speech in the preset language.
8. A voice recognition device, characterized in that, It includes: A language recognition unit for performing language recognition on the speech to be recognized to obtain the language features of the speech to be recognized; A speech decoding unit for performing parallel speech decoding on the encoded features of the speech to be recognized based on the language features to obtain the recognition texts of the speech to be recognized in the speech language and the preset language respectively. The speech language is the language indicated by the language features; The parallel speech decoding of the encoded features of the speech to be recognized includes: Applying the encoded features of the speech to be recognized and dividing them into two branches for parallel decoding. One branch performs speech decoding in the speech language, and the other branch performs speech decoding in the preset language. The preset language is a preset language, and the speech decoding of the speech language and the preset language shares the encoded features of the speech to be recognized.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the speech recognition method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the speech recognition method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Bilingual hybrid speech recognition method, device and equipment and storage medium
CN110634487A
Voice recognition system and method
CN111933146A
Voice recognition method and device, electronic equipment and storage medium
CN112382275A
Model training method and device, voice recognition method and device, electronic equipment and storage medium
CN112951240A