Phoneme Recognition Method, Device, Electronic Device, and Storage Medium
Through clustering and iterative training methods, phoneme nodes are screened and the first recognition model is constructed, which solves the problem of excessive scale and poor recognition effect of the phoneme recognition model, and realizes efficient phoneme recognition under multilingual conditions.
Patent Information
- Application Number
- CN202210855299.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-07-19
AI Technical Summary
In the prior art, as the language types increase, the scale of phoneme recognition models increases, affecting local chip deployment, and traditional methods have poorer recognition effects on languages with larger differences.
By clustering based on phoneme node similarity, the current phoneme nodes are selected, the redundant nodes are deleted, the first recognition model is constructed, and the iterative training of the phoneme classification layer is combined with feature extraction and iterative training of the phoneme classification layer, the model scale is reduced and the key phoneme nodes are retained.
The scale of the phoneme recognition model is reduced under multilingual conditions, while maintaining high accuracy, and can accurately distinguish phonemes from different languages.
Smart Images

Figure CN115359783B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and in particular, to a phoneme recognition method, apparatus, electronic device, and storage medium. Background Art
[0002] In the field of speech recognition, phonemes are the smallest units in speech. To improve the accuracy of speech recognition, it is necessary to improve the recognition accuracy of each phoneme in speech.
[0003] In actual application scenarios, speech corresponds to different languages. To accurately recognize speech in different languages, currently, a sub-model is trained for each language, and a phoneme recognition model is constructed based on these sub-models. Each sub-model in the phoneme recognition model is used to recognize the phonemes of the speech in each language, and then the corresponding speech recognition result is obtained according to the phoneme recognition result. However, as the number of language types increases, the number of sub-models also increases, resulting in an increase in the scale of the phoneme recognition model, which in turn affects the deployment of the phoneme recognition model on a local chip. Summary of the Invention
[0004] The present invention provides a phoneme recognition method, apparatus, electronic device, and storage medium to solve the defect of the large scale of the phoneme recognition model in the prior art.
[0005] The present invention provides a phoneme recognition method, including:
[0006] Determine the speech to be recognized;
[0007] Input the speech to be recognized into the phoneme recognition model to obtain the phoneme recognition result output by the phoneme recognition model;
[0008] The phoneme recognition model is obtained by training a first recognition model based on the sample speech of multiple languages and the phoneme-level labels of each sample speech. The first recognition model is obtained by screening the phoneme nodes in the second recognition model based on the similarity between the phonemes corresponding to each phoneme node. The second recognition model includes phoneme nodes corresponding to multiple languages respectively.
[0009] According to the phoneme recognition method provided by the present invention, the determining step of the first recognition model includes:
[0010] Cluster the phoneme nodes in the second recognition model based on the similarity between the phonemes corresponding to each phoneme node to obtain multiple clusters;
[0011] Select the current phoneme node from the phoneme nodes in each cluster, and delete the other phoneme nodes in each cluster except the current phoneme node to obtain the first recognition model.
[0012] A phoneme recognition method provided by the present invention, the second recognition model includes a feature extraction layer and phoneme classification layers corresponding to multiple languages respectively, and each phoneme classification layer is constructed based on phoneme nodes corresponding to each language;
[0013] The second recognition model is trained based on the following steps:
[0014] Input the sample speech of each language into the feature extraction layer of the second recognition model to obtain the first phoneme hidden layer features output by the feature extraction layer of the second recognition model;
[0015] Input the first phoneme hidden layer features into the phoneme classification layers of each language to obtain the first phoneme prediction results output by the phoneme classification layers of each language;
[0016] Based on the difference between the phoneme-level labels and the first phoneme prediction results, perform parameter iteration on the feature extraction layer of the second recognition model and the phoneme classification layers of each language to obtain the second recognition model.
[0017] A phoneme recognition method provided by the present invention, after obtaining the first phoneme hidden layer features output by the feature extraction layer of the second recognition model, further includes:
[0018] Based on the first phoneme hidden layer features, determine word-level hidden layer features and / or sentence-level hidden layer features;
[0019] Based on the difference between the word-level labels and word-level prediction results of the sample speech and / or the difference between the language labels and language prediction results of the sample speech, perform parameter iteration on the feature extraction layer of the second recognition model to obtain the second recognition model; the word-level prediction results are determined based on the word-level hidden layer features, and the language prediction results are determined based on the sentence-level hidden layer features.
[0020] A phoneme recognition method provided by the present invention, based on the difference between the word-level labels and word-level prediction results of the sample speech and / or the difference between the language labels and language prediction results of the sample speech, performing parameter iteration on the feature extraction layer of the second recognition model to obtain the second recognition model, includes:
[0021] Input the word-level hidden layer features into a word-level classification layer to obtain the word-level prediction results output by the word-level classification layer, and / or input the sentence-level hidden layer features into a language classification layer to obtain the language prediction results output by the language classification layer;
[0022] Iterate the parameters of the feature extraction layer of the second recognition model based on the difference between the character-level label and the character-level prediction result and / or the difference between the language label and the language prediction result, to obtain the second recognition model.
[0023] According to a phoneme recognition method provided by the present invention, the determining of the character-level hidden feature and / or the sentence-level hidden feature based on the first phoneme hidden feature includes:
[0024] Perform a sliding window on the first phoneme hidden feature to obtain the character-level hidden feature;
[0025] Perform pooling on the character-level hidden feature to obtain the sentence-level hidden feature.
[0026] According to a phoneme recognition method provided by the present invention, the phoneme recognition model is trained based on the following steps:
[0027] Fix the parameters of the feature extraction layer of the first recognition model;
[0028] Input the sample speech of each language into the feature extraction layer of the first recognition model to obtain the second phoneme hidden feature output by the feature extraction layer of the first recognition model;
[0029] Input the second phoneme hidden feature into the current phoneme classification layer to obtain the second phoneme prediction result output by the current phoneme classification layer; the current phoneme classification layer is constructed based on the phoneme nodes selected from the second recognition model;
[0030] Iterate the parameters of the current phoneme classification layer based on the difference between the phoneme-level label and the second phoneme prediction result, to obtain the phoneme recognition model.
[0031] The present invention also provides a phoneme recognition device, including:
[0032] A determination unit, configured to determine the speech to be recognized;
[0033] A recognition unit, configured to input the speech to be recognized into the phoneme recognition model to obtain the phoneme recognition result output by the phoneme recognition model;
[0034] The phoneme recognition model is trained by using the sample speech of multiple languages and the phoneme-level labels of each sample speech to train the first recognition model. The first recognition model is obtained by screening the phoneme nodes in the second recognition model based on the similarity between the phonemes corresponding to the phoneme nodes in the second recognition model. The second recognition model includes phoneme nodes corresponding to multiple languages respectively.
[0035] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the phoneme recognition method described in any one of the above is implemented.
[0036] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the phoneme recognition method described in any one of the above is implemented.
[0037] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the phoneme recognition method described in any one of the above is implemented.
[0038] The phoneme recognition method, device, electronic device, and storage medium provided by the present invention screen the phoneme nodes in the second recognition model based on the similarity between the phonemes corresponding to the phoneme nodes in the second recognition model to obtain the first recognition model. This not only reduces the scale of the first recognition model but also retains the phoneme nodes corresponding to different phonemes in the first recognition model. Furthermore, after training the first recognition model based on the sample voices in multiple languages and the phoneme-level labels of each sample voice, not only is the scale of the obtained phoneme recognition model smaller than that of the second recognition model, but also the phoneme recognition model can accurately distinguish phonemes in different languages. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0040] Figure 1 is a flowchart of the phoneme recognition method provided by the present invention;
[0041] Figure 2 is a flowchart of the method for determining the first recognition model provided by the present invention;
[0042] Figure 3 is a flowchart of the method for training the second recognition model provided by the present invention;
[0043] Figure 4 is a flowchart of another method for training the second recognition model provided by the present invention;
[0044] Figure 5 is a flowchart of the implementation manner of step 420 in another method for training the second recognition model provided by the present invention;
[0045] Figure 6 It is a schematic flowchart of the phoneme recognition model training method provided by the present invention;
[0046] Figure 7 It is a schematic flowchart of another second recognition model training method provided by the present invention;
[0047] Figure 8 It is a schematic structural diagram of the phoneme recognition device provided by the present invention;
[0048] Figure 9 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0049] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0050] Currently, when performing speech recognition on different languages, usually a sub-model is trained for each language, and a phoneme recognition model is constructed based on these sub-models, so as to use each sub-model in the phoneme recognition model to perform phoneme recognition on the speech of each language, and then obtain the corresponding speech recognition result according to the phoneme recognition result. However, as the number of language types increases, the number of sub-models will also increase, resulting in an increase in the scale of the phoneme recognition model, which in turn affects the deployment of the phoneme recognition model on the local chip.
[0051] In addition, in order to avoid increasing the scale of the phoneme recognition model, a language classification branch is also introduced, and then a gradient reversal layer is inserted between the language classification branch and the main branch, so that the phoneme recognition model can learn language-invariant features through gradient adversarial training. However, this method is applicable to languages with small differences, and has poor recognition effects on languages with large differences (such as Cantonese, Minnan dialect, etc., which are quite different from Mandarin).
[0052] In view of this, the present invention provides a phoneme recognition method. Figure 1 It is a schematic flowchart of the phoneme recognition method provided by the present invention. As Figure 1 shown, the method includes the following steps:
[0053] Step 110, determine the speech to be recognized.
[0054] Here, the speech to be recognized is the speech for which phoneme recognition needs to be performed. The speech to be recognized can be obtained through a sound pickup device. Here, the sound pickup device can be a smartphone, a tablet computer, or also a smart appliance such as a speaker, a TV, and an air conditioner, etc. After the sound pickup device obtains the speech to be recognized through the microphone array, it can also amplify and denoise the speech to be recognized. The embodiments of the present invention do not make specific limitations on this.
[0055] Step 120: Input the speech to be recognized into the phoneme recognition model to obtain the phoneme recognition result output by the phoneme recognition model;
[0056] The phoneme recognition model is obtained by training the first recognition model based on sample speeches of multiple languages and the phoneme-level labels of each sample speech. The first recognition model is obtained by screening the phoneme nodes in the second recognition model based on the similarity between the phonemes corresponding to each phoneme node in the second recognition model. The second recognition model includes phoneme nodes corresponding to multiple languages respectively.
[0057] Here, the phoneme nodes of each language respectively correspond to different phonemes of each language. For example, in Mandarin, "a" and "i" are different phonemes, so in Mandarin, "a" and "i" correspond to different phoneme nodes. The second recognition model includes phoneme nodes corresponding to multiple languages respectively, so the second recognition model can distinguish phonemes in different languages through the phoneme nodes corresponding to each language, that is, the second recognition model has the ability to accurately distinguish phonemes in different languages.
[0058] However, if there are a large number of different categories of languages, it will lead to too many phoneme nodes in the second recognition model, and then the number of calculation parameters of the second model will be relatively large. In addition, considering that there may be phonemes with high similarity in the same category of languages, and there may also be phonemes with high similarity between different categories of languages, that is, the phoneme nodes corresponding to each language in the second recognition model may be redundant.
[0059] In this regard, the embodiments of the present invention screen the phoneme nodes in the second recognition model based on the similarity between the phonemes corresponding to each phoneme node in the second recognition model. For example, based on the similarity between phonemes, cluster each phoneme node, cluster the phoneme nodes corresponding to phonemes with high similarity into one category, and then select any phoneme node in the same category as the current phoneme node of the first recognition model, and delete the remaining phoneme nodes in this category, so as to reduce the number of phoneme nodes in the first recognition model, reduce the scale of the first recognition model, and correspondingly also reduce the scale of the phoneme recognition model.
[0060] After screening the phoneme nodes under the second recognition model, redundant phoneme nodes under the second recognition model are filtered out, that is, phoneme nodes corresponding to similar phonemes are filtered out. That is, the current phoneme nodes included in the first recognition model also correspond to phonemes of different categories. Therefore, after training the first recognition model based on the sample voices of multiple languages and the phoneme-level labels of each sample voice, the obtained phoneme recognition model can accurately distinguish phonemes of different languages.
[0061] The phoneme recognition method provided by the embodiment of the present invention screens the phoneme nodes under the second recognition model based on the similarity between the phonemes corresponding to each phoneme node under the second recognition model to obtain the first recognition model. This not only reduces the scale of the first recognition model, but also retains the phoneme nodes corresponding to different phonemes in the first recognition model. Furthermore, after training the first recognition model based on the sample voices of multiple languages and the phoneme-level labels of each sample voice, not only is the scale of the obtained phoneme recognition model smaller than that of the second recognition model, but also the phoneme recognition model can accurately distinguish phonemes of different languages.
[0062] Based on the above embodiments, Figure 2 is a schematic flowchart of the method for determining the first recognition model provided by the present invention. As Figure 2 shown, the steps for determining the first recognition model include:
[0063] Step 210: Cluster each phoneme node under the second recognition model based on the similarity between the phonemes corresponding to each phoneme node to obtain multiple clusters;
[0064] Step 220: Screen the current phoneme node from the phoneme nodes in each cluster, and delete other phoneme nodes in each cluster except the current phoneme node to obtain the first recognition model.
[0065] Specifically, the similarity between phonemes is used to represent the probability that the corresponding two phonemes belong to the same category. The higher the similarity between phonemes, the higher the probability that the corresponding two phonemes belong to the same category; conversely, the lower the similarity between phonemes, the lower the probability that the corresponding two phonemes belong to the same category.
[0066] Based on the similarity between the phonemes corresponding to each phoneme node, the phoneme nodes corresponding to the same category of phonemes are clustered into one category to obtain multiple clusters, that is, the phoneme nodes included in each cluster correspond to the same or similar phoneme categories.
[0067] In addition, the current phoneme node refers to the phoneme node of the first recognition model, and the phoneme categories corresponding to the current phoneme node in different clusters are different. After obtaining multiple clusters, the current phoneme node is selected from the phoneme nodes in each cluster. The current phoneme node can be one or multiple, but the total number of current phoneme nodes in each cluster is less than the total number of phoneme nodes in the second recognition model.
[0068] Optionally, any one or multiple phoneme nodes in each cluster can be used as the current phoneme node, or the phoneme nodes whose distance from the center of each cluster is less than the threshold can be used as the current phoneme node, or the phoneme nodes closest to the center of each cluster can be used as the current phoneme node. The embodiments of the present invention do not make specific limitations on this.
[0069] After determining the current phoneme nodes in each cluster, the other phoneme nodes in each cluster have the same or similar categories as the current phoneme nodes, that is, the other phoneme nodes can be regarded as redundant factor nodes. In this regard, after obtaining the current phoneme nodes of each cluster, the embodiments of the present invention delete the other phoneme nodes in each cluster except the current phoneme nodes to obtain the first recognition model, that is, it can be understood that the first recognition model is obtained by deleting the redundant phoneme nodes from the second recognition model.
[0070] It can be seen that based on the similarity between the phonemes corresponding to each phoneme node, the embodiments of the present invention can accurately cluster each phoneme node under the second recognition model, and then obtain multiple clusters. At the same time, the embodiments of the present invention select the current phoneme nodes from the phoneme nodes in each cluster and delete the other phoneme nodes in each cluster except the current phoneme nodes, so as to reduce the number of phoneme nodes in the first recognition model, realize reducing the scale of the first recognition model, and then correspondingly reduce the scale of the phoneme recognition model.
[0071] As an alternative embodiment, when clustering each phoneme node under the second recognition model, the phonemes corresponding to each phoneme node can be clustered by using the Gaussian mixture model (GMM model) and the Expectation-Maximum (EM) algorithm.
[0072] For example, the second recognition model includes N a phoneme nodes corresponding to languages, and the number of phoneme nodes corresponding to each language is N c , that is, the total number of phoneme nodes included in the second recognition model is N a ×N c If the total number of phoneme nodes in the first recognition model is to be N cIf there are N such nodes, the GMM model can be used to cluster each phoneme node, and then the parameters of the GMM model can be continuously iteratively adjusted according to the clustering results until the clustering result is to divide each phoneme node in the second recognition model into N c cluster classes.
[0073] It should be noted that in the embodiments of the present invention, each phoneme node in the second recognition model can also be divided into other numbers of cluster classes according to actual needs, and the embodiments of the present invention do not make specific limitations on this.
[0074] Based on any of the above embodiments, the second recognition model includes a feature extraction layer and phoneme classification layers corresponding to multiple languages, and each phoneme classification layer is constructed based on the phoneme nodes corresponding to each language. Figure 3 is a schematic flowchart of the second recognition model training method provided by the present invention. As Figure 3 shown, the training steps of the second recognition model include:
[0075] Step 310: Input the sample speech of each language into the feature extraction layer of the second recognition model to obtain the first phoneme hidden layer feature output by the feature extraction layer of the second recognition model;
[0076] Step 320: Input the first phoneme hidden layer feature into the phoneme classification layer of each language to obtain the first phoneme prediction result output by the phoneme classification layer of each language;
[0077] Step 330: Based on the difference between the phoneme-level label and the first phoneme prediction result, perform parameter iteration on the feature extraction layer of the second recognition model and the phoneme classification layer of each language to obtain the second recognition model.
[0078] Specifically, the first phoneme hidden layer feature is used to represent the feature information of each phoneme in the sample speech, and it can be understood as the frame-level hidden layer feature. The feature extraction layer of the second recognition model is used to extract the first phoneme hidden layer feature corresponding to the sample speech of each language. Among them, the feature extraction layer in the second recognition model is shared, that is, the sample speech of each language can be feature-extracted by this feature extraction layer. In addition, the feature extraction layer of the second recognition model can use neural network models such as DNN (Deep Neural Network), RNN (Recurrent Neural Network), or CNN (Convolution Neural Network) to extract the first phoneme hidden layer feature, and the embodiments of the present invention do not make specific limitations on this.
[0079] In addition, the second recognition model further includes phoneme classification layers corresponding to multiple languages respectively, that is, the phoneme classification layers corresponding to each language are independent of each other. Thus, the phoneme classification layers corresponding to each language can independently learn the phoneme information of the corresponding language, avoiding the influence of pronunciation conflicts between different languages on phoneme recognition, and then accurately recognizing the phonemes in that language to obtain the first phoneme prediction result.
[0080] After obtaining the first phoneme prediction result, based on the difference between the phoneme-level label and the first phoneme prediction result, parameter iteration is performed on the feature extraction layer of the second recognition model and the phoneme classification layers of each language, so that the second recognition model can learn the information of different categories of phonemes in each language as much as possible during the training process, thereby enabling the second recognition model to accurately recognize the phonemes in each language.
[0081] It can be seen that based on the difference between the phoneme-level label and the first phoneme prediction result, parameter iteration is performed on the feature extraction layer of the second recognition model and the phoneme classification layers of each language in the embodiments of the present invention, which can enable the trained second recognition model to accurately recognize the phonemes in each language.
[0082] As an alternative embodiment, the feature extraction layer in the second recognition model may include a first encoding layer and a first attention layer. The first encoding layer is used to encode the sample speech of each language to obtain the first encoding features of each sample speech. The first attention layer is used to perform an attention transformation on the first encoding features of each sample speech based on the attention mechanism to obtain the first phoneme hidden layer features. In addition, the phoneme classification layers of each language in the second recognition model may include a first decoding layer and a first recognition layer. The first decoding layer is used to decode the first phoneme hidden layer features to obtain the first decoding features of each sample speech. The first recognition layer is used to perform phoneme recognition based on the first decoding features of each sample speech to obtain the first phoneme prediction result.
[0083] Based on any of the above embodiments, Figure 4 is a schematic flowchart of another second recognition model training method provided by the present invention. As Figure 4 shown, the training steps of the second recognition model include:
[0084] Step 410: After obtaining the first phoneme hidden layer features output by the feature extraction layer of the second recognition model, determine the word-level hidden layer features and / or the sentence-level hidden layer features based on the first phoneme hidden layer features;
[0085] Step 420: Based on the difference between the word-level label of the sample speech and the word-level prediction result and / or the difference between the language label of the sample speech and the language prediction result, perform parameter iteration on the feature extraction layer of the second recognition model to obtain the second recognition model. The word-level prediction result is determined based on the word-level hidden layer features, and the language prediction result is determined based on the sentence-level hidden layer features.
[0086] Specifically, the character-level hidden layer features are used to characterize the feature information of each character in the sample speech. Since each character is constructed based on multiple phonemes, when determining the character-level hidden layer features, it is necessary to determine them based on multiple first phoneme hidden layer features corresponding to the character. The sentence-level hidden layer features are used to characterize the feature information of each clause in the sample speech. Since each clause is constructed based on multiple characters, when determining the sentence-level hidden layer features, it is necessary to determine them based on multiple character-level hidden layer features corresponding to the clause.
[0087] Based on the difference between the character-level label and the character-level prediction result of the sample speech, when performing parameter iteration on the feature extraction layer of the second recognition model, the feature extraction layer of the second recognition model can learn different phoneme information in each language from the character level, and thus can accurately recognize different phonemes from the character level.
[0088] Based on the difference between the language label and the language prediction result of the sample speech, when performing parameter iteration on the feature extraction layer of the second recognition model, the feature extraction layer of the second recognition model can learn different phoneme information in each language from the sentence level, and thus can accurately recognize different phonemes from the sentence level.
[0089] It can be seen that based on the difference between the character-level label and the character-level prediction result of the sample speech and / or the difference between the language label and the language prediction result of the sample speech, performing parameter iteration on the feature extraction layer of the second recognition model can enable the second recognition model to also accurately recognize different phonemes from the character level and / or sentence level with a larger granularity, further improving the phoneme recognition effect of the second recognition model.
[0090] Based on any of the above embodiments, Figure 5 is a schematic flowchart of an implementation manner of step 420 in another second recognition model training method provided by the present invention, as Figure 5 shown, step 420 includes:
[0091] Step 421, inputting the character-level hidden layer features into the character-level classification layer to obtain the character-level prediction result output by the character-level classification layer, and / or inputting the sentence-level hidden layer features into the language classification layer to obtain the language prediction result output by the language classification layer;
[0092] Step 422, based on the difference between the character-level label and the character-level prediction result and / or the difference between the language label and the language prediction result, performing parameter iteration on the feature extraction layer of the second recognition model to obtain the second recognition model.
[0093] Specifically, the character-level classification layer is used to determine the character-level prediction result based on the character-level hidden layer features, and the language classification layer is used to determine the language prediction result based on the sentence-level hidden layer features. Among them, the character-level prediction result can be understood as the prediction result of each character in the sample speech, and the language prediction result can be understood as the language prediction result of each clause in the sample speech.
[0094] Optionally, embodiments of the present invention can perform parameter iteration on the feature extraction layer of the second recognition model based on the difference between the character-level label and the character-level prediction result of the sample speech, so that the feature extraction layer of the second recognition model can learn different phoneme information under each language from the character level, and thus can accurately recognize different phonemes from the character level.
[0095] Optionally, embodiments of the present invention can perform parameter iteration on the feature extraction layer of the second recognition model based on the difference between the language label and the language prediction result of the sample speech, so that the feature extraction layer of the second recognition model can learn different phoneme information under each language from the sentence level, and thus can accurately recognize different phonemes from the sentence level.
[0096] Optionally, embodiments of the present invention can perform parameter iteration on the feature extraction layer of the second recognition model based on the difference between the character-level label and the character-level prediction result of the sample speech and the difference between the language label and the language prediction result of the sample speech, so that the feature extraction layer of the second recognition model can learn different phoneme information under each language from the character level and the sentence level, and thus can accurately recognize different phonemes from the character level and the sentence level.
[0097] It should be noted that the character-level classification layer and the sentence-level classification layer can be set in the auxiliary model, that is, the second recognition model does not include the character-level classification layer and the sentence-level classification layer. This auxiliary model is used to assist in training the second recognition model from the character level and / or sentence level with a larger granularity, so that the second recognition model can further accurately recognize different phonemes from the character level and / or sentence level and improve the phoneme recognition effect.
[0098] As an optional embodiment, the auxiliary model may include a character-level feature extraction layer, a character-level classification layer, a sentence-level feature extraction layer, and a language classification layer. Among them, the character-level feature extraction layer is used to perform a sliding window on the first phoneme hidden layer features to obtain character-level hidden layer features. The character-level classification layer is used to perform character recognition based on the character-level hidden layer features to obtain a character-level prediction result. The sentence-level feature extraction layer is used to perform pooling on the character-level hidden layer features output by the character-level feature extraction layer to obtain sentence-level hidden layer features. The language classification layer is used to perform language recognition based on the sentence-level hidden layer features to obtain a language prediction result.
[0099] Based on any of the above embodiments, determining the character-level hidden layer features and / or the sentence-level hidden layer features based on the first phoneme hidden layer features includes:
[0100] Perform a sliding window on the first phoneme hidden layer features to obtain word-level hidden layer features;
[0101] Perform pooling on the word-level hidden layer features to obtain sentence-level hidden layer features.
[0102] Specifically, since the granularity of character recognition based on word-level hidden layer features is greater than that of phoneme recognition based on the first phoneme hidden layer features, it is necessary to perform a sliding window operation on the first phoneme hidden layer features. For example, the window length can be set to B, and each time B frames of the first phoneme hidden layer features are sent into the neural network. After the neural network abstracts the word-level hidden layer features, the word-level hidden layer features are then sent into the word-level classification layer to obtain the word-level prediction result. The granularity of language recognition based on sentence-level hidden layer features is larger than that of character recognition based on word-level hidden layer features. Therefore, the word-level hidden layer features can be used to generate sentence-level hidden layer features through multiple poolings of the neural network, and the sentence-level hidden layer features are input into the language classification layer to obtain the language prediction result.
[0103] Based on any of the above embodiments, Figure 6 is a schematic flowchart of the method for training a phoneme recognition model provided by the present invention. As Figure 6 shown, the training steps of the phoneme recognition model include:
[0104] Step 610: Fix the parameters of the feature extraction layer of the first recognition model;
[0105] Step 620: Input the sample voices of each language into the feature extraction layer of the first recognition model to obtain the second phoneme hidden layer features output by the feature extraction layer of the first recognition model;
[0106] Step 630: Input the second phoneme hidden layer features into the current phoneme classification layer to obtain the second phoneme prediction result output by the current phoneme classification layer; the current phoneme classification layer is constructed based on the phoneme nodes selected from the second recognition model;
[0107] Step 640: Based on the difference between the phoneme-level labels and the second phoneme prediction result, perform parameter iteration on the current phoneme classification layer to obtain the phoneme recognition model.
[0108] Specifically, the feature extraction layer of the first recognition model is the feature extraction layer of the trained second recognition model. Since the trained second recognition model has the ability to accurately extract features, after obtaining the first recognition model and fixing the parameters of the feature extraction layer of the first recognition model, the first recognition model can retain the feature extraction ability of the trained second recognition model.
[0109] The feature extraction layer of the first recognition model is used to extract the second phoneme hidden layer features corresponding to the sample voices of each language. Since the feature extraction layer of the first recognition model is the feature extraction layer of the second recognition model that has been trained, the feature extraction layer of the first recognition model can accurately extract the second phoneme hidden layer features.
[0110] The current phoneme classification layer is constructed based on the phoneme nodes screened from the second recognition model. That is, the current phoneme classification layer filters out redundant phoneme nodes on the basis of the phoneme classification layer of the second recognition model. That is, the number of phoneme nodes in the current phoneme classification layer is less than the number of phoneme nodes in the phoneme classification layer of the second recognition model. Similarly, the current phoneme classification layer is used to perform phoneme recognition based on the second phoneme hidden layer features to obtain the second phoneme prediction result.
[0111] After obtaining the second phoneme prediction result, based on the difference between the phoneme-level label and the second phoneme prediction result, parameter iteration is performed on the current phoneme classification layer, so that the phoneme recognition model can learn as much information about different categories of phonemes in each language as possible during training, so that the phoneme recognition model can accurately identify the phonemes in each language.
[0112] It can be seen that based on the difference between the phoneme-level label and the second phoneme prediction result, the present invention embodiment performs parameter iteration on the current phoneme classification layer, which can enable the trained phoneme recognition model to accurately identify the phonemes in each language.
[0113] As an alternative embodiment, the feature extraction layer in the first recognition model may include a second encoding layer and a second attention layer. The second encoding layer is used to encode the sample voices of each language to obtain the second encoding features of each sample voice. The second attention layer is used to perform an attention transformation on the second encoding features of each sample voice based on the attention mechanism to obtain the second phoneme hidden layer features. In addition, the current phoneme classification layer in the second recognition model may include a second decoding layer and a second recognition layer. The second decoding layer is used to decode the second phoneme hidden layer features to obtain the second decoding features of each sample voice. The second recognition layer is used to perform phoneme recognition based on the second decoding features of each sample voice to obtain the second phoneme prediction result.
[0114] Based on any of the above embodiments, the phoneme recognition model is trained based on the sample voices of multiple languages and the phoneme-level labels of each sample voice. The first recognition model is obtained by screening the phoneme nodes in the second recognition model based on the similarity between the phonemes corresponding to each phoneme node in the second recognition model. The present invention also provides a method for training a phoneme recognition model, and the method includes:
[0115] Figure 7 is a schematic flowchart of still another method for training the second recognition model provided by the present invention, asFigure 7 As shown, the second recognition model includes a feature extraction layer and a phoneme classification layer corresponding to each language. Each phoneme classification layer is constructed based on the phoneme nodes corresponding to each language. The second recognition model is trained based on the sample voices of multiple languages and the phoneme-level labels of each sample voice. Specifically: the sample voices of each language are input into the feature extraction layer of the second recognition model to obtain the first phoneme hidden layer features, and the first phoneme hidden layer features are input into the phoneme classification layers of each language to obtain the first phoneme prediction results. Based on the difference between the phoneme-level labels and the first phoneme prediction results, the feature extraction layer of the second recognition model and the phoneme classification layers of each language are iteratively parameterized to obtain the initial second recognition model.
[0116] After obtaining the initial second recognition model, based on the first phoneme hidden layer features, the word-level hidden layer features and the sentence-level hidden layer features are determined. The word-level hidden layer features are input into the word-level classification layer of the auxiliary model to obtain the word-level prediction results, and the sentence-level hidden layer features are input into the language classification layer of the auxiliary model to obtain the language prediction results. Based on the difference between the word-level labels and the word-level prediction results and the difference between the language labels and the language prediction results, the feature extraction layer of the initial second recognition model is iteratively parameterized to obtain the second recognition model.
[0117] Next, based on the similarity between the phonemes corresponding to each phoneme node, the phoneme nodes under the second recognition model are clustered to obtain multiple clusters, and any one phoneme node in each cluster is retained and the other phoneme nodes in each cluster are deleted to obtain the first recognition model.
[0118] After obtaining the first recognition model, the parameters of the feature extraction layer of the first recognition model are fixed. The sample voices of each language are input into the feature extraction layer of the first recognition model to obtain the second phoneme hidden layer features, and the second phoneme hidden layer features are input into the current phoneme classification layer to obtain the second phoneme prediction results; where the current phoneme classification layer is constructed based on the phoneme nodes selected from the second recognition model.
[0119] Finally, based on the difference between the phoneme-level labels and the second phoneme prediction results, the current phoneme classification layer is iteratively parameterized to obtain the phoneme recognition model, which not only has a small scale but also can accurately distinguish the phonemes of different languages.
[0120] The phoneme recognition device provided by the present invention will be described below. The phoneme recognition device described below can be mutually referred to the phoneme recognition method described above.
[0121] Based on any of the above embodiments, Figure 8 is a schematic structural diagram of the phoneme recognition device provided by the present invention, as Figure 8 shown, the device includes:
[0122] A determination unit 810, configured to determine the speech to be recognized;
[0123] An identification unit 820, configured to input the speech to be recognized into a phoneme recognition model, and obtain a phoneme recognition result output by the phoneme recognition model;
[0124] The phoneme recognition model is obtained by training a first recognition model based on sample speeches of multiple languages and phoneme-level labels of each sample speech. The first recognition model is obtained by screening phoneme nodes in the second recognition model based on the similarity between phonemes corresponding to each phoneme node in the second recognition model. The second recognition model includes phoneme nodes corresponding to multiple languages respectively.
[0125] Based on any of the above embodiments, the apparatus further includes:
[0126] A clustering unit, configured to cluster each phoneme node in the second recognition model based on the similarity between phonemes corresponding to each phoneme node, and obtain multiple clusters;
[0127] A pruning unit, configured to screen a current phoneme node from the phoneme nodes in each cluster, and delete other phoneme nodes in each cluster except the current phoneme node, to obtain the first recognition model.
[0128] Based on any of the above embodiments, the second recognition model includes a feature extraction layer and phoneme classification layers corresponding to multiple languages respectively. Each phoneme classification layer is constructed based on phoneme nodes corresponding to each language;
[0129] The apparatus further includes:
[0130] A first feature extraction unit, configured to input the sample speeches of each language into the feature extraction layer of the second recognition model, and obtain a first phoneme hidden layer feature output by the feature extraction layer of the second recognition model;
[0131] A first phoneme classification unit, configured to input the first phoneme hidden layer feature into the phoneme classification layers of each language, and obtain a first phoneme prediction result output by the phoneme classification layers of each language;
[0132] A first parameter iteration unit, configured to perform parameter iteration on the feature extraction layer of the second recognition model and the phoneme classification layers of each language based on the difference between the phoneme-level label and the first phoneme prediction result, to obtain the second recognition model.
[0133] Based on any of the above embodiments, the apparatus further includes:
[0134] A feature determination unit, configured to, after obtaining first phoneme hidden layer features output by a feature extraction layer of the second recognition model, determine word-level hidden layer features and / or sentence-level hidden layer features based on the first phoneme hidden layer features;
[0135] A second parameter iteration unit, configured to perform parameter iteration on the feature extraction layer of the second recognition model based on a difference between a word-level label and a word-level prediction result of the sample speech and / or a difference between a language label and a language prediction result of the sample speech, to obtain the second recognition model; the word-level prediction result is determined based on the word-level hidden layer features, and the language prediction result is determined based on the sentence-level hidden layer features.
[0136] Based on any of the above embodiments, the second parameter iteration unit includes:
[0137] An auxiliary prediction unit, configured to input the word-level hidden layer features into a word-level classification layer to obtain the word-level prediction result output by the word-level classification layer, and / or input the sentence-level hidden layer features into a language classification layer to obtain the language prediction result output by the language classification layer;
[0138] An auxiliary training unit, configured to perform parameter iteration on the feature extraction layer of the second recognition model based on a difference between the word-level label and the word-level prediction result and / or a difference between the language label and the language prediction result, to obtain the second recognition model.
[0139] Based on any of the above embodiments, the feature determination unit includes:
[0140] A sliding window unit, configured to perform sliding window on the first phoneme hidden layer features to obtain the word-level hidden layer features;
[0141] A pooling unit, configured to perform pooling on the word-level hidden layer features to obtain the sentence-level hidden layer features.
[0142] Based on any of the above embodiments, the apparatus further includes:
[0143] A parameter fixing unit, configured to fix parameters of the feature extraction layer of the first recognition model;
[0144] A second feature extraction unit, configured to input sample speeches of each language into the feature extraction layer of the first recognition model to obtain second phoneme hidden layer features output by the feature extraction layer of the first recognition model;
[0145] A second phoneme classification unit, configured to input the second phoneme hidden layer features into a current phoneme classification layer to obtain a second phoneme prediction result output by the current phoneme classification layer; the current phoneme classification layer is constructed based on phoneme nodes screened from the second recognition model;
[0146] A third parameter iteration unit, configured to perform parameter iteration on the current phoneme classification layer based on the difference between the phoneme-level label and the second phoneme prediction result, so as to obtain the phoneme recognition model.
[0147] Figure 9 is a schematic structural diagram of an electronic device provided by the present invention, as Figure 9 shown. The electronic device may include: a processor 910, a memory 920, a communication interface 930, and a communication bus 940. Among them, the processor 910, the memory 920, and the communication interface 930 communicate with each other through the communication bus 940. The processor 910 may call logic instructions in the memory 920 to execute a phoneme recognition method, which includes: determining a speech to be recognized; inputting the speech to be recognized into a phoneme recognition model to obtain a phoneme recognition result output by the phoneme recognition model; the phoneme recognition model is obtained by training a first recognition model based on sample speeches in multiple languages and phoneme-level labels of each sample speech, the first recognition model is obtained by screening phoneme nodes in the second recognition model based on the similarity between phonemes corresponding to each phoneme node in the second recognition model, and the second recognition model includes phoneme nodes corresponding to multiple languages respectively.
[0148] In addition, when the logic instructions in the above-mentioned memory 920 are implemented in the form of software function units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0149] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the phoneme recognition method provided by each of the above methods, and the method includes: determining the speech to be recognized; inputting the speech to be recognized into a phoneme recognition model to obtain a phoneme recognition result output by the phoneme recognition model; the phoneme recognition model is obtained by training a first recognition model based on sample speeches in multiple languages and phoneme-level labels of each sample speech, and the first recognition model is obtained by screening phoneme nodes in the second recognition model based on the similarity between phonemes corresponding to each phoneme node in the second recognition model, and the second recognition model includes phoneme nodes corresponding to multiple languages respectively.
[0150] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the phoneme recognition method provided by each of the above, and the method includes: determining the speech to be recognized; inputting the speech to be recognized into a phoneme recognition model to obtain a phoneme recognition result output by the phoneme recognition model; the phoneme recognition model is obtained by training a first recognition model based on sample speeches in multiple languages and phoneme-level labels of each sample speech, and the first recognition model is obtained by screening phoneme nodes in the second recognition model based on the similarity between phonemes corresponding to each phoneme node in the second recognition model, and the second recognition model includes phoneme nodes corresponding to multiple languages respectively.
[0151] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0152] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A phoneme recognition method, characterized in that, Including: Determine the speech to be recognized; Input the speech to be recognized into a phoneme recognition model to obtain a phoneme recognition result output by the phoneme recognition model; The phoneme recognition model is obtained by training a first recognition model based on sample speeches of multiple languages and phoneme-level labels of each sample speech. The first recognition model is obtained by screening phoneme nodes in the second recognition model based on the similarity between phonemes corresponding to each phoneme node. The second recognition model includes phoneme nodes corresponding to multiple languages respectively; The second recognition model includes a feature extraction layer, and the second recognition model is obtained by training based on the following steps: Input the sample speeches of each language into the feature extraction layer of the second recognition model to obtain first phoneme hidden layer features output by the feature extraction layer of the second recognition model; Based on the first phoneme hidden layer features, determine word-level hidden layer features and / or sentence-level hidden layer features; Based on the difference between the word-level label and the word-level prediction result of the sample speech and / or the difference between the language label and the language prediction result of the sample speech, perform parameter iteration on the feature extraction layer of the second recognition model to obtain the second recognition model; the word-level prediction result is determined based on the word-level hidden layer features, and the language prediction result is determined based on the sentence-level hidden layer features.
2. The phoneme recognition method according to claim 1, wherein The determination steps of the first recognition model include: Based on the similarity between phonemes corresponding to each phoneme node, cluster the phoneme nodes in the second recognition model to obtain multiple clusters; Select the current phoneme node from the phoneme nodes in each cluster, and delete other phoneme nodes in each cluster except the current phoneme node to obtain the first recognition model.
3. The phoneme recognition method according to claim 1, characterized in that, The second recognition model includes a feature extraction layer and phoneme classification layers corresponding to multiple languages respectively. Each phoneme classification layer is constructed based on the phoneme nodes corresponding to each language; The second recognition model is obtained by training based on the following steps: Input the sample speeches of each language into the feature extraction layer of the second recognition model to obtain first phoneme hidden layer features output by the feature extraction layer of the second recognition model; Input the first phoneme hidden layer features into the phoneme classification layers of each language to obtain first phoneme prediction results output by the phoneme classification layers of each language; Based on the difference between the phoneme-level label and the first phoneme prediction result, perform parameter iteration on the feature extraction layer of the second recognition model and the phoneme classification layers of each language to obtain the second recognition model.
4. The phoneme recognition method according to claim 1, wherein The performing parameter iteration on the feature extraction layer of the second recognition model based on the difference between the word-level label and the word-level prediction result of the sample speech and / or the difference between the language label and the language prediction result of the sample speech to obtain the second recognition model includes: Input the word-level hidden layer features into a word-level classification layer to obtain the word-level prediction result output by the word-level classification layer, and / or input the sentence-level hidden layer features into a language classification layer to obtain the language prediction result output by the language classification layer; Based on the difference between the character-level label and the character-level prediction result and / or the difference between the language label and the language prediction result, perform parameter iteration on the feature extraction layer of the second recognition model to obtain the second recognition model.
5. The phoneme recognition method according to claim 1, wherein The determining the character-level hidden layer feature and / or the sentence-level hidden layer feature based on the first phoneme hidden layer feature includes: Perform a sliding window on the first phoneme hidden layer feature to obtain the character-level hidden layer feature; Perform pooling on the character-level hidden layer feature to obtain the sentence-level hidden layer feature.
6. The phoneme recognition method according to claim 3, wherein The phoneme recognition model is trained based on the following steps: Fix the parameters of the feature extraction layer of the first recognition model; Input the sample speech of each language into the feature extraction layer of the first recognition model to obtain the second phoneme hidden layer feature output by the feature extraction layer of the first recognition model; Input the second phoneme hidden layer feature into the current phoneme classification layer to obtain the second phoneme prediction result output by the current phoneme classification layer; the current phoneme classification layer is constructed based on the phoneme nodes selected from the second recognition model; Based on the difference between the phoneme-level label and the second phoneme prediction result, perform parameter iteration on the current phoneme classification layer to obtain the phoneme recognition model.
7. A phoneme recognition device, characterized in that, Including: A determining unit, configured to determine the speech to be recognized; A recognition unit, configured to input the speech to be recognized into the phoneme recognition model to obtain the phoneme recognition result output by the phoneme recognition model; The phoneme recognition model is trained by using the sample speech of multiple languages and the phoneme-level labels of each sample speech to train the first recognition model. The first recognition model is obtained by screening the phoneme nodes in the second recognition model based on the similarity between the phonemes corresponding to the phoneme nodes in the second recognition model. The second recognition model includes phoneme nodes corresponding to multiple languages respectively; The second recognition model includes a feature extraction layer. The second recognition model is trained based on the following steps: Input the sample speech of each language into the feature extraction layer of the second recognition model to obtain the first phoneme hidden layer feature output by the feature extraction layer of the second recognition model; Based on the first phoneme hidden layer feature, determine the character-level hidden layer feature and / or the sentence-level hidden layer feature; Based on the difference between the character-level label of the sample speech and the character-level prediction result and / or the difference between the language label of the sample speech and the language prediction result, perform parameter iteration on the feature extraction layer of the second recognition model to obtain the second recognition model; the character-level prediction result is determined based on the character-level hidden layer feature, and the language prediction result is determined based on the sentence-level hidden layer feature.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the phoneme recognition method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the phoneme recognition method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-language speech recognition method based on language type and speech content collaborative classification
CN110895932A
Small data learning identification method for unknown language end-side command word
CN114420101A