Language recognition method, device, electronic device and storage medium

By using a backbone network and multiple data volume classification sample sets to train the model in language recognition, combined with posterior probability and mapping relationships, the problem of poor recognition effect under unbalanced language data distribution is solved, and high-accuracy recognition of languages ​​with a small distribution ratio is achieved.

CN114512116BActive Publication Date: 2025-09-19HEFEI IFLY DIGITAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210061164.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-19
Publication Date
2025-09-19
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

When the distribution of language data is unbalanced, existing language recognition methods cannot accurately identify languages ​​with a relatively small distribution, resulting in poor recognition effect and low accuracy.

Method used

By extracting language features based on the backbone network and using the full sample set and multiple data classification sample sets to train multiple language recognition models, the final language recognition result is determined by combining the posterior probabilities and language mapping relationships of different categories, reducing the impact of uneven data distribution on recognition.

Benefits of technology

The recognition rate of languages ​​with a small distribution ratio has been improved, the accuracy of language recognition has been improved, and the recognition effect has been further improved through verification of multiple model results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114512116B_ABST
    Figure CN114512116B_ABST
Patent Text Reader

Abstract

The present invention provides a language recognition method, device, electronic device and storage medium, wherein the method includes: extracting language features of speech to be recognized based on a backbone network; determining a first recognition result of the language features based on a full sample set; and / or determining a second recognition result of the language features based on multiple data volume classification sample sets; and determining a language recognition result based on the first recognition result and / or the second recognition result. The method, device, electronic device and storage medium provided by the present invention can determine the first recognition result through a mapping relationship obtained by training a full sample set with balanced speech distribution, and determine the second recognition result through a mapping relationship obtained by training multiple data volume classification sample sets, thereby improving the classification ability of language recognition and thus improving the recognition rate of language, and can determine the language recognition result by combining the first recognition result with the second recognition result, further improving the accuracy of language recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice interaction technology, and in particular to a language recognition method, device, electronic device and storage medium. Background Art

[0002] Language identification, also known as language recognition, refers to the process by which a machine automatically determines the language of a speech segment. After decades of development, language recognition technology has demonstrated tremendous application value and potential, and has been widely adopted.

[0003] The current mainstream language recognition methods work well and can reach a usable level when the language data is evenly distributed (for example, there are C languages, and each language accounts for around 1 / C). However, in real-world scenarios, balanced language distribution is rare, and more common are scenarios where the language distribution is uneven. Even in special circumstances, such as when there are many languages ​​and some languages ​​account for very small proportions, the recognition effect of current language recognition methods is significantly reduced, and they are unable to accurately identify languages ​​with a small distribution proportion. Summary of the Invention

[0004] The present invention provides a language recognition method, device, electronic device and storage medium to address the defects of the language recognition method in the prior art, such as poor recognition effect and low accuracy for languages ​​with a small distribution ratio when the distribution of language data is unbalanced.

[0005] The present invention provides a language recognition method, comprising:

[0006] Based on the backbone network, the language features of the speech to be recognized are extracted;

[0007] Determining a first recognition result of the language feature based on a full sample set, wherein the full sample set includes first sample speech of all languages, and the first sample speech of all languages ​​is evenly distributed;

[0008] and / or, determining a second recognition result of the language feature based on a plurality of data volume classification sample sets, wherein each data volume classification sample set includes a second sample speech of a language corresponding to a data volume category, and the plurality of data volume classification sample sets are obtained based on the data volume of the second sample speech of the full language;

[0009] A language recognition result is determined based on the first recognition result and / or the second recognition result.

[0010] According to a language recognition method provided by the present invention, determining a first recognition result of the language feature based on the full sample set includes:

[0011] Based on the first language recognition model, classify the language features into different languages ​​to obtain the first recognition result;

[0012] The first language recognition model is a first classification layer in a first classification network trained based on the full sample set, and the first classification network includes the backbone network and the first classification layer.

[0013] According to a language recognition method provided by the present invention, determining the second recognition result of the language feature based on multiple data volume classification sample sets includes:

[0014] Classifying the language features by data volume category to obtain a data volume category classification result of the language features;

[0015] performing language classification on the language feature based on second language recognition models corresponding to a plurality of data volume categories, thereby obtaining language classification results for the language feature under the plurality of data volume categories, wherein the second language recognition models corresponding to the plurality of data volume categories are trained based on the plurality of data volume classification sample sets;

[0016] The second recognition result is determined based on the data volume category classification result and the language classification results under the multiple data volume categories.

[0017] According to a language identification method provided by the present invention, determining the second identification result based on the data volume category classification result and the language classification results under the multiple data volume categories includes:

[0018] Determining a partial language recognition result for any data volume category based on the posterior probability of any data volume category in the data volume category classification results and the language classification result corresponding to the any data volume category;

[0019] The second recognition result is obtained based on the partial language recognition results of each data volume category.

[0020] According to a language recognition method provided by the present invention, the second language recognition models corresponding to the multiple data volume categories are trained based on the following steps:

[0021] determining a second classification network, wherein the second classification network includes the backbone network and a second classification layer;

[0022] Based on any data volume classification sample set, the second classification network is trained, and the second classification layer in the trained second classification network is used as the second language recognition model for the data volume category corresponding to the any data volume classification sample set.

[0023] According to a language recognition method provided by the present invention, determining a language recognition result based on the first recognition result and / or the second recognition result includes:

[0024] Determining a third recognition result of the language feature based on a plurality of feature classification sample sets, wherein each feature classification sample set includes a third sample speech of a language corresponding to a feature category, and the plurality of feature classification sample sets are obtained based on the language feature classification of the sample speech of the full set of languages;

[0025] The language recognition result is determined based on the first recognition result and / or the second recognition result, and the third recognition result.

[0026] According to a language recognition method provided by the present invention, determining the third recognition result of the language feature based on multiple feature classification sample sets includes:

[0027] Performing feature category classification on the language features to obtain feature category classification results of the language features;

[0028] performing language classification on the language features based on third language recognition models corresponding to a plurality of feature categories to obtain language classification results for the language features under the plurality of feature categories, wherein the third language recognition models corresponding to the plurality of feature categories are trained based on the plurality of feature classification sample sets;

[0029] The third recognition result is determined based on the feature category classification result and the language classification results under the multiple feature categories.

[0030] According to a language recognition method provided by the present invention, the third recognition result is determined based on the feature category classification result and the language classification results under the multiple feature categories, including:

[0031] determining a partial language recognition result for any feature category based on the posterior probability of any feature category in the feature category classification results and the language classification result corresponding to the any feature category;

[0032] The third recognition result is obtained based on the partial language recognition results of each feature category.

[0033] According to a language recognition method provided by the present invention, the third language recognition models corresponding to the plurality of feature categories are trained based on the following steps:

[0034] Determining a third classification network, wherein the third classification network includes the backbone network and a third classification layer;

[0035] Based on any feature classification sample set, the third classification network is trained, and the third classification layer in the trained third classification network is used as the third language recognition model for the feature category corresponding to the any feature classification sample set.

[0036] The present invention also provides a language recognition device, comprising: a feature determination module for extracting language features of a speech to be recognized based on a backbone network;

[0037] A language recognition module is configured to determine a first recognition result of the language feature based on a full sample set, wherein the full sample set includes first sample speech of all languages, and the first sample speech of all languages ​​is evenly distributed;

[0038] and / or, determining a second recognition result of the language feature based on a plurality of data volume classification sample sets, wherein each data volume classification sample set includes a second sample speech of a language corresponding to a data volume category, and the plurality of data volume classification sample sets are obtained based on the data volume of the second sample speech of the full language;

[0039] A result determination module is configured to determine a language recognition result based on the first recognition result and / or the second recognition result.

[0040] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above-described language recognition methods when executing the program.

[0041] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-mentioned language recognition methods.

[0042] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned language recognition methods are implemented.

[0043] The language recognition method, device, electronic device and storage medium provided by the present invention determine a first recognition result of a language feature by mapping the relationship between a full set of speech samples with balanced distribution obtained through training and the language recognition results, thereby reducing the situation where the language with a large distribution ratio dominates the network parameter training during random sampling, resulting in poor recognition of languages ​​with a small distribution ratio, thereby improving the recognition rate of languages ​​with a small distribution ratio; determine a second recognition result by mapping the relationship between multiple data volume classification sample sets obtained through training and the language recognition results, thereby reducing the situation where the language with a small distribution ratio due to an imbalance in the amount of training sample speech data leads to poor recognition of languages ​​with a small distribution ratio, thereby improving the recognition rate of languages ​​with a small distribution ratio in the amount of training sample speech data; and determine the language recognition result by combining the first recognition result with the second recognition result, thereby realizing mutual verification of the two results and further improving the accuracy of language recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0045] Figure 1 1 is a flow chart of the language identification method provided by the present invention;

[0046] Figure 2 This is one of the flow charts of the method for obtaining the second recognition result provided by the present invention;

[0047] Figure 3 This is the second flow chart of the method for obtaining the second recognition result provided by the present invention;

[0048] Figure 4 This is a flow chart of the second language recognition model training method corresponding to the data volume category provided by the present invention;

[0049] Figure 5 This is a flow chart of the method for obtaining language recognition results provided by the present invention;

[0050] Figure 6 It is a structural diagram of the backbone network training provided by the present invention;

[0051] Figure 7 It is a structural diagram of the first classification network training provided by the present invention;

[0052] Figure 8 It is a structural diagram of the second classification network training provided by the present invention;

[0053] Figure 9 It is a structural diagram of the third classification network training provided by the present invention;

[0054] Figure 10 It is a structural diagram of the language recognition device provided by the present invention;

[0055] Figure 11 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0056] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0057] Taking the TV (Total Variability) system as an example, the current mainstream language recognition method, from a technical perspective, has two crucial steps in its training phase: training the UBM (background model) and T (factor orthogonal space), and training the LDA (linear transform space). UBM and T training yield corresponding models that map SDC (Shifted Delta Cepstra) features into i-vectors of equal dimension. When the data distribution is imbalanced, due to the limited training data for the minority class, the learned UBM and T parameters tend to be biased toward the majority class. LDA training utilizes the original labeled i-vectors to generate a reduced dimensionality matrix, minimizing the distance between data in the same language and maximizing the distance between data in different languages. When the data distribution is imbalanced, the resulting reduced dimensionality matrix will also be biased toward the majority class. These two training steps significantly degrade language recognition performance in imbalanced data scenarios, rendering it unusable for languages ​​with a minority class.

[0058] Therefore, how to accurately identify the languages ​​​​including those with a small distribution ratio in the case of an imbalance in the proportion of language data is a technical problem that needs to be solved urgently.

[0059] In view of the above situation, an embodiment of the present invention provides a language identification method. Figure 1 Schematic diagram of the flow of the language recognition method provided by the present invention. Figure 1 As shown, the method includes:

[0060] Step 110: extracting language features of the speech to be recognized based on the backbone network;

[0061] Specifically, the language features of the speech to be recognized are obtained by inputting the speech to be recognized into the backbone network, extracting the features, and outputting them. The speech to be recognized can contain one or more languages, which is not limited in this embodiment of the present invention.

[0062] It should be noted that the initial backbone network and the linear fully connected layer are concatenated, and based on the samples collected in the sample set of all languages, the concatenated initial backbone network and the linear fully connected layer are trained by randomly selecting samples from the sample set of all languages. After the concatenated initial backbone network and the linear fully connected layer converge, the converged initial backbone network becomes the backbone network.

[0063] Step 120: determining a first recognition result of the language feature based on the full sample set; the full sample set includes first sample speech of all languages, and the first sample speech of all languages ​​is evenly distributed;

[0064] and / or, determining a second recognition result of the language feature based on a plurality of data volume classification sample sets, wherein each data volume classification sample set includes a second sample speech of a language corresponding to a data volume category, and the plurality of data volume classification sample sets are obtained based on the data volume division of the second sample speech of all languages;

[0065] Step 130: Determine a language recognition result based on the first recognition result and / or the second recognition result.

[0066] Specifically, the first recognition result is obtained by performing language classification on the language features of the speech to be recognized. The language classification on the language features of the speech to be recognized here can be obtained by performing language classification mapping on the language features of the speech to be recognized through a pre-acquired language mapping relationship. The language mapping relationship here can be specifically reflected in the first language recognition model obtained through model training.

[0067] Here, the first language recognition model is trained using the first sample speech of all languages ​​with a balanced speech distribution as the full sample set. The full language refers to each language collected for training based on language recognition requirements, and the full sample set is obtained by selecting an equal amount of sample speech of each language in the full language from a large number of samples. In particular, considering that existing language recognition methods have low recognition rates for languages ​​with a small distribution ratio, using sample speech with a balanced distribution when training the first language recognition model helps improve the recognition rate for languages ​​with a small distribution ratio.

[0068] It should be noted that the distribution balance in the first sample speech distribution balance refers to the balanced distribution of data volumes of various languages ​​in the sample speech. For example, if the sample speech contains three languages ​​A, B, and C, then the sample speech distribution balance means that the data volumes of the sample speech of the three languages ​​A, B, and C in all the sample speech are close to each other in the total sample data volume (e.g., the sample data volumes of A, B, and C account for about 33% of the total sample data volume).

[0069] The second recognition result is obtained by language classification of the language features of the speech to be recognized. The language classification of the language features of the speech to be recognized here can be a pre-acquired data volume category classification mapping relationship and a language mapping relationship corresponding to multiple data volume categories, and based on the data volume category classification mapping relationship and the language mapping relationship corresponding to multiple data volume categories, the language features of the speech to be recognized are language classified and mapped. Among them, the data volume category classification mapping relationship classifies the language features of the speech to be recognized into data volume categories to obtain the data volume classification result of the language features of the speech to be recognized, which can be specifically reflected in the data volume category classification model obtained through model training, and the second language mapping relationships corresponding to multiple data volume categories are respectively obtained by language classification of the language features of the speech to be recognized. The multiple mapping relationships here can specifically be reflected in the second language recognition models corresponding to multiple data volume categories obtained through model training.

[0070] Here, the second language recognition model corresponding to each data volume category is trained based on the second sample set of the language corresponding to that data volume category. The data volume category categorizes the second sample set of the full language based on the sample data volume. For example, data volumes of 101 to 200 items are in one data volume category, while data volumes of 201 to 300 items are in another data volume category. This is not a limitation of the present invention. It is necessary to ensure that the data volumes of the languages ​​within each data volume category are as close as possible. In particular, given that some language samples are difficult to obtain in real-world environments, existing language recognition methods can affect model training due to differences in language sample data volume, leading to low recognition rates for languages ​​with fewer samples. Therefore, classifying the second sample languages ​​of the full language category based on sample data volume and using the second sample languages ​​of the languages ​​within each data volume category to train the second language recognition model corresponding to that data volume category helps reduce the impact of sample data volume ratio on model training, thereby improving the model's language recognition rate.

[0071] It should be noted that the second recognition result can be obtained by first classifying the language features of the speech to be recognized into data volume categories through a data volume category classification model to obtain a data volume category classification result, and then classifying the language features of the speech to be recognized based on the second language recognition model corresponding to the data volume category classification result. It can also be obtained by classifying the language features of the speech to be recognized into data volume categories through a data volume category classification model to obtain a data volume category classification result, and performing language recognition on the language features of the speech to be recognized using second language recognition models corresponding to multiple data volume categories to obtain multiple partial language recognition results, and then obtaining it based on the data volume category classification result and the multiple partial language recognition results. The embodiment of the present invention does not impose any restrictions on this.

[0072] It should be noted that, for the first sample speech and the second sample speech of the same language, "first" and "second" are used to distinguish whether the sample speech belongs to the full sample set or the data volume classification sample set.

[0073] For any language, the content of the first and second sample speech samples of that language may be the same or different, and the data volume of the first and second sample speech samples may be the same or different. Furthermore, the first sample speech may be a portion of the sample speech selected from the sample speech of that language, and the second sample speech may be the entire sample speech of that language.

[0074] In step 130, the first recognition result or the second recognition result can be directly used as the final language recognition result, or the first recognition result can be combined with the second recognition result to obtain the final language recognition result. For example, the first recognition result and the second recognition result can be weighted or the average value of the corresponding language posterior probabilities in the two results can be calculated. This embodiment of the present invention is not limited to this.

[0075] The language recognition method provided by the embodiment of the present invention determines a first recognition result of the language characteristics of the speech to be recognized by mapping the relationship between the full sample set with balanced speech distribution obtained through training and the language recognition result, reduces the situation where the language with a large distribution ratio dominates the network parameter training during random sampling of samples, resulting in poor recognition of the language with a small distribution ratio, and improves the recognition rate of the language with a small distribution ratio. The second recognition result is determined by mapping the relationship between multiple data volume classification sample sets and the language recognition result obtained through training, reduces the situation where the language with a small distribution ratio due to the imbalance of the training sample speech data volume leads to poor recognition of the language with a small sample speech data volume, and improves the recognition rate of the language with a small training sample speech data volume. In addition, the language recognition result can be determined by combining the first recognition result with the second recognition result, so that the two results are mutually verified, further improving the accuracy of language recognition.

[0076] Based on the above embodiment, determining the first recognition result of the language feature based on the full sample set in step 120 includes:

[0077] Based on the first language recognition model, classify the language features into different languages ​​to obtain a first recognition result;

[0078] The first language recognition model is the first classification layer in the first classification network trained based on the full sample set. The first classification network includes a backbone network and the first classification layer.

[0079] Specifically, the training of the first classification network is to fix the parameters of the backbone network, use the full sample set to train the first classification network, and use the first classification layer in the trained first classification network as the first language recognition model. The first language recognition model is used to classify language features and output a first recognition result. The first language recognition model obtained by training with a full sample set with a uniform distribution in the embodiment of the present invention can reduce the situation in which the language with a large distribution proportion dominates the network parameter training during random sampling of samples, resulting in poor recognition of the language with a small distribution proportion, thereby improving the recognition rate of the language with a small distribution proportion.

[0080] Based on the above embodiments, Figure 2 This is one of the flow charts of the method for obtaining the second recognition result provided by the present invention. Figure 2 As shown, the second recognition result of the language feature is determined based on the multiple data volume classification sample sets in step 120, including:

[0081] Step 210, classifying the language features by data volume category to obtain a data volume category classification result of the language features;

[0082] Step 220: Classify the language features based on the second language recognition models corresponding to the multiple data volume categories to obtain language classification results for the language features under the multiple data volume categories. The second language recognition models corresponding to the multiple data volume categories are trained based on the multiple data volume classification sample sets.

[0083] Considering that if the language features of the speech to be recognized are first classified by data volume category to obtain a data volume category classification result, and then the second language recognition model corresponding to the data volume category determined in the data volume category classification result is used to perform language recognition on the language features to obtain a language recognition result, in this case, the language recognition of the language features of the speech to be recognized will be completely dependent on the data volume category classification result. Once the data volume category classification result is incorrect, it will directly lead to the failure of the subsequent language recognition task. Therefore, the embodiment of the present invention simultaneously considers the classification result output by the data volume category classification model and the partial language recognition results output by the second language recognition model corresponding to multiple data volume categories.

[0084] Specifically, step 210 inputs the language features of the speech to be recognized into the data volume category classification model to obtain the data volume category classification result of the language features of the speech to be recognized output by the data volume category classification model.

[0085] In step 220, the language features of the speech to be recognized are input into the second language recognition model corresponding to each data volume category. The second language recognition model corresponding to each data volume category outputs a language classification result for that data volume category. The second language recognition model corresponding to each data volume category is trained using the data volume classification samples corresponding to that data volume category.

[0086] It should be noted that the execution order of step 210 and step 220 is not particular, and step 210 and step 220 can be serial or parallel, and the embodiment of the present invention does not limit this.

[0087] Step 230 : Determine a second recognition result based on the data volume category classification result and the language classification results under the multiple data volume categories.

[0088] It should be noted that the second recognition result can sort the posterior probabilities of each data volume category in the data volume classification result from high to low, select the highest preset number of data volume categories, and then use the language with the highest posterior probability of language classification in the language classification results under these data volume categories as the second recognition result. The posterior probability of each data volume category in the data volume classification result can also be weighted or multiplied with the posterior probabilities of each language classification in the language classification results under the data volume category to obtain the partial language recognition results corresponding to each data volume category, and the partial language recognition results corresponding to each data volume category are combined to obtain the second recognition result. The embodiment of the present invention does not impose any restrictions on this.

[0089] Based on the above embodiments, Figure 3 This is the second flow chart of the method for obtaining the second recognition result provided by the present invention. Figure 3 As shown, step 230 includes:

[0090] Step 231, based on the posterior probability of any data volume category in the data volume category classification results and the language classification result corresponding to the data volume category, determine a partial language recognition result for the data volume category;

[0091] Step 232: Obtain a second recognition result based on the partial language recognition results of each data volume category.

[0092] Considering the language characteristics of the speech to be recognized, when performing language recognition, it is possible that the posterior probability of a certain data category in the data category classification results is very low, while the posterior probability of another language in the language classification results corresponding to that data category is very high, resulting in misclassification. Therefore, this embodiment of the present invention multiplies the posterior probability of each data category in the data category classification results with the posterior probability of each language in the language classification results corresponding to that data category.

[0093] Specifically, the posterior probability of each data volume category in the data volume category classification result is multiplied by the posterior probability of each language in the language classification result corresponding to the data volume category to obtain the partial language recognition results of each language data volume category. For example, the data volume category is A, and category A contains three languages ​​A1, A2 and A3. The posterior probability of the language feature of the speech to be recognized in the data volume category classification result is P A , the posterior probability of A1 in the language classification result corresponding to the A data volume category of this language feature is P A1 , the posterior probability of A2 is P A2 , and the posterior probability of A3 is P A3 , then the language recognition results of the A data volume category are: A1 language recognition result is P A ×P A1 , the language recognition result of A2 is P A ×P A2 , A3 language recognition result is P A ×P A3 .

[0094] Then, based on the partial language recognition results for each data volume category, a second recognition result is obtained. It should be noted that the second recognition result can be obtained directly from the union of the partial language recognition results for each data volume category, or after obtaining the partial language recognition results for each data volume category, each language recognition result in the partial language recognition results for each data volume category can be normalized to obtain a score for each language in the partial language recognition results for each data volume category, and then obtained based on the score for each language in the partial language recognition results for each data volume category. This embodiment of the present invention is not limited to this.

[0095] Based on the above embodiments, Figure 4 This is a flow chart of the second language recognition model training method corresponding to the data volume category provided by the present invention. Figure 4 As shown, the second language recognition models corresponding to the multiple data volume categories in step 220 are trained based on the following steps:

[0096] Step 410, determining a second classification network, the second classification network including a backbone network and a second classification layer;

[0097] Specifically, a second classification network including a backbone network and a second classification layer is determined, wherein the backbone network is used to extract language features of the sample, and the second classification layer is an untrained classification network.

[0098] Step 420: Based on any data volume classification sample set, the second classification network is trained, and the second classification layer in the trained second classification network is used as the second language recognition model for the data volume category corresponding to the data volume classification sample set.

[0099] Specifically, the training of the second classification network is to fix the network parameters of the backbone network, train the second classification network through each sample in any data volume classification sample set, and use the second classification layer in the trained second classification network as the second language recognition model of the data volume category corresponding to the data volume classification sample set.

[0100] It should be noted that each data volume classification sample set has a corresponding second classification network. For example, the data volume classification sample sets corresponding to the three data volume categories are recorded as A, B and C respectively, then the data volume classification sample set A corresponds to a second classification network, the data volume classification sample set B corresponds to a second classification network, and the data volume classification sample set C corresponds to a second classification network.

[0101] Based on the above embodiments, Figure 5 This is a flow chart of the method for obtaining language recognition results provided by the present invention. Figure 5 As shown, step 130 includes:

[0102] Step 131: Determine a third recognition result of the speech to be recognized based on multiple feature classification sample sets; each feature classification sample set includes a third sample speech in a language corresponding to a feature category, and the multiple feature classification sample sets are obtained based on the language feature classification of the sample speech in all languages;

[0103] Step 132: Determine a language recognition result based on the first recognition result and / or the second recognition result, and the third recognition result.

[0104] Specifically, the third recognition result is obtained by performing language classification on the language features of the speech to be recognized. The language classification of the language features of the speech to be recognized here can be a pre-acquired feature category classification mapping relationship and a language mapping relationship corresponding to multiple feature categories, and based on the feature category classification mapping relationship and the language mapping relationship corresponding to multiple feature categories, the language features of the speech to be recognized are language classified and mapped. Among them, the feature classification mapping relationship classifies the language features of the speech to be recognized into feature categories to obtain the feature classification result of the language features of the speech to be recognized, which can be specifically reflected in the feature category classification model obtained through model training, and the language mapping relationships corresponding to multiple feature categories are respectively obtained by performing language classification on the language features of the speech to be recognized. The multiple mapping relationships here can specifically be reflected in the third language recognition model corresponding to the multiple feature categories obtained through model training.

[0105] Here, the third language recognition model corresponding to each feature category is trained based on the third sample set of the language corresponding to that feature category, where the feature category is a similarity cluster classification of sample features. In particular, given that the language features of speech to be recognized with similar speech features are easily confused during language recognition, resulting in a low recognition rate when simultaneously recognizing languages ​​with similar speech features, the third sample languages ​​of all languages ​​are classified into feature categories based on their sample features, and the third sample languages ​​of the languages ​​within each feature category are used to train the third language recognition model corresponding to that feature category. This helps improve the model's recognition rate for easily confused languages.

[0106] It should be noted that the third recognition result can be obtained by first performing feature category classification on the language features of the speech to be recognized through a feature category classification model to obtain a feature category classification result, and then performing language classification on the language features of the speech to be recognized based on the third language recognition model corresponding to the feature category classification result. It can also be obtained by performing feature category classification on the language features of the speech to be recognized through a feature category classification model to obtain a feature category classification result, performing language recognition on the language features of the speech to be recognized using third language recognition models corresponding to multiple feature categories to obtain multiple partial language recognition results, and then obtaining it based on the feature category classification result and the multiple partial language recognition results. The embodiment of the present invention does not impose any restrictions on this.

[0107] In addition, for the same language, the "third" here is used to distinguish it from the "first" and "second" mentioned above. It is used to indicate that the sample speech belongs to a feature classification sample set. For the same language, the third sample speech of the language can be the same or different in content from the first sample speech or the second sample speech mentioned above, the third sample speech can be the same or different in feature category from the first sample speech or the second sample speech mentioned above, and the third sample speech can be all the sample speech of the language or a portion of the sample speech selected from the language, which is not limited in this embodiment of the present invention.

[0108] In step 132, the first recognition result or the second recognition result can be combined with the third recognition result as the final language recognition result, or the first recognition result, the second recognition result and the third recognition result can be combined to obtain the final language recognition result. For example, the first recognition result, the second recognition result and the third recognition result can be weighted or the average value of the corresponding language posterior probabilities in the three results can be calculated. This embodiment of the present invention is not limited to this.

[0109] The language recognition method provided by the embodiment of the present invention improves the language recognition rate for easily confused language features by determining the third recognition result through the mapping relationship between multiple feature classification sample sets and language recognition results obtained through training, and further improves the language recognition result by jointly determining the language recognition result through the first recognition result, the second recognition result and the third recognition result, thereby realizing mutual verification of the three results and further improving the accuracy of language recognition.

[0110] Based on the above embodiment, determining the third recognition result of the language feature based on the multiple feature classification sample sets in step 131 includes:

[0111] Step 610, classifying the language features into feature categories to obtain a feature category classification result of the language features;

[0112] Step 620: Classify the language features based on the third language recognition models corresponding to the multiple feature categories to obtain language classification results for the language features under the multiple feature categories. The third language recognition models corresponding to the multiple feature categories are trained based on the multiple feature classification sample sets.

[0113] Considering that if the language features of the speech to be recognized are first classified by feature category to obtain a feature category classification result, and then the language features are subjected to language recognition based on the third language recognition model corresponding to the feature category determined in the feature category classification result to obtain a language recognition result, in this case, the language recognition of the language features of the speech to be recognized will be completely dependent on the feature category classification result. Once the feature category classification result is incorrect, it will directly lead to the failure of the subsequent language recognition task. Therefore, the embodiment of the present invention simultaneously considers the classification result output by the feature category classification model and the partial language recognition results output by the third language recognition model corresponding to multiple feature categories.

[0114] Specifically, step 610 inputs the language features of the speech to be recognized into the feature category classification model to obtain the feature category classification result of the language features of the speech to be recognized output by the feature category classification model.

[0115] In step 620, the language features of the speech to be recognized are input into the third language recognition model corresponding to each feature category. The third language recognition model corresponding to each feature category outputs a language classification result for that feature category. The third language recognition model corresponding to each feature category is trained based on the feature classification samples corresponding to that feature category.

[0116] It should be noted that the execution order of step 610 and step 620 is not particular, and step 610 and step 620 can be serial or parallel, and the embodiment of the present invention does not limit this.

[0117] Step 630: Determine a third recognition result based on the feature category classification result and the language classification results under the multiple feature categories.

[0118] It should be noted that the third recognition result can sort the posterior probabilities of each feature category in the feature classification result from high to low, select the highest preset number of feature categories, and then use the language with the highest posterior probability of language classification in the language classification results under these feature categories as the third recognition result. The posterior probability of each feature category in the feature classification result can also be weighted or multiplied with the posterior probabilities of each language classification in the language classification results under the feature category to obtain the partial language recognition results corresponding to each feature category, and the partial language recognition results corresponding to each feature category are combined to obtain the third recognition result. The embodiment of the present invention does not limit this.

[0119] Based on the above embodiment, step 630 includes:

[0120] Step 631, based on the posterior probability of any feature category in the feature category classification results and the language classification result corresponding to the feature category, determine a partial language recognition result for the feature category;

[0121] Step 632: Obtain a third recognition result based on the partial language recognition results of each feature category.

[0122] Considering that when performing language recognition based on the language features of the speech to be recognized, it is possible that the posterior probability of a certain feature category in the feature category classification result is very low, while the posterior probability of a certain language in the language classification result corresponding to that feature category is very high, resulting in misclassification. Therefore, the embodiment of the present invention multiplies the posterior probability of each feature category in the feature category classification result with the posterior probability of each language in the language classification result corresponding to that feature category.

[0123] Specifically, the posterior probability of each feature category in the feature category classification result is multiplied by the posterior probability of each language in the language classification result corresponding to the feature category to obtain the partial language recognition results of each language feature category. For example, the feature category is A, and category A contains three languages ​​A1, A2 and A3. The posterior probability of the language feature of the speech to be recognized in the feature category classification result is P A , the posterior probability of A1 in the language classification result corresponding to the language feature category A is P A1 , the posterior probability of A2 is P A2 , and the posterior probability of A3 is P A3 , then the language recognition results of some of the A feature categories are: A1 language recognition result is P A ×P A1 , the language recognition result of A2 is P A ×PA2 , the A3 language recognition result is P A ×P A3 .

[0124] Then, based on the partial language recognition results of each feature category, a second recognition result is obtained.

[0125] It should be noted that the second recognition result can be obtained directly based on the union of the partial language recognition results of each feature category, or after obtaining the partial language recognition results of each feature category, the language recognition results in the partial language recognition results of each feature category can be normalized to obtain the scores of each language in the partial language recognition results of each feature category, and the scores of each language in the partial language recognition results of each feature category can be obtained based on the scores of each language in the partial language recognition results of the feature category. The embodiments of the present invention do not limit this.

[0126] Based on the above embodiment, the third language recognition models corresponding to the multiple feature categories in step 620 are trained based on the following steps:

[0127] Step 810, determining a third classification network, where the third classification network includes a backbone network and a third classification layer;

[0128] Specifically, a third classification network including a backbone network and a third classification layer is determined, wherein the backbone network is used to extract language features of the sample, and the third classification layer is an untrained classification network.

[0129] Step 820 : Based on any feature classification sample set, the third classification network is trained, and the third classification layer in the trained third classification network is used as the third language recognition model for the feature category corresponding to the feature classification sample set.

[0130] Specifically, the training of the third classification network is to fix the network parameters of the backbone network, train the third classification network through each sample in any feature classification sample set, and use the third classification layer in the trained third classification network as the third language recognition model of the feature category corresponding to the feature classification sample set.

[0131] It should be noted that each feature classification sample set has a corresponding third classification network. For example, the feature classification sample sets corresponding to the three feature categories are recorded as A, B and C respectively, then the feature classification sample set A corresponds to a third classification network, the feature classification sample set B corresponds to a third classification network, and the feature classification sample set C corresponds to a third classification network.

[0132] Based on the above embodiments, embodiments of the present invention also provide a language recognition method capable of accurately identifying the target language in scenarios with an imbalanced distribution of language data and a large number of language categories. Embodiments of the present invention perform speech language recognition using a language recognition system. The system includes a backbone network and a first language recognition model, a second language recognition model, and a third language recognition model. The first language recognition model corresponds to the language recognition model obtained through Strategy 1 described below, the second language recognition model corresponds to the language recognition model obtained through Strategy 2 described below, and the third language recognition model corresponds to the language recognition model obtained through Strategy 3 described below. The system employs a ResNet (Residual Network) language recognition model and uses a full set of annotated language samples to train three language recognition models using three strategies. Finally, the scores of the three language recognition models are fused to improve language recognition performance. Here, the three language recognition models are the first, second, and third language recognition models, and the scores of the three models are the first, second, and third language recognition results.

[0133] Before training the three language recognition models, we first used a full set of annotated language samples and trained the backbone network using the CE loss (cross entropy loss function) to extract language features. After that, we fixed the backbone network parameters and used different strategies to train the first, second, and third language recognition models.

[0134] In strategy 1, a first classification network is constructed based on the backbone network and the first classification layer. A balanced sample set of all languages, i.e., a full sample set, is obtained by class-balanced sampling. A decoupled training method is adopted to fix the backbone network in the first classification network, and the first classification layer is trained based on the full sample set to obtain a first language recognition model, thereby reducing the situation where the speech of the language with a large sample data volume dominates the network parameter training during random sampling, resulting in poor recognition of the language with a small sample data volume; in strategy 2, the sample set of all languages ​​is divided into M data volume categories according to the sample data volume of each language, and the sample data volume distribution of each language in each data volume category is made as close as possible. A second classification network corresponding to each data volume category is constructed based on the backbone network and the second classification layer corresponding to each data volume category, and the method of strategy 1 is adopted to train the second classification layer corresponding to each data volume category based on each data volume classification sample set to obtain a second language recognition model corresponding to each data volume category, so as to reduce the impact of the imbalance of sample data volume on the model training effect. At the same time, an initial data volume category classification model is constructed based on the backbone network and the data volume category classification layer, and a training set is trained based on each data volume category sample. The initial data volume category classification model is trained to obtain the data volume category classification model, and the data volume category classification model is used to obtain the data volume category classification results of speech during recognition; in strategy 3, the traditional TV technology is used to extract the i-vector of the samples in the sample set of all languages, and then the threshold is set. The samples of all languages ​​are clustered into several feature categories by clustering. The third classification network corresponding to each feature category is constructed based on the backbone network and the third classification layer corresponding to each feature category, and the strategy 1 is adopted to train the third classification layer corresponding to each feature category based on each feature classification sample set to obtain the third language recognition model corresponding to each feature category, so as to improve the recognition effect of easily confused languages ​​in category C languages. At the same time, the initial feature category classification model is constructed based on the backbone network and the feature category classification layer, and the initial feature category classification model is trained based on each feature classification sample set to obtain the feature category classification model. The feature category classification model is used to obtain the feature category classification results of speech during recognition. During recognition, Strategy 2 and Strategy 3 models combine the scores of the macro-classification model (the data volume classification model in Strategy 2; the feature classification model in Strategy 3) and the sub-classification model (the second language recognition model in Strategy 2; the third language recognition model in Strategy 3). Specifically, Strategy 2 combines the scores of the data volume classification results with the recognition results of the second language recognition model corresponding to each data volume category, while Strategy 3 combines the scores of the feature classification results with the recognition results of the third language recognition model corresponding to each feature category. This reduces the possibility of language misclassification. The last three strategies each provide a speech score, which is then averaged. The language type corresponding to the highest score is the language recognition result.

[0135] The above language recognition system is divided into two parts: training and recognition, as follows:

[0136] Training phase:

[0137] The training is divided into two stages. The first stage mainly trains the backbone network (Resnet BackBone) parameters for extracting language features; the second stage fixes the Resnet BackBone parameters and trains the language recognition models of the three strategies respectively.

[0138] Step 1: Prepare a sample set of all languages ​​for training. Collect samples for each language, requiring at least 1 hour of samples for each language.

[0139] Step 2: For the samples in the sample set of all languages ​​in step 1, filter out invalid samples such as silence and noise, and extract SDC features.

[0140] Step 3: Figure 6 This is a schematic diagram of the backbone network training structure provided by the present invention. Figure 6 As shown in the figure, the backbone network is concatenated with the linear fully connected layer 1 (FC layer 1) and trained using the CE loss. In this training step, data is randomly sampled. The network's initial learning rate is set to 0.1, and the loss function is used to iterate step 3 until the concatenated backbone network and the linear fully connected layer 1 converge.

[0141] Step 4: Figure 7 This is a schematic diagram of the structure of the first classification network training provided by the present invention. Figure 7 As shown in the figure, based on step 3, the first classification network is trained based on the full sample set. Specifically, the network parameters of the trained backbone network are fixed, and the language features output by the backbone network are embedded into the first classification layer, the linear fully connected layer 2 (FC2 layer), for classification. The network is trained using CE loss until convergence at the first classification layer.

[0142] During this training step, it is particularly important to note that the full sample set is selected from the sample set of all languages ​​using a class-balanced sampler to ensure that the proportion of sample data for each language is basically consistent. This reduces the dominance of languages ​​with large sample data volumes on the first classification layer, allowing languages ​​with small sample data volumes to be fully trained.

[0143] Step 5: Repeat step 4 until the CE Loss stabilizes or the maximum number of iterations is reached. At this point, the training of the first classification layer of strategy 1 is completed, and FC layer 2 is used as the first language recognition model.

[0144] Step 6: Figure 8 This is a schematic diagram of the structure of the second classification network training provided by the present invention. Figure 8As shown, according to the sample data volume of each language category, M data volume categories are divided, and based on the M data volume classification sample sets, the data volume category classification model and the second language recognition model corresponding to the M data volume categories are trained respectively, wherein the data volume category classification model is used to classify speech into a category among the M data volume categories, and the second language recognition model corresponding to the M data volume categories is used to identify the language category of the speech. The data volume category classification model is trained based on the M data volume classification sample sets. Specifically, the backbone network parameters are fixed, and the language features output by the backbone network are embedded into the data volume category classification layer (linear fully connected layer) of size M×1. The second classification network corresponding to the i-th data volume category is trained based on the i-th data volume classification sample set. Specifically, assuming that the i-th data volume category contains Ti types of languages, the backbone network parameters are fixed, and the language features output by the backbone network are embedded into the second classification layer (linear fully connected layer) of the second classification network corresponding to the i-th data volume category of size Ti×1.

[0145] The training method for this step is consistent with the model training method for Strategy 1, using class-balanced sampling and the CE loss criterion. It is important to note that when dividing the data into categories based on sample data volume, the sample data volume of each language within each category should be as balanced as possible. Repeated iterations are performed until the second classification layer of the second classification network for each data category converges, resulting in the second language recognition model for each data category in Strategy 2.

[0146] Step 7: Using the features in step 2, use the EM algorithm (maximum expectation algorithm) to train the UBM and T (factor orthogonal space) matrix, and extract the i-vector of the sample speech.

[0147] Step 8: Use the AP (Affinity Propagation) clustering algorithm to cluster all i-vectors from Step 7. Set a reasonable clustering threshold. After clustering is complete, count the language categories within each cluster's feature category. If i-vectors for a given language's speech sample appear in multiple feature categories, only the feature category with the largest number of i-vectors for that language's speech sample is selected as the feature category for that language. Assume that after clustering, a total of N feature categories are clustered.

[0148] Step 9: Figure 9 This is a schematic diagram of the structure of the third classification network training provided by the present invention. Figure 9As shown, for the N feature categories divided by clustering in step 8, a feature category classification model and a third language recognition model corresponding to each of the N feature categories are trained based on the N feature classification sample sets. The feature category classification model is used to classify speech into one of the N feature categories, while the third language recognition model corresponding to each feature category is used to identify the language category of the speech. Similarly, training is performed according to the model training method of Strategy 1, and iteratively iterates until the third classification layer in the third classification network corresponding to each feature category converges, resulting in the third language recognition model corresponding to each feature category of Strategy 3.

[0149] Identification phase:

[0150] Step 1: Extract sdc features, i.e. language features, from the speech to be recognized;

[0151] Step 2: Input the language features extracted in step 1 into the first language recognition model in strategy 1 to obtain the posterior probability of each language category output by the first language recognition model in strategy 1. Normalize the posterior probability using the softmax function to obtain the score of each language classification, which is recorded as the first recognition result Pstrategy_1 = {Pstrategy_1_1, ..., Pstrategy_1_j, ..., Pstrategy_1_C}, where C represents the number of training language categories, and 1 <j<C。

[0152] Step 3: Load the data volume category classification model and the second language recognition model corresponding to the M data volume categories in strategy 2. First, use the data volume category classification model to calculate the posterior probability of the language features extracted in step 1 on the M data volume categories, and normalize the posterior probability using the softmax function, denoted as fNumSplit = {fNumSplit_1, ..., fNumSplit_i, ..., fNumSplit_M}. Then, based on the language features extracted in step 1, calculate the posterior probability of each language on the second language recognition model corresponding to the M data volume categories. Specifically, assuming that there are Ti language categories on the second language recognition model corresponding to the i-th data volume category, the posterior probability of the language features extracted in step 1 on the second language recognition model corresponding to the i-th data volume category is gNumSplit_i = {gNumSplit_i_1, ..., gNumSplit_i_k, ..., gNumSplit_i _Ti}, then the posterior probability output by strategy 2 is uNumSplit_i=fNumSplit_i*gNumSplit_i={fNumSplit_i*gNumSplit_i_1,…,fNumSplit_i*gNumSplit_i_k,…,fNumSplit_i*gNumSplit_i_Ti}, then uNumSplit_i is normalized using the softmax function to obtain the score of the language features extracted in step 1 after passing through the second language recognition model corresponding to the i-th data volume category, pNumSplit_i={pNumSplit_i_1,…,pNumSplit_i_k,…,pNumSplit_i_Ti}, where pNumSplit_i_k (if the k-th language of the i-th feature category has the language index j in strategy 2, it is recorded as Pstrategy_2_j in strategy 2) is calculated as follows:

[0153]

[0154] According to this method, the scores of C language categories are solved in sequence and recorded as the second recognition result Pstrategy_2 = {Pstrategy_2_1, ..., Pstrategy_2_j, ..., Pstrategy_2_C}.

[0155] The reason why the classification probability of the data volume category classification model is used in this step is mainly to prevent the situation where the language feature extracted in step 1 has a very low score in the i-th data volume category, but the score after softmax of a certain language category in the second language recognition model corresponding to the i-th data volume category is very high, resulting in the misclassification of the speech to be recognized by only looking at the scores on the second language recognition model corresponding to M data volume categories.

[0156] Step 4: Load the feature category classification model in strategy 3 and the third language recognition model corresponding to N feature categories. First, use the feature category classification model to calculate the posterior probability of the language features extracted in step 1 on the N feature categories, and normalize the posterior probability using the softmax function, denoted as fClusterSplit = {fClusterSplit_1, ..., fClusterSplit_i, ..., fClusterSplit_N}. Then, based on the language features extracted in step 1, calculate the posterior probability of each language on the third language recognition model corresponding to the N feature categories. Specifically, assuming that there are Li language categories on the third language recognition model corresponding to the i-th feature category, the posterior probability of the language features extracted in step 1 on the third language recognition model corresponding to the i-th feature category is gClusterSplit_i = {gClusterSplit_i_1, ..., gClusterSplit_i_k, ..., gClusterSplit_i_Li}, then the posterior probability output by strategy 3 is u ClusterSplit_i = fClusterSplit_i*gClusterSplit_i = {fClusterSplit_i*gClusterSplit_i_1, …, fClusterSplit_i*gClusterSplit_i_k, …, fClusterSplit_i*gClusterSplit_i_Li}, then normalize uClusterSplit_i using the softmax function to obtain the score of the language features extracted in step 1 after passing through the third language recognition model corresponding to the i-th feature category, pClusterSplit_i = {pClusterSplit_i_1, …, pClusterSplit_i_k, …, pClusterSplit_i_Li}, where pClusterSplit_i_k (if the k-th language of the i-th feature category has the language index j in strategy 3, it is denoted as Pstrategy_3_j in strategy 3) is calculated as follows:

[0157]

[0158] According to this method, the scores of C language categories are solved in sequence and recorded as the third recognition result Pstrategy_3 = {Pstrategy_3_1, ..., Pstrategy_3_j, ..., Pstrategy_3_C}.

[0159] The reason why the classification probability of the feature category classification model is used in this step is mainly to prevent the situation where the language feature extracted in step 1 has a very low score in the i-th feature category, but the score after softmax of a certain language category in the third language recognition model corresponding to the i-th feature category is very high, resulting in the misclassification of the speech to be recognized by only looking at the scores on the third language recognition model corresponding to N feature categories.

[0160] Step 5: Calculate the average of the scores for each language category obtained in steps 2 to 4, that is, the scores for the first, second, and third recognition results, and record it as Paverage = {Paverage_1, ..., Paverage_i, ..., Paverage_C}, where Paverage_i is calculated as follows:

[0161]

[0162] The language category corresponding to the largest score in Paverage is taken as the final language recognition result of the speech to be recognized.

[0163] The language recognition device provided by the present invention is described below. The language recognition device described below and the language recognition method described above can be referenced to each other.

[0164] Figure 10 This is a schematic diagram of the structure of the language recognition device provided by the present invention. Figure 10 As shown, the device includes: a feature determination module 1010, a language recognition module 1020 and a result determination module 1030.

[0165] in,

[0166] A feature determination module 1010 is used to extract language features of the speech to be recognized based on the backbone network;

[0167] The language identification module 1020 is configured to determine a first recognition result of the language feature based on the full sample set, wherein the full sample set includes first sample speech of all languages, and the first sample speech of all languages ​​is evenly distributed;

[0168] and / or, determining a second recognition result of the language feature based on a plurality of data volume classification sample sets, wherein each data volume classification sample set includes a second sample speech of a language corresponding to a data volume category, and the plurality of data volume classification sample sets are obtained based on the data volume division of the second sample speech of all languages;

[0169] The result determination module 1030 is configured to determine a language recognition result based on the first recognition result and / or the second recognition result.

[0170] In the embodiment of the present invention, the feature determination module 1010 is used to determine the language feature of the language feature;

[0171] The language recognition module 1020 is used to determine a first recognition result of the language feature based on the full sample set; the full sample set includes the first sample speech of the full language, and the first sample speech of the full language is evenly distributed; and / or, based on multiple data volume classification sample sets, determine a second recognition result of the language feature; each data volume classification sample set includes the second sample speech of the language of the corresponding data volume category, and the multiple data volume classification sample sets are obtained based on the data volume division of the second sample speech of the full language; the result determination module 1030 is used to determine the language recognition result based on the first recognition result and / or the second recognition result, reduce the situation where the language with a large distribution proportion dominates the network parameter training during random sampling of samples, resulting in poor recognition of the language with a small distribution proportion, improve the recognition rate of the language with a small distribution proportion, and reduce the situation where the imbalance of the training sample speech data volume leads to poor recognition of the language with a small sample speech data volume, improve the recognition rate of the language with a small training sample speech data volume, and can determine the language recognition result by combining the first recognition result with the second recognition result, so as to realize mutual verification of the two results and further improve the accuracy of language recognition.

[0172] Based on any of the above embodiments, the language identification module 1020 includes:

[0173] Determine a first recognition result submodule, configured to classify the language features based on the first language recognition model to obtain a first recognition result;

[0174] The first language recognition model training module is used to train the first classification layer in the first classification network obtained based on the full sample set as the first language recognition model; the first classification network includes a backbone network and a first classification layer.

[0175] Based on any of the above embodiments, the language identification module 1020 includes:

[0176] The data volume category classification submodule is used to classify the language features into data volume categories and obtain the data volume category classification results of the language features;

[0177] A data volume classification language identification submodule is used to classify language features based on second language identification models corresponding to multiple data volume categories, thereby obtaining language classification results for the language features under multiple data volume categories. The second language identification models corresponding to the multiple data volume categories are trained based on multiple data volume classification sample sets.

[0178] The second recognition result determination submodule is used to determine the second recognition result based on the data volume category classification result and the language classification results under multiple data volume categories.

[0179] Based on any of the above embodiments, the second recognition result determination submodule includes:

[0180] A submodule for determining partial language recognition results for a data volume category, configured to determine partial language recognition results for the data volume category based on the posterior probability of any data volume category in the data volume category classification results and the language classification results corresponding to the data volume category;

[0181] The second recognition result calculation submodule is used to obtain a second recognition result based on the partial language recognition results of each data volume category.

[0182] Based on any of the above embodiments, the data volume classification and language identification submodule includes:

[0183] A second classification network determination submodule, used to determine a second classification network, where the second classification network includes a backbone network and a second classification layer;

[0184] The language recognition model training submodule for data volume category is used to train the second classification network based on any data volume classification sample set, and use the second classification layer in the trained second classification network as the second language recognition model for the data volume category corresponding to the data volume classification sample set.

[0185] Based on any of the above embodiments, the result determination module 1030 includes:

[0186] A third recognition result determination submodule is configured to determine a third recognition result of language features based on a plurality of feature classification sample sets, wherein each feature classification sample set includes a third sample speech of a language corresponding to a feature category, and the plurality of feature classification sample sets are obtained based on the language features of the sample speech of all languages;

[0187] The language recognition result determination submodule is configured to determine a language recognition result based on the first recognition result and / or the second recognition result, and the third recognition result.

[0188] Based on any of the above embodiments, the third recognition result determination submodule includes:

[0189] The feature classification category classification submodule is used to classify the language features into feature categories and obtain the feature category classification results of the language features;

[0190] A feature classification language identification submodule is used to classify language features based on third language identification models corresponding to multiple feature categories, thereby obtaining language classification results for the language features under the multiple feature categories. The third language identification models corresponding to the multiple feature categories are trained based on multiple feature classification sample sets.

[0191] The third recognition result determination submodule is configured to determine a third recognition result based on the feature category classification result and the language classification results under the multiple feature categories.

[0192] Based on any of the above embodiments, determining the third recognition result submodule includes:

[0193] A submodule for determining partial language recognition results of a feature category determines partial language recognition results of the feature category based on the posterior probability of any feature category in the feature category classification results and the language classification results corresponding to the feature category;

[0194] The third recognition result calculation submodule is used to obtain a third recognition result based on the partial language recognition results of each feature category.

[0195] Based on any of the above embodiments, the feature classification language identification submodule includes:

[0196] A third classification network determination submodule, used for a third classification network, where the third classification network includes a backbone network and a third classification layer;

[0197] The language recognition model training submodule for feature categories is used to train the third classification network based on any feature classification sample set, and use the third classification layer in the trained third classification network as the third language recognition model for the feature category corresponding to the feature classification sample set.

[0198] Figure 11 An example of a physical structure diagram of an electronic device is shown below. Figure 11 As shown, the electronic device may include: a processor 1110, a communication interface 1120, a memory 1130, and a communication bus 1140, wherein the processor 1110, the communication interface 1120, and the memory 1130 communicate with each other via the communication bus 1140. The processor 1110 may call logic instructions in the memory 1130 to execute a language recognition method, which includes: extracting language features of a speech to be recognized based on a backbone network; determining a first recognition result of the language features based on a full sample set; the full sample set includes first sample speech of all languages, and the first sample speech of all languages ​​is evenly distributed; and / or determining a second recognition result of the language features based on multiple data volume classification sample sets; each data volume classification sample set includes second sample speech of a language corresponding to a data volume category, and the multiple data volume classification sample sets are obtained based on the data volume division of the second sample speech of the full language; and determining a language recognition result based on the first recognition result and / or the second recognition result.

[0199] In addition, the logic instructions in the above-mentioned memory 1130 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0200] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the language recognition method provided by the above-mentioned methods, which includes: extracting the language features of the speech to be recognized based on the backbone network; determining a first recognition result of the language features based on the full sample set; the full sample set includes the first sample speech of the full language, and the first sample speech of the full language is evenly distributed; and / or determining a second recognition result of the language features based on multiple data volume classification sample sets; each data volume classification sample set includes the second sample speech of the language of the corresponding data volume category, and the multiple data volume classification sample sets are obtained based on the data volume division of the second sample speech of the full language; determining the language recognition result based on the first recognition result and / or the second recognition result.

[0201] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the language recognition method provided by the above-mentioned methods, the method comprising: extracting the language features of the speech to be recognized based on a backbone network; determining a first recognition result of the language features based on a full sample set; the full sample set includes the first sample speech of the full language, and the first sample speech of the full language is evenly distributed; and / or determining a second recognition result of the language features based on multiple data volume classification sample sets; each data volume classification sample set includes the second sample speech of the language of the corresponding data volume category, and the multiple data volume classification sample sets are obtained based on the data volume division of the second sample speech of the full language; determining the language recognition result based on the first recognition result and / or the second recognition result.

[0202] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0203] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0204] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A language identification method, characterized in that: include: Based on the backbone network, the language features of the speech to be recognized are extracted; Determining a first recognition result of the language feature based on a full sample set; the full sample set includes first sample speech of all languages, and the first sample speech of all languages ​​is evenly distributed; the full sample set is obtained by selecting an equal number of sample speech of each language in the full language from a large number of samples; Classifying the language features by data volume category to obtain a data volume category classification result of the language features; performing language classification on the language feature based on second language recognition models corresponding to a plurality of data volume categories, thereby obtaining language classification results for the language feature under the plurality of data volume categories, wherein the second language recognition models corresponding to the plurality of data volume categories are trained based on a plurality of data volume classification sample sets; Determining a partial language recognition result for any data volume category based on the posterior probability of any data volume category in the data volume category classification results and the language classification result corresponding to the any data volume category; Based on the partial language recognition results of each data volume category, a second recognition result is obtained; each data volume classification sample set includes a second sample speech of the language of the corresponding data volume category, and the multiple data volume classification sample sets are obtained based on the data volume of the second sample speech of the full language; Determining a third recognition result of the language feature based on a plurality of feature classification sample sets, wherein each feature classification sample set includes a third sample speech of a language corresponding to a feature category, and the plurality of feature classification sample sets are obtained based on the language feature classification of the sample speech of the full set of languages; The language recognition result is determined based on the first recognition result, the second recognition result, and the third recognition result.

2. The language identification method according to claim 1, wherein: Determining the first recognition result of the language feature based on the full sample set includes: Based on the first language recognition model, classify the language features into different languages ​​to obtain the first recognition result; The first language recognition model is a first classification layer in a first classification network trained based on the full sample set, and the first classification network includes the backbone network and the first classification layer.

3. The language identification method according to claim 1, wherein: The second language recognition models corresponding to the multiple data volume categories are trained based on the following steps: determining a second classification network, wherein the second classification network includes the backbone network and a second classification layer; Based on any data volume classification sample set, the second classification network is trained, and the second classification layer in the trained second classification network is used as the second language recognition model for the data volume category corresponding to the any data volume classification sample set.

4. The language identification method according to claim 1, wherein: The determining of the third recognition result of the language feature based on the plurality of feature classification sample sets includes: Performing feature category classification on the language features to obtain feature category classification results of the language features; performing language classification on the language features based on third language recognition models corresponding to a plurality of feature categories to obtain language classification results for the language features under the plurality of feature categories, wherein the third language recognition models corresponding to the plurality of feature categories are trained based on the plurality of feature classification sample sets; The third recognition result is determined based on the feature category classification result and the language classification results under the multiple feature categories.

5. The language identification method according to claim 4, characterized in that: Determining the third recognition result based on the feature category classification result and the language classification results under the multiple feature categories includes: determining a partial language recognition result for any feature category based on the posterior probability of any feature category in the feature category classification results and the language classification result corresponding to the any feature category; The third recognition result is obtained based on the partial language recognition results of each feature category.

6. The language identification method according to claim 4 or 5, characterized in that: The third language recognition models corresponding to the multiple feature categories are trained based on the following steps: Determining a third classification network, wherein the third classification network includes the backbone network and a third classification layer; Based on any feature classification sample set, the third classification network is trained, and the third classification layer in the trained third classification network is used as the third language recognition model for the feature category corresponding to the any feature classification sample set.

7. A language recognition device, characterized in that: include: The feature determination module is used to extract the language features of the speech to be recognized based on the backbone network; a language recognition module configured to determine a first recognition result of the language feature based on a full sample set, wherein the full sample set includes first sample speech of all languages, and the first sample speech of all languages ​​is evenly distributed; the full sample set is obtained by selecting an equal number of sample speech of each language in the full sample set from a large number of samples; A category classification module, configured to classify the language features into data volume categories to obtain a data volume category classification result of the language features; a language classification module configured to classify the language features based on second language recognition models corresponding to a plurality of data volume categories, thereby obtaining language classification results for the language features under the plurality of data volume categories, wherein the second language recognition models corresponding to the plurality of data volume categories are trained based on a plurality of data volume classification sample sets; a partial language recognition module, configured to determine a partial language recognition result for any data volume category based on the posterior probability of any data volume category in the data volume category classification results and the language classification result corresponding to the any data volume category; A second result determination module is used to obtain a second recognition result based on the partial language recognition results of each data volume category; a third result determination module, configured to determine a third recognition result of the language feature based on a plurality of feature classification sample sets, each feature classification sample set comprising a third sample speech of a language corresponding to a feature category, the plurality of feature classification sample sets being obtained based on the language feature classification of the sample speech of the full set of languages; The language recognition result determination module is configured to determine the language recognition result based on the first recognition result, the second recognition result, and the third recognition result.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the language recognition method according to any one of claims 1 to 6 are implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the language recognition method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Intelligent cross-language speech recognition and conversion method

    CN107945805A

  • Text classification method and device, storage medium and computer equipment

    CN110287311A

  • Language recognition method and device and language recognition model training method and device

    CN113724700A