Language recognition method, training method, equipment and storage medium

By introducing constraint loss adjustment parameters into the language recognition model, increasing the phonetic feature distance between different languages ​​and reducing the distance within the same language, the problem of low accuracy in the recognition of target languages ​​in the prior art is solved, and higher recognition accuracy is achieved.

CN119993121AActive Publication Date: 2025-05-13HEFEI IFLY DIGITAL TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510017109.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-13
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The existing language recognition methods have limitations in dealing with pronunciation characteristics between different languages ​​and within languages, resulting in low accuracy in recognition of target languages.

Method used

By introducing constraint loss, the language recognition model is adjusted, the pronunciation feature distance between different second languages ​​is increased, and the pronunciation feature distance within the same second language is reduced, thereby improving the recognition ability and recognition sensitivity of the target language.

Benefits of technology

Even when training the language recognition model, the model can accurately recognize the pronunciation of the target language, improving the accuracy of the recognition of the target language.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993121A_ABST
    Figure CN119993121A_ABST
Patent Text Reader

Abstract

The invention discloses a language recognition method, a training method, equipment and a storage medium. The method comprises the following steps: acquiring to-be-recognized voice data; performing feature extraction on the to-be-recognized voice data by using the language recognition model to obtain to-be-recognized voice features; a language recognition model is used for recognition based on the to-be-recognized voice features, a first language recognition result is obtained, and the first language recognition result is used for representing the language to which the to-be-recognized voice data belong; wherein the language recognition model is obtained by at least utilizing constraint loss parameter adjustment, the constraint loss is determined based on a first sample feature extracted from the first training set by the language recognition model, and the target of utilizing the constraint loss parameter adjustment is to increase the voice feature distance between different second languages and reduce the voice feature distance in the same second language. Through the method, the recognition accuracy of the target language can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech processing, and in particular to a language recognition method, training method, device and storage medium. Background Art

[0002] Language recognition is the process of determining the language to which a speech segment belongs. Specifically, the speech data to be recognized can be subjected to language recognition to obtain the language category to which the speech data to be recognized belongs. However, general language recognition methods have limitations in processing speech features between different languages ​​and within a language, resulting in low recognition accuracy of the target language. Summary of the invention

[0003] The main technical problem solved by the present application is to provide a language recognition method, training method, device and storage medium, which can improve the recognition accuracy of the target language.

[0004] In order to solve the above technical problems, a technical solution adopted in the present application is: to provide a language recognition method, the method comprising: obtaining speech data to be recognized; using a language recognition model to perform feature extraction on the speech data to be recognized to obtain speech features to be recognized; using the language recognition model to perform recognition based on the speech features to be recognized to obtain a first language recognition result, the first language recognition result is used to characterize the language to which the speech data to be recognized belongs; wherein the language recognition model can recognize a number of first languages, the number of first languages ​​include at least one target language, the language recognition model is obtained by adjusting parameters using at least a constraint loss, the constraint loss is determined based on a first sample feature extracted from a first training set by the language recognition model, the first training set includes sample speech data of at least two second languages, and at least two second languages ​​include the target language, and the goal of adjusting parameters using the constraint loss is to increase the speech feature distance between different second languages ​​and reduce the speech feature distance within the same second language.

[0005] In order to solve the above technical problems, another technical solution adopted by the present application is: to provide a training method for a language recognition model, the method comprising: obtaining a first training set and a second training set, wherein the first training set includes sample speech data of at least two second languages, and the at least two second languages ​​include a target language, and the second training set includes sample speech data corresponding to at least one of the first languages; using the language recognition model to extract features from each sample speech data in the first training set to obtain first sample features of each sample speech data in the first training set, and using each first sample feature to determine the constraint loss of the language recognition model; and using the language recognition model to extract features from each sample speech data in the first training set to obtain first sample features of each sample speech data in the first training set, and using each first sample feature to determine the constraint loss of the language recognition model. The model performs feature extraction on each sample speech data in the second training set to obtain the second sample features of each sample speech data in the second training set, and performs recognition based on each second sample feature to obtain the predicted language corresponding to each sample speech data in the second training set; based on the difference between the predicted language and the labeled language corresponding to each sample speech data in the second training set, the prediction loss of the language recognition model is determined; the constraint loss and prediction loss of the language recognition model are used to adjust the network parameters of the language recognition model, wherein the goal of adjusting the parameters using the constraint loss is to increase the speech feature distance between different second languages ​​and reduce the speech feature distance within the same second language.

[0006] To solve the above technical problems, another technical solution adopted by the present application is: to provide a language recognition device, which includes: a speech data acquisition module for acquiring speech data to be recognized; a speech feature acquisition module for extracting features of the speech data to be recognized using a language recognition model to obtain speech features to be recognized; a language recognition module for performing recognition based on the speech features to be recognized using the language recognition model to obtain a first language recognition result, and the first language recognition result is used to characterize the language to which the speech data to be recognized belongs, wherein the language recognition model can recognize several first languages, and the several first languages ​​include at least one target language, and the language recognition model is obtained by adjusting parameters using at least a constraint loss, and the constraint loss is determined based on the first sample features extracted from the first training set by the language recognition model, and the first training set includes sample speech data of at least two second languages, and at least the two second languages ​​include the target language, and the goal of adjusting parameters using the constraint loss is to increase the speech feature distance between different second languages ​​and reduce the speech feature distance within the same second language.

[0007] In order to solve the above technical problems, another technical solution adopted in the present application is: to provide a language recognition device, including a memory and a processor, the memory stores program instructions, and the processor is used to execute the program instructions to implement the above language recognition method.

[0008] In order to solve the above technical problems, another technical solution adopted in the present application is: providing a computer-readable storage medium, which is used to store program instructions, and the program instructions can be executed to implement the above language recognition method.

[0009] The above scheme uses a language recognition model to extract features of speech data to be recognized, obtains speech features to be recognized, and recognizes based on the speech features to be recognized to obtain a first language recognition result, which is used to characterize the language to which the speech data to be recognized belongs. Among them, the language recognition model is adjusted by introducing a constraint loss to increase the speech feature distance between different second languages ​​and reduce the speech feature distance within the same second language. The second language includes the target language, that is, the constraint loss amplifies the recognition ability and recognition sensitivity of the language recognition model to the target language. Therefore, even if the samples of the target language are small when training the language recognition model, the language recognition model obtained by training can still accurately recognize the speech of the target language. In addition, when using the language recognition model for language recognition, even if the speech of the target language to be recognized is relatively short, the language recognition model can still accurately recognize the speech. Therefore, it has a good recognition effect on the target language and improves the recognition accuracy of the target language. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a flow chart of an embodiment of a language identification method provided by the present application;

[0011] Figure 2 It is a flowchart of an embodiment of a training method for a language recognition model provided by the present application;

[0012] Figure 3 It is a schematic diagram of the framework of an embodiment of the language identification device of the present application;

[0013] Figure 4 It is a schematic diagram of the framework of an embodiment of a language recognition device of the present application;

[0014] Figure 5 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0015] In order to make the purpose, technical solution and effect of the present application clearer and more specific, the present application is further described in detail below with reference to the accompanying drawings and examples.

[0016] It should be noted that the term "several" in this article means at least one, and the terms "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. The term "and / or" is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of types. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0017] See also Figure 1 , Figure 1 is a flow chart of an embodiment of the language recognition method provided by the present application. It should be noted that if there are substantially the same results, this embodiment does not Figure 1 The process sequence shown is limited. Figure 1 As shown, this embodiment includes:

[0018] Step S11: Acquire speech data to be recognized.

[0019] The speech data to be recognized may be provided by the target user or determined by the language recognition device itself. Specifically, the language recognition device may be provided with a human-computer interaction interface, and the speech data to be recognized input by the user may be obtained through the human-computer interaction interface. Of course, in other embodiments, the language recognition device has a communication module, and the language recognition device receives the speech data to be recognized sent by the terminal device through the communication module. The terminal device has a human-computer interaction function, and the user uses the human-computer interaction function to input the speech data to be recognized to be sent to the language recognition device on the terminal device. Among them, the speech data to be recognized may be long speech data or short speech data.

[0020] Step S12: extracting features of the speech data to be recognized using the language recognition model to obtain features of the speech to be recognized.

[0021] In one embodiment, the speech data to be recognized can be divided into several segments of speech data and input into the above-mentioned language recognition model for feature extraction. For example, the speech data to be recognized can be divided into several segments of speech data of fixed length and input into the language recognition model for feature extraction.

[0022] In another embodiment, the speech data to be recognized may be preprocessed to reduce invalid data therein, for example, invalid data such as silence and noise in the speech data to be recognized may be filtered out.

[0023] The language recognition model in this article can be a pre-trained model, that is, the speech data to be recognized can be input into the pre-trained model for feature extraction to obtain the speech features to be recognized. In one embodiment, feature extraction can be performed using a feature extraction network in the language recognition model. For example, a time delay neural network (TDNN), a convolutional neural network (CNN), a residual network (ResNet), etc.

[0024] In a specific implementation, the speech data to be recognized is preprocessed to remove invalid data. The preprocessed speech data to be recognized is divided into several segments of speech data of fixed length, and each segment is called a frame. That is, for each frame of speech data to be recognized, the filter group feature corresponding to each frame of speech data to be recognized is obtained through the time-delay neural network in the language recognition model.

[0025] (Filter Bank), the filter bank features are used as the features of the language to be recognized. It can be understood that when the speech data to be recognized is divided into several segments of speech data, there are also several speech features to be recognized, and each segment of speech data to be recognized corresponds to each speech feature to be recognized.

[0026] Step S13: using a language recognition model to perform recognition based on the features of the speech to be recognized, and obtaining a first language recognition result, wherein the first language recognition result is used to characterize the language to which the speech data to be recognized belongs.

[0027] Among them, the language recognition model can recognize several first languages, and the several first languages ​​include at least one target language. The language recognition model is obtained by adjusting parameters using at least constrained loss. The constrained loss is determined based on the first sample feature extracted from the first training set by the language recognition model. The first training set includes sample speech data of at least two second languages, and at least two second languages ​​include the target language. The goal of adjusting parameters using constrained loss is to increase the speech feature distance between different second languages ​​and reduce the speech feature distance within the same second language.

[0028] The target language refers to a minority language, that is, a language with a relatively small amount of data in the training data set for constructing a language recognition model, or a language with relatively few users. In one embodiment, before training the language recognition model, sample speech data corresponding to each first language is pre-stored as a training data set. In the training data set, a language with less sample speech data is determined as the target language. For example, in the training data set, the duration of the sample speech data corresponding to each first language is at least one hour. Among them, the sample speech data corresponding to Chinese is 4 hours, and the sample speech data corresponding to Czech is 1 hour. Czech can be used as the above-mentioned target language.

[0029] The second language refers to each target language, or the target language and the remaining other languages ​​excluding the target language. That is, the constraint loss parameter is used to increase the speech feature distance between different target languages ​​and reduce the speech feature distance within the same target language. Alternatively, the constraint loss parameter is used to increase the speech feature distance between different target languages ​​and other languages, reduce the speech feature distance within the target language, and reduce the speech feature distance within other languages. For example, the target languages ​​are Czech and Malay, and the constraint loss parameter is used to increase the speech feature distance between Czech and Malay, reduce the speech feature distance within Czech, and reduce the speech feature distance within Malay. For another example, the first language is Czech, Chinese, and English. The target language is Czech, and the other languages ​​are Chinese and English. The constraint loss parameter is used to increase the speech feature distance between Czech and other languages, reduce the speech feature distance within Czech, and reduce the speech feature distance within other languages.

[0030] The target training set can be repeatedly obtained from the training data set, and the language recognition model can be trained using the target training set, or the language recognition model can be directly trained using the training data set. In one embodiment, the target training set is a first training set and a second training set, and the first training set and the second training set are obtained, wherein the second training set includes at least one sample speech data corresponding to the first language. The language recognition model is used to extract features from each sample speech data in the first training set to obtain the first sample features of each sample speech data in the first training set, and the constraint loss of the language recognition model is determined using each first sample feature. Also, the language recognition model is used to extract features from each sample speech data in the second training set to obtain the second sample features of each sample speech data in the second training set, and recognition is performed based on each second sample feature to obtain the predicted language corresponding to each sample speech data in the second training set. Based on the difference between the predicted language and the marked language corresponding to each sample speech data in the second training set, the prediction loss of the language recognition model is determined. The constraint loss and prediction loss of the language recognition model are used to adjust the network parameters of the language recognition model. Repeat the above steps to iteratively train the language recognition model. Among them, the feature extraction network in the language recognition model can be used to extract features from the sample speech data in the first training set and the second training set.

[0031] The sample speech data of each second language can be divided into first sample data and second sample data to calculate the constraint loss. In one embodiment, the sample speech data of each second language is divided into first sample speech data and second sample speech data, and the first sample feature of the first sample speech data of the second language is used to calculate the language sample feature of the second language. The constraint loss is positively correlated with the first feature distance of each second sample speech data and negatively correlated with the second feature distance of the second sample speech data. The first feature distance represents the distance between the first sample feature of the second sample speech data and the language sample feature corresponding to the second language to which it belongs, and the second feature distance represents the distance between the first sample feature of the second sample speech data and the language sample feature corresponding to the second language not belonging to it. Wherein, the above-mentioned first sample feature and second sample feature can be sample feature vectors.

[0032] In a specific embodiment, the constraint loss is the sum of the total feature distances of each second sample speech data, and the total feature distance of the second sample speech data is equal to the first feature distance of the second sample speech data minus the second feature distance between the second sample speech data and each non-belonging second language.

[0033] In another specific implementation, the language sample feature of the second language is the central tendency value of the first sample feature of each first sample speech data of the second language. For example, the average value of the first sample feature of each first sample speech data of the second language is used as the central tendency value. For another example, the standard deviation of the first sample feature of each first sample speech data of the second language is used as the central tendency value. Of course, other values ​​that can characterize the changing trend of each first sample feature can be used as the central tendency value, and there is no limitation here.

[0034] The language recognition model can recognize multiple languages ​​or one first language.

[0035] In the case where there are multiple first languages, that is, the language recognition model can recognize multiple first languages. Among them, the first training set and the second training set can be obtained from the pre-stored training data set. In one embodiment, from the pre-stored first number of sample speech data corresponding to each target language, a second number of sample speech data of the target language is randomly selected as the sample speech data corresponding to each second language in the first training set, and the number of each second language extracted is the third number. And from the pre-stored sample speech data corresponding to each first language, a fourth number of sample speech data of the first language is randomly selected to obtain the second training set, wherein the sample speech data extracted from each first language is the fifth number. In a specific embodiment, the total amount of sample speech data in the first training set obtained by each iterative training is fixed and consistent.

[0036] Of course, in the case where the first language is one, that is, the language recognition model can recognize one first language. In one embodiment, the first training set and the second training set can be randomly extracted from pre-stored sample speech data corresponding to the first language and the other language. In a specific embodiment, the total amount of sample speech data corresponding to the first language and the other language obtained in the first training set or the second training set is the same.

[0037] Taking the case where there are multiple first languages, obtaining the first training set as an example, randomly extracting sample speech data corresponding to the second number of N target languages ​​from the sample speech data corresponding to the first number of M target languages, and extracting a third number of L sample speech data from the second number of N target languages, where N*L=D, and the size of D remains unchanged. The sample speech data is used as the first training set.

[0038] The constraint loss and prediction loss of the language recognition model can be used to adjust the network parameters of the language recognition model. In a specific implementation, the weights of the constraint loss and prediction loss of the language recognition model in this iteration are obtained respectively, wherein the sum of the weight corresponding to the constraint loss and the weight corresponding to the prediction loss is a fixed value. When the current iteration is within the first number, the weight corresponding to the prediction loss is a preset weight value. When the current iteration is greater than the first number and less than the maximum number of iterations, the weight corresponding to the prediction loss is the difference between the preset weight value and the number ratio, and the number ratio is the ratio between the current iteration number and the maximum iteration number. The constraint loss and prediction loss of the language recognition model are weighted and summed using the weights to obtain the total loss of the language recognition model. The total loss of the language recognition model is used to adjust the network parameters of the language recognition model.

[0039] Among them, when the value of the total loss tends to be stable or the number of iterations reaches the maximum number of iterations, the adjustment of the network parameters of the language recognition model is completed.

[0040] By using the language recognition model to perform recognition based on the speech features to be recognized, a corresponding first language recognition result can be obtained. If the target language exists in the language corresponding to the corresponding first language recognition result, the first language recognition result can be reconfirmed.

[0041] In one embodiment, there are multiple first languages, and the language recognition model corresponds to a multi-language recognition model, and each second language is a target language. That is, the multi-language recognition model can recognize multiple first languages. The multi-language recognition model is used to recognize the speech features to be recognized, and the corresponding first language recognition results are obtained. For a specific description of the multi-language recognition model, reference can be made to the relevant description of the language recognition model in step S13.

[0042] In the case where the target language exists in the language corresponding to the corresponding first language recognition result, the language recognition model corresponding to the target language can be used to recognize the speech features to be recognized for secondary confirmation. In one embodiment, the first language is one and is a target language, and the language recognition model corresponds to a binary classification model corresponding to the target language. That is, the binary classification model corresponding to the target language can recognize a first language (target language), and the second language is the target language and other languages, wherein the first language except the target language is regarded as other languages. The goal of using constraint loss parameter adjustment is to increase the speech feature distance between the target language and other languages, reduce the speech feature distance within the target language, and reduce the speech feature distance within other languages.

[0043] In a specific embodiment, each target language is used as a language to be trained. A third training set is obtained, wherein the third training set includes sample speech data of at least two first languages, and at least two first languages ​​include the language to be trained, and the first languages ​​other than the language to be trained in the third training set are all used as other languages. The feature extraction network of the language recognition model is used to extract features of each sample speech data in the third training set, and the third sample features of each sample speech data in the third training set are obtained. The binary classification model corresponding to the language to be trained is used to perform recognition based on each third sample feature, and the language prediction result corresponding to each sample speech data in the third training set is obtained, and the language prediction result represents whether the sample speech data belongs to the language to be trained or the other language. The constraint loss of the binary classification model is determined using each third sample feature. And, based on the language prediction result and the labeled language result corresponding to each sample speech data in the third training set, the prediction loss of the binary classification model is determined. The network parameters of the binary classification model are adjusted using the constraint loss and prediction loss of the binary classification model. Repeat the aforementioned steps of obtaining the third training set and the subsequent steps to iteratively train the binary classification model corresponding to the language to be trained. The third sample feature may be a sample feature vector.

[0044] The number of samples of the sample speech data selected for the language to be trained and other languages ​​corresponding to the third training set can be the same. For example, the first languages ​​are Chinese, English, and Czech, where Czech is the language to be trained. Randomly extract N sample speech data from the sample speech data corresponding to Czech, and randomly extract a total of N sample speech data from the sample speech data corresponding to Chinese and English. The sample speech data extracted above is used as the corresponding third training set.

[0045] In another embodiment, the first language is multiple, that is, a multi-language recognition model is used for language recognition. In response to the first language recognition result indicating that the speech data to be recognized belongs to the target language, a second recognition is performed on the speech features to be recognized using a binary classification model corresponding to the target language to obtain a second language recognition result, and the second language recognition result indicates whether the speech data to be recognized belongs to the target language to be recognized. In response to the first language recognition result indicating that the speech data to be recognized does not belong to the target language to be recognized, the speech data to be recognized is determined to be the language represented by the first language recognition result.

[0046] The speech data to be recognized is processed in segments, and there may be several corresponding speech features to be recognized, that is, there may be several corresponding first language recognition results. In the above several corresponding first language recognition results, if the target language exists in the languages ​​corresponding to the several corresponding first language recognition results, the first language recognition result may be confirmed for a second time.

[0047] In one embodiment, the first language recognition result includes the confidence that the speech data to be recognized belongs to each first language. After obtaining the second language recognition result, in response to the second language recognition result indicating that the speech data to be recognized belongs to the recognized target language, the speech data to be recognized is determined to be the recognized target language. In response to the second language recognition result indicating that the speech data to be recognized does not belong to the recognized target language, the confidence of each remaining first language other than the recognized target language is obtained from the first language recognition result, and the language to which the speech data to be recognized belongs is determined based on the confidence of each remaining first language.

[0048] In one specific implementation, the average value of the confidence of each first language can be obtained, and the first language corresponding to the largest average value is selected to be determined as the language to which the speech data to be recognized belongs. In another specific implementation, the first language corresponding to the largest confidence is selected to be determined as the language to which the speech data to be recognized belongs.

[0049] For example, Czech is the target language. The speech data to be recognized is segmented and feature extracted to obtain a number of speech features to be recognized. The multilingual recognition model is used to recognize each speech data to be recognized, and the first language recognition results obtained are 90 points for Chinese, 50 points for English, 80 points for Chinese, and 50 points for Czech. The speech features to be recognized corresponding to Czech are input into the binary classification model corresponding to Czech for recognition. When the second language recognition result is other languages, the language to which the speech data to be recognized belongs is determined from the first language recognition results of 90 points for Chinese, 50 points for English, and 80 points for Chinese. Among them, the average confidence value corresponding to Chinese is 85 points, and the average confidence value corresponding to English is 50 points, so Chinese is taken as the language to which the speech data to be recognized belongs.

[0050] See also Figure 2 , Figure 2 1 is a flow chart of an embodiment of a training method for a language recognition model provided in the present application. The language recognition model can recognize a plurality of first languages, wherein the plurality of first languages ​​include at least one target language. It should be noted that if there are substantially the same results, this embodiment does not use Figure 2 The process sequence shown is limited. Figure 2 As shown, this embodiment includes:

[0051] Step S21: Obtain a first training set and a second training set.

[0052] The first training set includes sample speech data of at least two second languages, and the at least two second languages ​​include the target language, and the second training set includes sample speech data corresponding to at least one first language.

[0053] The first training set and the second training set can be obtained from sample speech data in a pre-stored training data set, wherein the sample speech data in the pre-stored training data set includes sample speech data corresponding to the target language.

[0054] Before pre-storing the sample speech data corresponding to the training data set, each sample speech data may be pre-processed to reduce invalid data. The first training set and the second training set are obtained from the sample speech data corresponding to the pre-processed training data set.

[0055] The target language is determined based on the data volume of the sample speech data corresponding to the training data set. The method for determining the target language may refer to the relevant description of step S13.

[0056] After the sample speech data is preprocessed, the sample speech data corresponding to the target language can be speed-changed. Of course, the sample speech data corresponding to the target language can also be speed-changed before the sample speech data is preprocessed to obtain more sample speech data. There is no limitation here. For example, the sample speech data corresponding to the target language is speed-changed by 0.9 times and 1.1 times, and the sample speech data corresponding to 0.9 times speed, 1.1 times speed, and 1 times speed are collectively used as the sample speech data corresponding to the target language. The sample speech data corresponding to the target language that has been speed-changed is preprocessed.

[0057] For the method of obtaining the first training set and the second training set, reference may be made to the relevant description in step S13, which will not be described in detail here.

[0058] Taking the acquisition of the first training set as an example when there are multiple first languages, when there are multiple first languages, the language recognition model corresponds to a multilingual recognition model, and the second languages ​​are all target languages, the first languages ​​are Chinese, English, Japanese, Czech, Malay, and Arabic. The target languages ​​are Czech and Malay. Randomly select 2 from Czech, Malay, and Arabic, such as Malay and Arabic. Then extract 6 sample speech data from Malay and Arabic respectively as the first training set. Alternatively, randomly select 3 from Czech, Malay, and Arabic, such as Malay, Czech and Arabic. Then extract 4 sample speech data from Czech, Malay, and Arabic respectively as the first training set. Among them, the sample speech data extracted in each iteration remains unchanged at 12.

[0059] Step S22: extracting features from each sample speech data in the first training set using the language recognition model to obtain first sample features of each sample speech data in the first training set, and determining the constraint loss of the language recognition model using each first sample feature.

[0060] The feature extraction network of the language recognition model can be used to extract features from each sample speech data in the first training set to obtain first sample features corresponding to each sample speech data.

[0061] For details on determining the constraint loss of the language recognition model using the first sample feature, please refer to the relevant description in step S13.

[0062] Taking the acquisition of constraint loss as an example, when there are multiple first languages, the language recognition model corresponds to a multilingual recognition model, and the second languages ​​are all target languages, 6 sample speech data are extracted from Malay and Arabic as the first training set. The 6 sample speech data corresponding to Malay are divided into the first sample speech data (3 sample speech data) and the second sample speech data (another 3 sample speech data). The 6 sample data corresponding to Arabic are also divided into the first sample speech data and the second sample speech data in the above manner. The first sample features corresponding to the first sample data corresponding to Malay are averaged to obtain the language sample features. In the same way, the language sample features corresponding to Arabic are obtained. The cosine distance between each second sample feature corresponding to Malay and the language sample feature corresponding to Malay is used as the first feature distance. The cosine distance between each second sample feature corresponding to Malay and the language sample feature corresponding to Arabic is used as the second feature distance. In the same way as above, the first feature distance and the second feature distance corresponding to Arabic are obtained. The sum of the first feature distance corresponding to Malay and the first feature distance corresponding to Arabic is subtracted from the sum of the second feature distance corresponding to Malay and the second feature distance corresponding to Arabic, and the final value is calculated as the sum of the total features, which is the constraint loss.

[0063] Step S23: Use the language recognition model to extract features from each sample speech data in the second training set to obtain second sample features of each sample speech data in the second training set, and perform recognition based on each second sample feature to obtain the predicted language corresponding to each sample speech data in the second training set. Based on the difference between the predicted language and the labeled language corresponding to each sample speech data in the second training set, determine the prediction loss of the language recognition model.

[0064] For details on determining the prediction loss of the language recognition model using the second sample feature, please refer to the relevant description in step S13.

[0065] Step S24: using the constraint loss and prediction loss of the language recognition model to adjust the network parameters of the language recognition model.

[0066] Among them, the goal of using constrained loss parameter adjustment is to increase the speech feature distance between different second languages ​​and reduce the speech feature distance within the same second language.

[0067] For a detailed description of adjusting the network parameters of the language recognition model by using the constraint loss and prediction loss of the language recognition model, please refer to the relevant description in step S13.

[0068] The total loss can be obtained by using the constraint loss and prediction loss. Taking the acquisition of the total loss Loss_all as an example, the weight of the prediction loss LimitLoss is θ, and the weight of the constraint loss CELoss is 1-θ. The expression of the total loss is Loss_all = θ*CELoss + (1-θ)*LimitLoss. Among them, θ is a piecewise function, specifically: In the above formula, iter is the current iteration number, and iter_max is the set maximum iteration number.

[0069] Taking the adjustment of network parameters as an example, when the value of the total loss Loss_all tends to be stable, or the current iteration number iter reaches the maximum iteration number iter_max, the adjustment of the network parameters of the language recognition model is completed.

[0070] The piecewise function can be changed according to the actual situation, that is, the first number can be changed. For example, the first number is changed from 5 to 2, that is, when iter>2, it starts to enter Attenuation is performed, and no restriction is imposed here.

[0071] Before training the language recognition model, an initial learning rate can be set in advance, and then the network parameters of the model can be dynamically adjusted and optimized based on the total loss value of the model.

[0072] See also Figure 3 , Figure 3 It is a schematic diagram of a framework of an embodiment of a language recognition device of the present application. The language recognition device 300 includes a speech data acquisition module 310, a speech feature acquisition module 320, and a language recognition module 330. The speech data acquisition module 310 is used to acquire speech data to be recognized. The speech feature acquisition module 320 is used to extract features of the speech data to be recognized using a language recognition model to obtain speech features to be recognized. The language recognition module 330 is used to perform recognition based on the speech features to be recognized using the language recognition model to obtain a first language recognition result, and the first language recognition result is used to characterize the language to which the speech data to be recognized belongs, wherein the language recognition model can recognize a plurality of first languages, and the plurality of first languages ​​include at least one target language, and the language recognition model is obtained by adjusting parameters using at least a constraint loss, and the constraint loss is determined based on the first sample feature extracted from the first training set by the language recognition model, and the first training set includes sample speech data of at least two second languages, and at least two second languages ​​include the target language, and the goal of adjusting parameters using the constraint loss is to increase the speech feature distance between different second languages ​​and reduce the speech feature distance within the same second language.

[0073] In some embodiments, the language recognition device 300 further includes a training module. Before the speech feature acquisition module 320 performs feature extraction on the speech data to be recognized using the language recognition model to obtain the speech features to be recognized, the training module performs acquisition of a first training set and a second training set, wherein the second training set includes at least one sample speech data corresponding to the first language. The language recognition model is used to perform feature extraction on each sample speech data in the first training set to obtain a first sample feature of each sample speech data in the first training set, and the constraint loss of the language recognition model is determined using each first sample feature. Also, the language recognition model is used to perform feature extraction on each sample speech data in the second training set to obtain a second sample feature of each sample speech data in the second training set, and recognition is performed based on each second sample feature to obtain a predicted language corresponding to each sample speech data in the second training set. Based on the difference between the predicted language and the marked language corresponding to each sample speech data in the second training set, the prediction loss of the language recognition model is determined. The constraint loss and prediction loss of the language recognition model are used to adjust the network parameters of the language recognition model. The aforementioned steps are repeated to iteratively train the language recognition model.

[0074] In some embodiments, each sample speech data of the second language is divided into first sample speech data and second sample speech data, the first sample feature of the first sample speech data of the second language is used to calculate the language sample feature of the second language, the constraint loss is positively correlated with the first feature distance of each second sample speech data, and negatively correlated with the second feature distance of the second sample speech data, the first feature distance represents the distance between the first sample feature of the second sample speech data and the language sample feature corresponding to the second language, and the second feature distance represents the distance between the first sample feature of the second sample speech data and the language sample feature corresponding to the second language not belonging to the second language.

[0075] In some embodiments, the constraint loss is the sum of the total feature distances of each second sample speech data, and the total feature distance of the second sample speech data is equal to the first feature distance of the second sample speech data minus the second feature distance between the second sample speech data and each non-belonging second language.

[0076] In some embodiments, the language sample feature of the second language is a central tendency value of the first sample feature of each first sample speech data of the second language.

[0077] In some embodiments, the training module executes the acquisition of the first training set and the second training set, including: randomly extracting a second number of sample speech data of the target language from the pre-stored first number of sample speech data corresponding to each target language as the sample speech data corresponding to each second language in the first training set, and the number of each second language extracted is the third number. And randomly extracting a fourth number of sample speech data of the first language from the pre-stored sample speech data corresponding to each first language to obtain the second training set, wherein the sample speech data extracted from each first language is the fifth number.

[0078] In some embodiments, the training module uses the constraint loss and prediction loss of the language recognition model to adjust the network parameters of the language recognition model, including: respectively obtaining the weights of the constraint loss and prediction loss of the language recognition model in this iteration, wherein the sum of the weight corresponding to the constraint loss and the weight corresponding to the prediction loss is a fixed value, and when the current iteration is within the first number, the weight corresponding to the prediction loss is a preset weight value, and when the current iteration is greater than the first number and less than the maximum number of iterations, the weight corresponding to the prediction loss is the difference between the preset weight value and the number ratio, and the number ratio is the ratio between the current iteration number and the maximum number of iterations. The constraint loss and prediction loss of the language recognition model are weighted and summed using the weights to obtain the total loss of the language recognition model. The network parameters of the language recognition model are adjusted using the total loss of the language recognition model.

[0079] In some embodiments, there are multiple first languages, the language recognition model is a corresponding multi-language recognition model, and each second language is the target language.

[0080] In some embodiments, the first language is one and is a target language, and the language recognition model corresponds to a binary classification model corresponding to the target language.

[0081] In some embodiments, the first language is multiple. After the language recognition module 330 uses the language recognition model to perform recognition based on the speech features to be recognized and obtains the first language recognition result, in response to the first language recognition result indicating that the speech data to be recognized belongs to the target language, a second recognition is performed on the speech features to be recognized using a binary classification model corresponding to the target language to obtain a second language recognition result, and the second language recognition result indicates whether the speech data to be recognized belongs to the recognized target language. In response to the first language recognition result indicating that the speech data to be recognized does not belong to the recognized target language, the speech data to be recognized is determined to be the language indicated by the first language recognition result.

[0082] In some embodiments, the first language recognition result includes the confidence that the speech data to be recognized belongs to each of the first languages. After the language recognition module 330 obtains the second language recognition result, it is executed in response to the second language recognition result indicating that the speech data to be recognized belongs to the recognized target language, and the speech data to be recognized is determined to be the recognized target language. In response to the second language recognition result indicating that the speech data to be recognized does not belong to the recognized target language, the confidence of each remaining first language other than the recognized target language is obtained from the first language recognition result, and the language to which the speech data to be recognized belongs is determined based on the confidence of each remaining first language.

[0083] In some embodiments, the language recognition device 300 further includes a training module, which is used to respectively use each target language as a language to be trained. A third training set is obtained, wherein the third training set includes sample speech data of at least two first languages, and at least two first languages ​​include the language to be trained, and the first languages ​​other than the language to be trained in the third training set are all used as other languages. The feature extraction network of the language recognition model is used to extract features from each sample speech data in the third training set, and a third sample feature of each sample speech data in the third training set is obtained. The binary classification model corresponding to the language to be trained is used to perform recognition based on each third sample feature, and a language prediction result corresponding to each sample speech data in the third training set is obtained, and the language prediction result represents whether the sample speech data belongs to the language to be trained or other languages. The constraint loss of the binary classification model is determined using each third sample feature. And, based on the language prediction result and the labeled language result corresponding to each sample speech data in the third training set, the prediction loss of the binary classification model is determined. The network parameters of the binary classification model are adjusted using the constraint loss and prediction loss of the binary classification model. Repeat the aforementioned steps of obtaining the third training set and subsequent steps to iteratively train the binary classification model corresponding to the training language.

[0084] See also Figure 4 , Figure 4 4 is a schematic diagram of a framework of an embodiment of a language recognition device of the present application. The language recognition device 40 includes a memory 41 and a processor 42 coupled to each other, and the processor 42 is used to execute program instructions stored in the memory 41 to implement the steps in any of the above-mentioned language recognition method embodiments. In a specific implementation scenario, the language recognition device 40 may include but is not limited to: a microcomputer, a server, and in addition, the language recognition device 40 may also include a mobile device such as a laptop computer and a tablet computer, which is not limited here.

[0085] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned language recognition method embodiments. The processor 42 can also be called a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 42 can be implemented by an integrated circuit chip.

[0086] See also Figure 5 , Figure 5 The computer-readable storage medium 50 stores program instructions 51 that can be executed by a processor, and the program instructions 51 are used to implement the steps in any of the above-mentioned language recognition method embodiments.

[0087] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0088] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0089] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0090] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0091] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

Claims

1. A language recognition method, characterized in that: The method comprises: Obtaining voice data to be recognized; Extracting features of the speech data to be recognized using a language recognition model to obtain features of the speech to be recognized; Using the language recognition model to perform recognition based on the features of the speech to be recognized, to obtain a first language recognition result, wherein the first language recognition result is used to characterize the language to which the speech data to be recognized belongs; In which, the language recognition model can recognize several first languages, and the several first languages ​​include at least one target language. The language recognition model is obtained by adjusting parameters using at least constraint loss, and the constraint loss is determined based on the first sample feature extracted from the first training set by the language recognition model. The first training set includes sample speech data of at least two second languages, and the at least two second languages ​​include the target language. The goal of adjusting parameters using the constraint loss is to increase the speech feature distance between different second languages ​​and to reduce the speech feature distance within the same second language.

2. The method according to claim 1, characterized in that Before extracting features of the speech data to be recognized by using the language recognition model to obtain the features of the speech to be recognized, the method further includes: Acquire the first training set and the second training set, wherein the second training set includes at least one sample speech data corresponding to the first language; Using the language recognition model to extract features from each of the sample speech data in the first training set, obtaining first sample features of each of the sample speech data in the first training set, and using each of the first sample features to determine the constraint loss of the language recognition model; and, Using the language recognition model to extract features from each of the sample speech data in the second training set, obtain second sample features of each of the sample speech data in the second training set, and perform recognition based on each of the second sample features to obtain a predicted language corresponding to each of the sample speech data in the second training set; based on the difference between the predicted language and the labeled language corresponding to each of the sample speech data in the second training set, determine the prediction loss of the language recognition model; Using the constraint loss and prediction loss of the language recognition model, adjusting the network parameters of the language recognition model; Repeat the above steps to iteratively train the language recognition model.

3. The method according to claim 1 or 2, characterized in that: Each sample speech data of the second language is divided into first sample speech data and second sample speech data. The first sample feature of the first sample speech data of the second language is used to calculate the language sample feature of the second language. The constraint loss is positively correlated with the first feature distance of each second sample speech data and negatively correlated with the second feature distance of the second sample speech data. The first feature distance represents the distance between the first sample feature of the second sample speech data and the language sample feature corresponding to the second language. The second feature distance represents the distance between the first sample feature of the second sample speech data and the language sample feature corresponding to the second language.

4. The method according to claim 3, characterized in that: The constraint loss is the sum of the total feature distances of each of the second sample speech data, and the total feature distance of the second sample speech data is equal to the first feature distance of the second sample speech data minus the second feature distance between the second sample speech data and each non-second language; And / or, the language sample feature of the second language is a central tendency value of the first sample feature of each first sample speech data of the second language.

5. The method according to claim 2, characterized in that: The obtaining of the first training set and the second training set comprises: Randomly extracting a second number of sample speech data in the target language from the first number of pre-stored sample speech data corresponding to each of the target languages ​​as the sample speech data corresponding to each of the second languages ​​in the first training set, the number of sample speech data extracted for each of the second languages ​​being a third number; and Randomly extracting a fourth number of sample speech data in the first language from the pre-stored sample speech data corresponding to each of the first languages ​​to obtain the second training set, wherein the sample speech data extracted in each of the first languages ​​is a fifth number; And / or, using the constraint loss and prediction loss of the language recognition model to adjust the network parameters of the language recognition model includes: Obtain the weights of the constraint loss and the prediction loss of the language recognition model in this iteration respectively, wherein the sum of the weight corresponding to the constraint loss and the weight corresponding to the prediction loss is a fixed value, and when this iteration is within the first number, the weight corresponding to the prediction loss is a preset weight value, and when this iteration is greater than the first number and less than the maximum number of iterations, the weight corresponding to the prediction loss is the difference between the preset weight value and the number ratio, and the number ratio is the ratio between the number of this iteration and the maximum number of iterations; Using the weights, weighted summation is performed on the constraint loss and the prediction loss of the language recognition model to obtain a total loss of the language recognition model; The total loss of the language recognition model is used to adjust the network parameters of the language recognition model.

6. The method according to claim 1, characterized in that There are multiple first languages, and the language recognition model corresponds to a multi-language recognition model, and each of the second languages ​​is the target language; or, there is only one first language, and it is the target language, and the language recognition model corresponds to a binary classification model corresponding to the target language.

7. The method according to claim 6, characterized in that The first language is multiple; After the language recognition model is used to perform recognition based on the speech feature to be recognized to obtain a first language recognition result, the method further includes: In response to the first language recognition result indicating that the speech data to be recognized belongs to the target language, a binary classification model corresponding to the target language is used to perform secondary recognition on the speech features to be recognized, so as to obtain a second language recognition result, wherein the second language recognition result indicates whether the speech data to be recognized belongs to the target language to be recognized; In response to the first language recognition result indicating that the speech data to be recognized does not belong to the target language to be recognized, it is determined that the speech data to be recognized is of the language indicated by the first language recognition result.

8. The method according to claim 7, characterized in that The first language recognition result includes the confidence level that the speech data to be recognized belongs to each of the first languages; After obtaining the second language recognition result, the method further includes: In response to the second language recognition result indicating that the speech data to be recognized belongs to the target language to be recognized, determining that the speech data to be recognized is the target language to be recognized; In response to the second language recognition result indicating that the speech data to be recognized does not belong to the target language to be recognized, the confidences of the remaining first languages ​​other than the recognized target language are obtained from the first language recognition result, and the language to which the speech data to be recognized belongs is determined based on the confidences of the remaining first languages.

9. The method according to claim 7, characterized in that: The method further comprises: respectively taking each of the target languages ​​as the languages ​​to be trained; Acquire a third training set, wherein the third training set includes sample speech data of at least two first languages, and the at least two first languages ​​include the language to be trained, and the first languages ​​in the third training set other than the language to be trained are all regarded as other languages; Using the feature extraction network of the language recognition model to extract features from each of the sample speech data in the third training set, to obtain third sample features of each of the sample speech data in the third training set; Using the binary classification model corresponding to the language to be trained to perform identification based on each of the third sample features, a language prediction result corresponding to each of the sample speech data in the third training set is obtained, wherein the language prediction result indicates whether the sample speech data belongs to the language to be trained or the other language; Determining the constraint loss of the binary classification model using each of the third sample features; and determining the prediction loss of the binary classification model based on the language prediction results and the labeled language results corresponding to each of the sample speech data in the third training set; Using the constraint loss and prediction loss of the binary classification model, adjusting the network parameters of the binary classification model; Repeat the aforementioned steps of obtaining the third training set and subsequent steps to iteratively train the binary classification model corresponding to the language to be trained.

10. A method for training a language recognition model, characterized in that: The language recognition model is capable of recognizing a plurality of first languages, wherein the plurality of first languages ​​includes at least one target language; and the method comprises: Acquire a first training set and a second training set, wherein the first training set includes sample speech data in at least two second languages, and the at least two second languages ​​include the target language, and the second training set includes sample speech data corresponding to at least one of the first languages; Using the language recognition model to extract features from each of the sample speech data in the first training set, obtaining first sample features of each of the sample speech data in the first training set, and using each of the first sample features to determine the constraint loss of the language recognition model; and, Using the language recognition model to extract features from each of the sample speech data in the second training set, obtain second sample features of each of the sample speech data in the second training set, and perform recognition based on each of the second sample features to obtain a predicted language corresponding to each of the sample speech data in the second training set; based on the difference between the predicted language and the labeled language corresponding to each of the sample speech data in the second training set, determine the prediction loss of the language recognition model; The constraint loss and prediction loss of the language recognition model are used to adjust the network parameters of the language recognition model, wherein the goal of adjusting the parameters using the constraint loss is to increase the speech feature distance between different second languages ​​and to reduce the speech feature distance within the same second language.

11. A language recognition device, characterized in that: The device comprises: A voice data acquisition module, used to acquire voice data to be recognized; A speech feature acquisition module, used to extract features of the speech data to be recognized using a language recognition model to obtain speech features to be recognized; A language recognition module is used to use the language recognition model to perform recognition based on the speech features to be recognized to obtain a first language recognition result, wherein the first language recognition result is used to characterize the language to which the speech data to be recognized belongs, wherein the language recognition model can recognize several first languages, and the several first languages ​​include at least one target language. The language recognition model is obtained by adjusting parameters using at least a constraint loss, and the constraint loss is determined based on the first sample features extracted from a first training set by the language recognition model. The first training set includes sample speech data of at least two second languages, and the at least two second languages ​​include the target language. The goal of adjusting parameters using the constraint loss is to increase the speech feature distance between different second languages ​​and to reduce the speech feature distance within the same second language.

12. A language recognition device, characterized in that: The language recognition device includes a memory and a processor, the memory stores program instructions, and the processor is used to execute the program instructions to implement the language recognition method according to any one of claims 1 to 9.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program instructions, and the program instructions can be executed to implement the language recognition method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Model training method, language identification method, language identification device and equipment

    CN110838286A

  • Language recognition method, related equipment and readable storage medium

    CN111724766A

  • Language recognition method and device and language recognition model training method and device

    CN113724700A

  • Label checking method, related device, electronic equipment and storage medium

    CN115050350A

  • Speech recognition model training method, speech recognition method and related equipment

    CN117854486A