A language recognition method, a training method, a device, and a storage medium
By introducing constraint loss parameter tuning into the language recognition model, the distance between speech features of different languages is increased, and the distance between speech features of the same language is reduced, which solves the problem of low accuracy in language recognition in the existing technology and achieves efficient recognition of the target language.
Patent Information
- Application Number
- CN202510017109.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-01-06
AI Technical Summary
Existing language recognition methods have limitations in processing speech features between and within different languages, resulting in low accuracy in identifying the target language.
A language recognition model is used for feature extraction and recognition. By adjusting the parameters of the constraint loss model, the distance between speech features of different languages is increased, while the distance between speech features of the same language is reduced. By introducing constraint loss to adjust the parameters of the language recognition model, the recognition ability and sensitivity of the target language are enhanced.
Even with a limited number of target language samples, the model can still accurately recognize speech, improving the accuracy of target language recognition.
Smart Images

Figure CN119993121B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of speech processing, and in particular to a language recognition method, a training method, a device and a storage medium. BACKGROUND
[0002] Language recognition is a process of determining a language category to which a speech segment belongs. Specifically, speech data to be recognized can be subjected to language recognition, so as to obtain a language category to which the speech data to be recognized belongs. However, a general language recognition method has limitations in processing speech features between different languages and within a language, resulting in low recognition accuracy for a target language. SUMMARY
[0003] The technical problem solved by the present application is to provide a language recognition method, a training method, a device and a storage medium, which can improve the recognition accuracy for a target language.
[0004] To solve the above technical problem, one technical solution adopted by the present application is to provide a language recognition method, which comprises: obtaining speech data to be recognized; extracting features of the speech data to be recognized by using a language recognition model, to obtain speech features to be recognized; and recognizing, by using the language recognition model, based on the speech features to be recognized, to obtain a first language recognition result, the first language recognition result being used to represent a language to which the speech data to be recognized belongs; wherein the language recognition model can recognize a plurality of first languages, the plurality of first languages including at least one target language, the language recognition model being obtained by at least using a constraint loss to adjust parameters, the constraint loss being determined based on first sample features extracted by the language recognition model from a first training set, the first training set including sample speech data of at least two second languages, and the at least two second languages including the target language, and the target of using the constraint loss to adjust parameters being to increase a distance between speech features of different second languages and to reduce a distance between speech features within a same second language.
[0005] To solve the above technical problems, another technical solution adopted by the present application is to provide a language recognition model training method, which comprises: obtaining a first training set and a second training set, wherein the first training set comprises sample voice data of at least two second languages, and the at least two second languages include a target language, and the second training set comprises sample voice data corresponding to at least one first language; using a language recognition model to extract features from each sample voice data in the first training set to obtain first sample features of each sample voice data in the first training set, and using each first sample feature to determine a constraint loss of the language recognition model; and using the language recognition model to extract features from each sample voice data in the second training set to obtain second sample features of each sample voice data in the second training set, and performing recognition based on each second sample feature respectively to obtain a predicted language corresponding to each sample voice data in the second training set; determining a prediction loss of the language recognition model based on the difference between the predicted language and the labeled language corresponding to each sample voice data in the second training set; and adjusting the network parameters of the language recognition model using the constraint loss and the prediction loss, wherein the target of adjusting the parameters using the constraint loss is to increase the voice feature distance between different second languages and to reduce the voice feature distance within the same second language.
[0006] To solve the above technical problems, still another technical solution adopted by the present application is to provide a language recognition device, which comprises: a voice data acquisition module for acquiring voice data to be recognized; a voice feature acquisition module for using a language recognition model to extract features from the voice data to be recognized to obtain voice features to be recognized; and a language recognition module for using the language recognition model to recognize based on the voice features to be recognized to obtain a first language recognition result, wherein the first language recognition result is used to represent the language to which the voice data to be recognized belongs, the language recognition model can recognize a plurality of first languages, the plurality of first languages include at least one target language, the language recognition model is obtained by adjusting parameters using a constraint loss, the constraint loss is determined based on first sample features extracted from a first training set by the language recognition model, the first training set comprises sample voice data of at least two second languages, and the at least two second languages include the target language, and the target of adjusting the parameters using the constraint loss is to increase the voice feature distance between different second languages and to reduce the voice feature distance within the same second language.
[0007] To solve the above technical problems, still another technical solution adopted by the present application is to provide a language recognition device, which comprises: a voice data acquisition module for acquiring voice data to be recognized; a voice feature acquisition module for using a language recognition model to extract features from the voice data to be recognized to obtain voice features to be recognized; and a language recognition module for using the language recognition model to recognize based on the voice features to be recognized to obtain a first language recognition result, wherein the first language recognition result is used to represent the language to which the voice data to be recognized belongs, the language recognition model can recognize a plurality of first languages, the plurality of first languages include at least one target language, the language recognition model is obtained by adjusting parameters using a constraint loss, the constraint loss is determined based on first sample features extracted from a first training set by the language recognition model, the first training set comprises sample voice data of at least two second languages, and the at least two second languages include the target language, and the target of adjusting the parameters using the constraint loss is to increase the voice feature distance between different second languages and to reduce the voice feature distance within the same second language.
[0008] To solve the above technical problems, the application adopts another technical solution: providing a computer readable storage medium for storing program instructions, which can be executed to implement the above language recognition method.
[0009] The above scheme uses a language recognition model to extract features from the to-be-recognized voice data, obtains to-be-recognized voice features, and performs recognition based on the to-be-recognized voice features to obtain a first language recognition result, which is used to represent the language to which the to-be-recognized voice data belongs. The language recognition model is adjusted by introducing a constraint loss to increase the voice feature distance between different second languages and reduce the voice feature distance within the same second language, and the second language includes a target language, that is, the constraint loss increases the recognition ability and recognition sensitivity of the language recognition model for the target language. Therefore, even if the sample of the target language is small when training the language recognition model, the trained language recognition model can still accurately recognize the voice of the target language. In addition, when using the language recognition model to recognize the language, even if the voice of the target language to be recognized is relatively short, the language recognition model can still accurately recognize the voice. Therefore, the target language has good recognition effect, and the recognition accuracy of the target language is improved. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1 is a flowchart of an embodiment of the language recognition method provided by the application;
[0011] Figure 2 is a flowchart of an embodiment of the training method of the language recognition model provided by the application;
[0012] Figure 3 is a framework diagram of an embodiment of the language recognition device provided by the application;
[0013] Figure 4 is a framework diagram of an embodiment of the language recognition device provided by the application;
[0014] Figure 5 is a framework diagram of an embodiment of the computer readable storage medium provided by the application. DETAILED DESCRIPTION
[0015] To make the purpose, technical solutions and effects of the application clearer and more explicit, the application is further described in detail below with reference to the drawings and embodiments.
[0016] It should be noted that the term "several" herein represents at least one, and the terms "first", "second", etc. are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. The term "and / or" is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in this paper generally represents an "or" relationship between the front and rear associated objects. In addition, the term "at least one" herein represents any one of a plurality of combinations of at least two of any one or more, for example, including at least one of A, B, and C, which can represent any one or more elements selected from the set consisting of A, B, and C.
[0017] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the language recognition method provided by the present application. It should be noted that the order of the flowchart shown in the embodiment is not limited if there is substantially the same result. As shown in Figure 1 , the embodiment includes the following steps: Figure 1
[0018] Step S11: obtaining voice data to be recognized.
[0019] The voice data to be recognized can be provided by a target user or determined by the language recognition device itself. Specifically, the language recognition device can have a human-computer interaction interface, and then the voice data to be recognized input by the user can be obtained through the human-computer interaction interface. Of course, in other embodiments, the language recognition device has a communication module, and the language recognition device receives the voice data to be recognized sent by the terminal device through the communication module. The terminal device has a human-computer interaction function, and the user inputs the voice data to be recognized to be sent to the language recognition device on the terminal device by using the human-computer interaction function. The voice data to be recognized can be long voice data or short voice data.
[0020] Step S12: using a language recognition model to extract features from the voice data to be recognized, to obtain voice features to be recognized.
[0021] In an embodiment, the voice data to be recognized can be divided into several segments of voice data and input into the above language recognition model for feature extraction. For example, the voice data to be recognized is divided into several segments of voice data with fixed length, and input into the language recognition model for feature extraction.
[0022] In another embodiment, the voice data to be recognized can be preprocessed to reduce invalid data therein. For example, the invalid data such as silence and noise in the voice data to be recognized is filtered out.
[0023] The language recognition model in this paper can be a pre-trained model, that is, the to-be-recognized speech data can be input into the pre-trained model for feature extraction to obtain to-be-recognized speech features. In an embodiment, feature extraction can be performed using a feature extraction network in the language recognition model. For example, a time-delay neural network (TDNN), a convolutional neural network (CNN), a residual network (ResNet), and the like.
[0024] In a specific embodiment, the to-be-recognized speech data is preprocessed to remove invalid data. The preprocessed to-be-recognized speech data is divided into a plurality of segments of fixed length, and a segment is referred to as a frame. That is, for the to-be-recognized speech data corresponding to each frame, a time-delay neural network in the language recognition model is used to obtain filter bank features corresponding to each frame of to-be-recognized speech data
[0025] (Filter Bank), and the filter bank features are used as to-be-recognized language features. It can be understood that in the case where the to-be-recognized speech data is divided into a plurality of segments of speech data, a plurality of to-be-recognized speech features are obtained, and each segment of to-be-recognized speech data corresponds to one to-be-recognized speech feature.
[0026] Step S13: using the language recognition model to recognize based on the to-be-recognized speech features to obtain a first language recognition result, and the first language recognition result is used to represent the language to which the to-be-recognized speech data belongs.
[0027] The language recognition model can recognize a plurality of first languages, the plurality of first languages include at least one target language, the language recognition model is obtained by at least adjusting parameters based on a constraint loss, the constraint loss is determined based on first sample features extracted by the language recognition model from a first training set, the first training set includes sample speech data of at least two second languages, and the at least two second languages include the target language, and the target of adjusting the parameters based on the constraint loss is to increase the distance between the speech features of different second languages and to reduce the distance between the speech features of the same second language.
[0028] The target language refers to a minority language, that is, a language with a relatively small amount of data in the training data set used to build the language recognition model, or a language that is relatively less used by the user. In an embodiment, before training the language recognition model, sample speech data corresponding to each first language is pre-stored as a training data set. In the training data set, a language with less sample speech data is determined as the target language. For example, in the training data set, the duration of the sample speech data corresponding to each first language is at least one hour. Among them, the sample speech data corresponding to Chinese is 4 hours, and the sample speech data corresponding to Czech is 1 hour. Czech can be used as the target language.
[0029] The second language refers to each target language, or the target language and the remaining other languages except the target language. That is, the constraint loss is used to adjust the parameters to increase the speech feature distance between different target languages, and to reduce the speech feature distance within the same target language. Alternatively, the constraint loss is used to adjust the parameters to increase the speech feature distance between the target language and other languages, to reduce the speech feature distance within the target language, and to reduce the speech feature distance within the other languages. For example, the target languages are Czech and Malay, and the constraint loss is used to adjust the parameters to increase the speech feature distance between Czech and Malay, to reduce the speech feature distance within Czech, and to reduce the speech feature distance within Malay. For another example, the first languages are Czech, Chinese, and English. The target language is Czech, and the other languages are Chinese and English. The constraint loss is used to adjust the parameters to increase the speech feature distance between Czech and the other languages, to reduce the speech feature distance within Czech, and to reduce the speech feature distance within the other languages.
[0030] The target training set can be repeatedly obtained from the training data set, and the language recognition model can be trained using the target training set, or the language recognition model can be directly trained using the training data set. In an embodiment, the target training set is the first training set and the second training set, and the first training set and the second training set are obtained, wherein the second training set includes sample speech data corresponding to at least one first language. The feature extraction network in the language recognition model is used to extract features of each sample speech data in the first training set, to obtain first sample features of each sample speech data in the first training set, and to determine the constraint loss of the language recognition model based on the first sample features. The feature extraction network in the language recognition model is also used to extract features of each sample speech data in the second training set, to obtain second sample features of each sample speech data in the second training set, and to identify the predicted language corresponding to each sample speech data in the second training set based on the second sample features. The prediction loss of the language recognition model is determined based on the difference between the predicted language and the labeled language corresponding to each sample speech data in the second training set. The network parameters of the language recognition model are adjusted based on the constraint loss and the prediction loss of the language recognition model. The above steps are repeated to iteratively train the language recognition model. The feature extraction network in the language recognition model can be used to extract features of the sample speech data in the first training set and the second training set.
[0031] The sample voice data of each second language can be divided into first sample data and second sample data for calculating the constraint loss. In an embodiment, the sample voice data of each second language is divided into first sample voice data and second sample voice data, the first sample feature of the first sample voice data of the second language is used to calculate the language sample feature of the second language, the constraint loss is positively correlated with the first feature distance of each second sample voice data and negatively correlated with the second feature distance of the second sample voice data, the first feature distance represents the distance between the first sample feature of the second sample voice data and the language sample feature corresponding to the second language, and the second feature distance represents the distance between the first sample feature of the second sample voice data and the language sample feature corresponding to the non-second language. Wherein, the first sample feature and the second sample feature can be sample feature vectors.
[0032] In a specific embodiment, the constraint loss is the sum of the total feature distances of each second sample voice data, and the total feature distance of the second sample voice data is equal to the first feature distance of the second sample voice data minus the second feature distance between the second sample voice data and each non-second language.
[0033] In another specific embodiment, the language sample feature of the second language is the central tendency value of the first sample feature of each first sample voice data of the second language. For example, the average value of the first sample feature of each first sample voice data of the second language is used as the central tendency value. For another example, the standard deviation of the first sample feature of each first sample voice data of the second language is used as the central tendency value. Of course, other values that can represent the trend of each first sample feature can also be used as the central tendency value, which is not limited here.
[0034] The language recognition model can recognize multiple or one first language.
[0035] In the case of multiple first languages, i.e., the language recognition model can recognize multiple first languages. The first training set and the second training set can be obtained from the pre-stored training data set. In an embodiment, from the pre-stored sample voice data corresponding to the first number of target languages, the sample voice data of the second number of target languages is randomly extracted as the sample voice data corresponding to each second language in the first training set, and the number of each second language extracted is the third number. And from the pre-stored sample voice data corresponding to each first language, the sample voice data of the fourth number of first languages is randomly extracted to obtain the second training set, wherein the sample voice data of each first language extracted is the fifth number. In a specific embodiment, the total amount of sample voice data in the first training set obtained by each iteration training is fixed and consistent.
[0036] Of course, in the case of a first language being one, i.e., the language recognition model can recognize one first language. In an embodiment, the first training set and the second training set can be randomly extracted from pre-stored sample voice data corresponding to the first language and other languages. In a specific embodiment, the total amount of sample voice data corresponding to the first language and other languages obtained in the first training set or the second training set is consistent.
[0037] In the case of the first language being multiple, taking the first training set as an example, second quantity N types of sample voice data corresponding to target languages are randomly extracted from the first quantity M of sample voice data corresponding to each target language. Third quantity L of sample voice data is extracted from the second quantity N types of target languages, wherein N*L=D, and the size of D remains unchanged. The above sample voice data is used as the first training set.
[0038] The constraint loss and the prediction loss of the language recognition model can be used to adjust the network parameters of the language recognition model. In a specific embodiment, the weight of the constraint loss and the prediction loss of the language recognition model obtained in this iteration is obtained respectively, wherein the sum of the weight corresponding to the constraint loss and the weight corresponding to the prediction loss is a fixed value. In the case of the first number of iterations, the weight corresponding to the prediction loss is a preset weight value. In the case of the number of iterations being greater than the first number and less than the maximum number of iterations, the weight corresponding to the prediction loss is the difference between the preset weight value and the number ratio, and the number ratio is the ratio between the number of iterations and the maximum number of iterations. The constraint loss and the prediction loss of the language recognition model are weighted and summed using the weight to obtain the total loss of the language recognition model. The network parameters of the language recognition model are adjusted using the total loss of the language recognition model.
[0039] In the case of the value of the total loss tending to be stable or the number of iterations reaching the maximum number of iterations, the adjustment of the network parameters of the language recognition model is completed.
[0040] The language recognition model can be used to recognize the corresponding first language recognition result based on the to-be-recognized voice feature. In the case that the corresponding first language recognition result corresponds to a target language, the first language recognition result can be confirmed again.
[0041] In an embodiment, the first language is multiple, the language recognition model corresponds to a multi-language recognition model, and each second language is a target language. That is, the multi-language recognition model can recognize multiple first languages. The multi-language recognition model is used to recognize the to-be-recognized voice feature to obtain the corresponding first language recognition result. For specific description of the multi-language recognition model, reference can be made to the related description in the language recognition model in step S13.
[0042] In a case where the target language exists in the language corresponding to the first language recognition result, the target language corresponding language recognition model can be used to recognize the to-be-recognized speech feature, so as to perform secondary confirmation. In an embodiment, the first language is one, and the target language is one. The language recognition model corresponds to a binary classification model corresponding to the target language. That is, the binary classification model corresponding to the target language can recognize a first language (target language) and a second language, which is the target language and other languages. In this case, the first language other than the target language is regarded as the other language. The target of the constraint loss parameter adjustment is to increase the speech feature distance between the target language and the other language, to reduce the speech feature distance within the target language, and to reduce the speech feature distance within the other language.
[0043] In a specific embodiment, each target language is regarded as a to-be-trained language respectively. A third training set is obtained, wherein the third training set includes sample speech data of at least two first languages, and the at least two first languages include the to-be-trained language. In the third training set, the first language other than the to-be-trained language is regarded as the other language. The feature extraction network of the language recognition model is used to extract features of each sample speech data in the third training set, to obtain third sample features of each sample speech data in the third training set. The binary classification model corresponding to the to-be-trained language is used to recognize each third sample feature respectively, to obtain a language prediction result corresponding to each sample speech data in the third training set. The language prediction result represents whether the sample speech data belongs to the to-be-trained language or the other language. The constraint loss of the binary classification model is determined by using each third sample feature. The prediction loss of the binary classification model is determined based on the language prediction result corresponding to each sample speech data in the third training set and the labeled language result. The network parameters of the binary classification model are adjusted by using the constraint loss and the prediction loss of the binary classification model. The foregoing step of obtaining the third training set and the subsequent steps are repeated to iteratively train the binary classification model corresponding to the to-be-trained language. The third sample feature can be a sample feature vector.
[0044] The sample number of the sample speech data selected in the third training set corresponding to the to-be-trained language and the other language can be consistent. For example, the first language is Chinese, English, and Czech. The Czech language is the to-be-trained language. N sample speech data are randomly selected from the sample speech data corresponding to the Czech language, and a total of N sample speech data are randomly selected from the sample speech data corresponding to Chinese and English. The sample speech data selected above are regarded as the corresponding third training set.
[0045] In yet another embodiment, the first language is multiple, i.e., the multi-language recognition model is used for language recognition. In response to the first language recognition result indicating that the to-be-recognized speech data belongs to a target language, a binary classification model corresponding to the target language is used to perform secondary recognition on the to-be-recognized speech feature, to obtain a second language recognition result indicating whether the to-be-recognized speech data belongs to the recognized target language. In response to the first language recognition result indicating that the to-be-recognized speech data does not belong to the recognized target language, the to-be-recognized speech data is determined to be the language indicated by the first language recognition result.
[0046] The to-be-recognized speech data is segmented and processed, and there can be a plurality of corresponding to-be-recognized speech features, i.e., a plurality of corresponding first language recognition results. In the above plurality of corresponding first language recognition results, there can be a target language, and the first language recognition result corresponding to the target language can be confirmed again.
[0047] In an embodiment, the first language recognition result includes a confidence level of the to-be-recognized speech data belonging to each first language. After obtaining the second language recognition result, in response to the second language recognition result indicating that the to-be-recognized speech data belongs to the recognized target language, the to-be-recognized speech data is determined to be the recognized target language. In response to the second language recognition result indicating that the to-be-recognized speech data does not belong to the recognized target language, the confidence levels of each remaining first language except the recognized target language are obtained from the first language recognition result, and the language to which the to-be-recognized speech data belongs is determined based on the confidence levels of each remaining first language.
[0048] In a specific embodiment, an average value of the confidence levels of each first language can be obtained, and the corresponding first language with the maximum average value is selected as the language to which the to-be-recognized speech data belongs. In another specific embodiment, the corresponding first language with the maximum confidence level is selected as the language to which the to-be-recognized speech data belongs.
[0049] For example, Czech is the target language. After the to-be-recognized speech data is segmented and processed for feature extraction, a plurality of to-be-recognized speech features are obtained. Each to-be-recognized speech data is recognized by using the multi-language recognition model, and each first language recognition result is Chinese 90 points, English 50 points, Chinese 80 points, and Czech 50 points. The to-be-recognized speech feature corresponding to the recognized Czech is input into the binary classification model corresponding to Czech for recognition. In the case that the second language recognition result is other language, the language to which the to-be-recognized speech data belongs is determined from the first language recognition results Chinese 90 points, English 50 points, and Chinese 80 points. Among them, the average value of the confidence levels corresponding to Chinese is 85 points, and the average value of the confidence levels corresponding to English is 50 points, so Chinese is selected as the language to which the to-be-recognized speech data belongs.
[0050] Please refer toFigure 2 , Figure 2 is a flowchart of an embodiment of a method for training a language identification model provided by the present application. The language identification model can identify a plurality of first languages, and the plurality of first languages includes at least one target language. It should be noted that the embodiment is not limited to the order of the flowchart shown in Figure 2 . As shown in Figure 2 , the embodiment includes the following steps:
[0051] Step S21: Obtain a first training set and a second training set.
[0052] The first training set includes sample voice data of at least two second languages, and the at least two second languages include the target language. The second training set includes sample voice data corresponding to at least one first language.
[0053] The first training set and the second training set can be obtained from sample voice data in a pre-stored training data set. In the pre-stored training data set, there is sample voice data corresponding to the target language.
[0054] The sample voice data can be pre-processed before the pre-stored training data set corresponding sample voice data, to reduce invalid data. The first training set and the second training set are obtained from the pre-processed training data set corresponding sample voice data.
[0055] The target language is determined based on the data amount corresponding to the sample voice data corresponding to the training data set. The determination method of the target language can refer to the related description of step S13.
[0056] The sample voice data corresponding to the target language can be processed at a variable speed after being pre-processed, or it can be processed at a variable speed before being pre-processed, to obtain more sample voice data, which is not limited here. For example, the sample voice data corresponding to the target language is processed at 0.9 times and 1.1 times, and the sample voice data corresponding to 0.9 times, the sample voice data corresponding to 1.1 times, and the sample voice data corresponding to 1 times are collectively used as sample voice data corresponding to the target language. The sample voice data corresponding to the target language after variable speed processing is pre-processed.
[0057] The obtaining method of the first training set and the second training set can refer to the related description in step S13, which is not repeated here.
[0058] In the case of multiple first languages, taking the acquisition of the first training set as an example, in the case of multiple first languages, the language recognition model corresponds to a multi-language recognition model, and the second language is the target language, each first language is Chinese, English, Japanese, Czech, Malay, and Arabic. The target language is Czech and Malay. Randomly select 2 from Czech, Malay and Arabic, such as Malay and Arabic. Then extract 6 sample speech data from Malay and Arabic respectively as the first training set. Alternatively, randomly select 3 from Czech, Malay and Arabic, such as Malay, Czech and Arabic. Then extract 4 sample speech data from Czech, Malay and Arabic respectively as the first training set. Among them, the sample speech data extracted each time remains at 12.
[0059] Step S22: using the language recognition model to perform feature extraction on each sample speech data in the first training set to obtain first sample features of each sample speech data in the first training set, and using each first sample feature to determine the constraint loss of the language recognition model.
[0060] The feature extraction network of the language recognition model can be used to perform feature extraction on each sample speech data in the first training set to obtain the corresponding first sample features of each sample speech data.
[0061] The specific description of using the first sample features to determine the constraint loss of the language recognition model can refer to the related description in step S13.
[0062] Taking the constraint loss as an example, in the case that the first language is multiple, the language recognition model corresponds to a multi-language recognition model, the second language is the target language, and 6 sample voice data are extracted from Malay and Arabic as the first training set. The 6 sample voice data corresponding to the Malay language are divided into first sample voice data (3 sample voice data) and second sample voice data (another 3 sample voice data). The 6 sample data corresponding to the Arabic language are also divided into first sample voice data and second sample voice data in the above manner. The first sample features corresponding to the first sample data of the Malay language are averaged to obtain language sample features. In the same way, the language sample features corresponding to the Arabic language are obtained. The cosine distance between each second sample feature corresponding to the Malay language and the language sample feature corresponding to the Malay language is taken as the first feature distance. The cosine distance between each second sample feature corresponding to the Malay language and the language sample feature corresponding to the Arabic language is taken as the second feature distance. In the same way, the first feature distance and the second feature distance corresponding to the Arabic language are obtained. The sum of the first feature distance corresponding to the Malay language and the first feature distance corresponding to the Arabic language is subtracted from the sum of the second feature distance corresponding to the Malay language and the second feature distance corresponding to the Arabic language. The final value calculated is taken as the total feature sum, that is, the constraint loss.
[0063] Step S23: Feature extraction is performed on each sample voice data in the second training set using the language recognition model to obtain second sample features of each sample voice data in the second training set, and recognition is performed based on each second sample feature to obtain the predicted language corresponding to each sample voice data in the second training set. Based on the difference between the predicted language and the labeled language corresponding to each sample voice data in the second training set, the prediction loss of the language recognition model is determined.
[0064] The specific description of determining the prediction loss of the language recognition model using the second sample features can be referred to the related description in step S13.
[0065] Step S24: The network parameters of the language recognition model are adjusted using the constraint loss and the prediction loss of the language recognition model.
[0066] The target of adjusting the parameters using the constraint loss is to increase the voice feature distance between different second languages and to reduce the voice feature distance within the same second language.
[0067] The specific description of adjusting the network parameters of the language recognition model using the constraint loss and the prediction loss of the language recognition model can be referred to the related description in step S13.
[0068] The total loss can be obtained using constraint loss and prediction loss. Taking the total loss Loss_all as an example, the weight of prediction loss LimitLoss is θ, and the weight of constraint loss CELoss is 1-θ. The expression for the total loss is Loss_all = θ * CELoss + (1-θ) * LimitLoss. Where θ is a piecewise function, specifically: In the above formula, iter is the current iteration number, and iter_max is the maximum iteration number set.
[0069] Taking network parameter adjustment as an example, the network parameters of the language recognition model are adjusted when the total loss Loss_all value tends to stabilize, or when the current iteration number iter reaches the maximum iteration number iter_max.
[0070] The piecewise function can be modified according to the actual situation, that is, the first number can be changed. For example, the first number can be changed from 5 to 2, meaning that if iter > 2, then the process begins. Attenuation is performed, and no restrictions are imposed here.
[0071] Before training the language recognition model, an initial learning rate can be preset, and then the network parameters of the model can be dynamically adjusted and optimized by combining the model's total loss value.
[0072] See Figure 3 , Figure 3 This is a schematic diagram of the framework of an embodiment of the language recognition device of this application. The language recognition device 300 includes a speech data acquisition module 310, a speech feature acquisition module 320, and a language recognition module 330. The speech data acquisition module 310 is used to acquire speech data to be recognized. The speech feature acquisition module 320 is used to extract features from the speech data to be recognized using a language recognition model to obtain speech features to be recognized. The language recognition module 330 is used to recognize based on the speech features to be recognized using the language recognition model to obtain a first language recognition result. The first language recognition result is used to characterize the language to which the speech data to be recognized belongs. The language recognition model can recognize several first languages, including at least one target language. The language recognition model is obtained by at least using constraint loss parameter tuning. The constraint loss is determined based on the first sample features extracted by the language recognition model from a first training set. The first training set includes sample speech data of at least two second languages, and the at least two second languages include the target language. The goal of using constraint loss parameter tuning is to increase the speech feature distance between different second languages and to reduce the speech feature distance within the same second language.
[0073] In some embodiments, the language recognition device 300 further comprises a training module, before the speech feature acquisition module 320 performs feature extraction on the to-be-recognized speech data by using the language recognition model to obtain the to-be-recognized speech features, the training module performs obtaining a first training set and a second training set, wherein the second training set comprises sample speech data corresponding to at least one first language. The language recognition model is used to perform feature extraction on each sample speech data in the first training set to obtain first sample features of each sample speech data in the first training set, and the constraint loss of the language recognition model is determined by using each first sample feature. And the language recognition model is used to perform feature extraction on each sample speech data in the second training set to obtain second sample features of each sample speech data in the second training set, and the predicted language corresponding to each sample speech data in the second training set is obtained by identifying based on each second sample feature respectively. Based on the difference between the predicted language and the labeled language corresponding to each sample speech data in the second training set, the prediction loss of the language recognition model is determined. The network parameters of the language recognition model are adjusted by using the constraint loss and the prediction loss of the language recognition model. The foregoing steps are repeated to iteratively train the language recognition model.
[0074] In some embodiments, the sample speech data of each second language is divided into first sample speech data and second sample speech data, the first sample features of the first sample speech data of the second language are used to calculate the language sample features of the second language, the constraint loss is positively correlated with the first feature distance of each second sample speech data and negatively correlated with the second feature distance of the second sample speech data, the first feature distance represents the distance between the first sample feature of the second sample speech data and the language sample feature corresponding to the second language to which the second sample speech data belongs, and the second feature distance represents the distance between the first sample feature of the second sample speech data and the language sample feature corresponding to the second language to which the second sample speech data does not belong.
[0075] In some embodiments, the constraint loss is the sum of the total feature distances of each second sample speech data, and the total feature distance of the second sample speech data is equal to the first feature distance of the second sample speech data minus the second feature distance between the second sample speech data and each second language to which the second sample speech data does not belong.
[0076] In some embodiments, the language sample feature of the second language is a central tendency value of the first sample features of each first sample speech data of the second language.
[0077] In some embodiments, the training module performs the obtaining the first training set and the second training set by: randomly extracting, from pre-stored sample voice data corresponding to each target language of the first quantity, sample voice data of a second quantity of target languages as sample voice data corresponding to each second language in the first training set, the quantity of each second language being a third quantity; and randomly extracting, from pre-stored sample voice data corresponding to each first language, sample voice data of a fourth quantity of first languages to obtain the second training set, wherein the sample voice data of each first language extracted is a fifth quantity.
[0078] In some embodiments, the training module performs the adjusting the network parameters of the language identification model using the constraint loss and the prediction loss of the language identification model by: respectively obtaining the weight of the constraint loss and the prediction loss of the language identification model in the current iteration, wherein the sum of the weight corresponding to the constraint loss and the weight corresponding to the prediction loss is a fixed value, in the case of the current iteration being less than a first number of times, the weight corresponding to the prediction loss is a preset weight value, in the case of the current iteration being greater than the first number of times and less than a maximum number of iterations, the weight corresponding to the prediction loss is the difference between the preset weight value and the number ratio, the number ratio being the ratio between the current number of iterations and the maximum number of iterations. The constraint loss and the prediction loss of the language identification model are weighted and summed using the weights to obtain the total loss of the language identification model. The network parameters of the language identification model are adjusted using the total loss of the language identification model.
[0079] In some embodiments, the first language is multiple, and the language identification model corresponds to a multi-language identification model, and each second language is the target language.
[0080] In some embodiments, the first language is one, and is a target language, and the language identification model corresponds to a binary classification model corresponding to the target language.
[0081] In some embodiments, the first language is multiple, and after the language identification module 330 performs the identification based on the to-be-identified voice feature using the language identification model to obtain the first language identification result, the following is performed: in response to the first language identification result indicating that the to-be-identified voice data belongs to the target language, performing secondary identification on the to-be-identified voice feature using a binary classification model corresponding to the target language to obtain a second language identification result, the second language identification result indicating whether the to-be-identified voice data belongs to the identified target language; and in response to the first language identification result indicating that the to-be-identified voice data does not belong to the identified target language, determining that the to-be-identified voice data is the language indicated by the first language identification result.
[0082] In some embodiments, the first language recognition result comprises a confidence level of the to-be-recognized speech data belonging to each of the first languages. After the language recognition module 330 obtains the second language recognition result, the step of determining that the to-be-recognized speech data belongs to the recognized target language is performed in response to the second language recognition result indicating that the to-be-recognized speech data belongs to the recognized target language. In response to the second language recognition result indicating that the to-be-recognized speech data does not belong to the recognized target language, the confidence levels of each of the remaining first languages except the recognized target language are obtained from the first language recognition result, and the language to which the to-be-recognized speech data belongs is determined based on the confidence levels of each of the remaining first languages.
[0083] In some embodiments, the language recognition apparatus 300 further comprises a training module configured to take each target language as a to-be-trained language respectively. A third training set is obtained, wherein the third training set comprises sample speech data of at least two first languages, and the at least two first languages comprise the to-be-trained language, and the first languages other than the to-be-trained language in the third training set are taken as other languages. The feature extraction network of the language recognition model is used to perform feature extraction on each sample speech data in the third training set to obtain third sample features of each sample speech data in the third training set. The binary classification model corresponding to the to-be-trained language is used to perform recognition based on each third sample feature respectively to obtain a language prediction result corresponding to each sample speech data in the third training set, which indicates whether the sample speech data belongs to the to-be-trained language or other languages. The constraint loss of the binary classification model is determined using each third sample feature. The prediction loss of the binary classification model is determined based on the language prediction result corresponding to each sample speech data in the third training set and the labeled language result. The network parameters of the binary classification model are adjusted using the constraint loss and the prediction loss of the binary classification model. The foregoing step of obtaining the third training set and the subsequent steps are repeated to iteratively train the binary classification model corresponding to the to-be-trained language.
[0084] Please refer to Figure 4 , Figure 4 is a schematic diagram of an embodiment of a language recognition device. The language recognition device 40 comprises a memory 41 and a processor 42 coupled to each other. The processor 42 is configured to execute program instructions stored in the memory 41 to implement the steps in any of the above language recognition method embodiments. In a specific implementation scenario, the language recognition device 40 can include but is not limited to a microcomputer, a server, and in addition, the language recognition device 40 can also include a notebook computer, a tablet computer, and other mobile devices, which are not limited here.
[0085] Specifically, the processor 42 is configured to control itself and the memory 41 to implement the steps in any of the above language recognition method embodiments. The processor 42 can also be referred to as a CPU (Central Processing Unit). The processor 42 can be an integrated circuit chip having a processing capability of signals. The processor 42 can also be a general processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The general processor can be a microprocessor or the processor can also be any conventional processor or the like. In addition, the processor 42 can be jointly implemented by integrated circuit chips.
[0086] Please refer to Figure 5 , Figure 5 is a schematic diagram of a framework of an embodiment of the computer readable storage medium of the present application. The computer readable storage medium 50 stores program instructions 51 capable of being executed by the processor, and the program instructions 51 are configured to implement the steps in any of the above language recognition method embodiments.
[0087] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, details are not described here.
[0088] The above description of various embodiments tends to emphasize the differences between various embodiments, and the same or similar parts can be mutually referred to. For brevity, details are not described here.
[0089] In several embodiments provided in the present application, it should be understood that the disclosed method and device can be implemented in other ways. For example, the above-described device implementation is only schematic; for example, the division of the modules or units is only a logical function division, and there can be another division manner in actual implementation; for example, a unit or component can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0090] In addition, each of the functional units in the various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0091] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application, essentially or in part, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, and various other media that can store program codes.
Claims
1. A language identification method characterized by, The method comprises: acquiring voice data to be recognized; extracting features from the voice data to be recognized by using a language recognition model to obtain voice features to be recognized; recognizing based on the voice features to be recognized by using the language recognition model to obtain a first language recognition result, the first language recognition result being used to represent a language to which the voice data to be recognized belongs; wherein the language recognition model can recognize a plurality of first languages, the plurality of first languages including at least one target language, the language recognition model being obtained by adjusting parameters based on a constraint loss, the constraint loss being determined based on first sample features extracted from a first training set by the language recognition model, the first training set including sample voice data of at least two second languages, and the at least two second languages including the target language, the target of adjusting the parameters based on the constraint loss being to increase voice feature distances between different second languages and to reduce voice feature distances within the same second language, sample voice data of each second language being divided into first sample voice data and second sample voice data, first sample features of the first sample voice data of the second language being used to calculate language sample features of the second language, the constraint loss being positively correlated with first feature distances of the second sample voice data and being negatively correlated with second feature distances of the second sample voice data, the first feature distance representing a distance between the first sample features of the second sample voice data and language sample features corresponding to the second language to which the second sample voice data belongs, the second feature distance representing a distance between the first sample features of the second sample voice data and language sample features corresponding to a second language other than the second language to which the second sample voice data belongs, the constraint loss being a sum of total feature distances of the second sample voice data, the total feature distance of the second sample voice data being equal to the first feature distance minus the second feature distance.
2. The method of claim 1, wherein, Before the step of extracting features from the voice data to be recognized by using the language recognition model to obtain voice features to be recognized, the method further comprises: acquiring the first training set and a second training set, wherein the second training set includes sample voice data corresponding to at least one first language; extracting features from each sample voice data in the first training set by using the language recognition model to obtain first sample features of each sample voice data in the first training set, and determining a constraint loss of the language recognition model based on each first sample feature; and extracting features from each sample voice data in the second training set by using the language recognition model to obtain second sample features of each sample voice data in the second training set, and recognizing based on each second sample feature to obtain predicted languages corresponding to each sample voice data in the second training set; determining a prediction loss of the language recognition model based on differences between the predicted languages and labeled languages corresponding to each sample voice data in the second training set; adjusting network parameters of the language recognition model based on the constraint loss and the prediction loss of the language recognition model. The foregoing steps are repeated to iteratively train the language recognition model.
3. The method of claim 1, wherein, The language sample feature of the second language is a central tendency value of first sample features of each first sample voice data of the second language.
4. The method of claim 2, wherein, The obtaining the first training set and the second training set comprises: randomly extracting a second number of sample voice data of the target language from a pre-stored first number of sample voice data corresponding to each of the target language, as the first training set of sample voice data corresponding to each of the second language, and the number of each of the second language extracted is a third number; and randomly extracting a fourth number of sample voice data of the first language from the pre-stored sample voice data corresponding to each of the first language to obtain the second training set, wherein the sample voice data extracted from each of the first language is a fifth number; And / or, the adjusting the network parameters of the language recognition model using the constraint loss and the prediction loss of the language recognition model comprises: respectively obtaining the weight of the constraint loss and the prediction loss of the language recognition model in this iteration, wherein the sum of the weight corresponding to the constraint loss and the weight corresponding to the prediction loss is a fixed value, in the case of the first number of iterations, the weight corresponding to the prediction loss is a preset weight value, in the case of more than the first number of iterations and less than the maximum number of iterations, the weight corresponding to the prediction loss is the difference between the preset weight value and the number ratio, and the number ratio is the ratio between the number of iterations and the maximum number of iterations; performing weighted summation on the constraint loss and the prediction loss of the language recognition model using the weight to obtain the total loss of the language recognition model; adjusting the network parameters of the language recognition model using the total loss of the language recognition model.
5. The method of claim 1, wherein, The first language is multiple, the language recognition model corresponds to a multi-language recognition model, and each of the second language is the target language; or the first language is one, and is the target language, and the language recognition model corresponds to a two-classification model corresponding to the target language.
6. The method of claim 5, wherein, The first language is multiple; After the language recognition model is used to recognize based on the to-be-recognized voice feature to obtain a first language recognition result, the method further comprises: in response to the first language recognition result representing that the to-be-recognized voice data belongs to the target language, using a two-classification model corresponding to the target language to perform secondary recognition on the to-be-recognized voice feature to obtain a second language recognition result, the second language recognition result representing whether the to-be-recognized voice data belongs to the recognized target language; in response to the first language recognition result representing that the to-be-recognized voice data does not belong to the recognized target language, determining that the to-be-recognized voice data is the language represented by the first language recognition result.
7. The method of claim 6, wherein, The first language recognition result comprises a confidence degree of the to-be-recognized voice data belonging to each of the first language; After obtaining the second language recognition result, the method further comprises: In response to the second language recognition result indicating that the to-be-recognized speech data belongs to the recognized target language, it is determined that the to-be-recognized speech data is of the recognized target language; In response to the second language recognition result indicating that the to-be-recognized speech data does not belong to the recognized target language, the confidence degrees of the remaining first languages except the recognized target language are obtained from the first language recognition result, and the language to which the to-be-recognized speech data belongs is determined based on the confidence degrees of the remaining first languages.
8. The method of claim 6, wherein, The method further comprises: Each of the target languages is taken as a to-be-trained language respectively; A third training set is obtained, wherein the third training set comprises sample speech data of at least two first languages, and the at least two first languages comprise the to-be-trained language, and each first language in the third training set except the to-be-trained language is taken as another language; Feature extraction is performed on each sample speech data in the third training set by using a feature extraction network of the language recognition model, to obtain third sample features of each sample speech data in the third training set; Each third sample feature is used to perform recognition based on the corresponding binary classification model of the to-be-trained language, to obtain a language prediction result corresponding to each sample speech data in the third training set, which indicates whether the sample speech data belongs to the to-be-trained language or the another language; A constraint loss of the binary classification model is determined by using each third sample feature, and a prediction loss of the binary classification model is determined based on the language prediction result corresponding to each sample speech data in the third training set and a labeled language result; The network parameters of the binary classification model are adjusted by using the constraint loss and the prediction loss of the binary classification model; The foregoing step of obtaining the third training set and the subsequent steps are repeated to iteratively train the binary classification model corresponding to the to-be-trained language. 9.A method for training a language identification model, the method comprising: The language recognition model can recognize a plurality of first languages, and the plurality of first languages comprise at least one target language; the method comprises: A first training set and a second training set are obtained, wherein the first training set comprises sample speech data of at least two second languages, and the at least two second languages comprise the target language, and the second training set comprises sample speech data corresponding to at least one first language; Feature extraction is performed on each sample speech data in the first training set by using the language recognition model, to obtain first sample features of each sample speech data in the first training set, and a constraint loss of the language recognition model is determined by using each first sample feature; and Feature extraction is performed on each sample speech data in the second training set by using the language recognition model, to obtain second sample features of each sample speech data in the second training set, and each second sample feature is used to perform recognition, to obtain a predicted language corresponding to each sample speech data in the second training set; a prediction loss of the language recognition model is determined based on a difference between the predicted language corresponding to each sample speech data in the second training set and a labeled language; and The network parameters of the language recognition model are adjusted by using the constraint loss and the prediction loss of the language recognition model. The network parameters of the language recognition model are adjusted by using constraint loss and prediction loss of the language recognition model, wherein the constraint loss is used for adjusting the parameters, and a target of the constraint loss is to increase the speech feature distance between different second languages and to reduce the speech feature distance within the same second language.
10. A language identification apparatus characterized by comprising: The device comprises: The voice data acquisition module is configured to acquire voice data to be recognized. The voice feature acquisition module is configured to extract features from the voice data to be recognized by using a language recognition model to obtain voice features to be recognized. The language recognition module is configured to recognize the voice features to be recognized by using the language recognition model to obtain a first language recognition result, wherein the first language recognition result is used to represent a language to which the voice data to be recognized belongs, the language recognition model can recognize a plurality of first languages, the plurality of first languages include at least one target language, the language recognition model is obtained by using at least constraint loss to adjust parameters, the constraint loss is determined based on first sample features extracted from a first training set by the language recognition model, the first training set includes sample voice data of at least two second languages, and the at least two second languages include the target language, a target of the constraint loss is to increase the speech feature distance between different second languages and to reduce the speech feature distance within the same second language, sample voice data of each second language is divided into first sample voice data and second sample voice data, first sample features of the first sample voice data of the second language are used to calculate language sample features of the second language, the constraint loss is positively correlated with first feature distances of the second sample voice data and is negatively correlated with second feature distances of the second sample voice data, the first feature distance represents a distance between the first sample features of the second sample voice data and language sample features corresponding to the second language to which the second sample voice data belongs, the second feature distance represents a distance between the first sample features of the second sample voice data and language sample features corresponding to a second language other than the second language to which the second sample voice data belongs, and the constraint loss is a sum of total feature distances of the second sample voice data, and the total feature distance of the second sample voice data is equal to the first feature distance minus the second feature distance.
11. A language identification device, characterized by The language recognition device comprises a memory and a processor, the memory stores program instructions, and the processor is configured to execute the program instructions to implement the language recognition method of any one of claims 1-8.
12. A computer-readable storage medium, characterized in that, The computer readable storage medium is configured to store program instructions, and the program instructions can be executed to implement the language recognition method of any one of claims 1-8.
Citation Information
Patent Citations
Model training method, language identification method, language identification device and equipment
CN110838286A
Language recognition method, related equipment and readable storage medium
CN111724766A