Language recognition method, computer device, storage medium, and computer program product
The method improves language identification accuracy by using a pre-trained model to extract features from audio data across multiple languages, determining language based on codebook feature distributions, effectively recognizing new languages without retraining.
Patent Information
- Application Number
- CN202211190072.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-28
AI Technical Summary
The existing language recognition technology model needs to be retrained when facing new languages, resulting in low recognition accuracy.
The model is extracted through pre-trained audio feature, and unsupervised learning is used to use sample audio from different languages, audio features are extracted and language categories are determined according to the distribution of feature vectors, avoiding model retraining.
It improves the recognition accuracy of new languages without retraining the model, and enhances the recognition ability of new language categories.
Smart Images

Figure CN115762474B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and particularly to a language identification method, a computer device, a storage medium, and a computer program product. Background Art
[0002] With the rapid development of computer technologies and the increasingly close international exchanges, audio data in multiple languages is used in multiple fields, and language identification of audio data has also become an important technology in various fields.
[0003] Currently, the models trained by existing language identification technologies can usually only identify fixed types of languages, which is related to the types of language datasets used in the model training process. When facing new languages, the model needs to be retrained, resulting in a relatively low recognition accuracy of the trained model for new languages. Summary of the Invention
[0004] Based on this, it is necessary to provide a language identification method, a computer device, a computer-readable storage medium, and a computer program product that can improve the language identification accuracy for the above technical problems.
[0005] In a first aspect, the present application provides a language identification method. The method includes:
[0006] Inputting the audio to be identified into a pre-trained audio feature extraction model to obtain the audio features of the audio to be identified; the pre-trained audio feature extraction model is trained by sample audios in different languages;
[0007] Obtaining, from the audio codebook included in the pre-trained audio feature extraction model, a target codebook feature corresponding to the audio features of the audio to be identified; the audio codebook includes codebook feature vectors in different languages;
[0008] Obtaining a distribution feature vector of the audio to be identified according to the distribution of each codebook feature vector in the target codebook feature;
[0009] Determining the language category corresponding to the target distribution feature vector whose distance from the distribution feature vector of the audio to be identified in the distribution feature vectors of the sample audios meets a preset distance condition as the language category of the audio to be identified.
[0010] In one embodiment, the pre-trained audio feature extraction model is trained in the following manner:
[0011] Inputting sample audios in different languages into an audio encoding model and a speaker encoding model in the audio feature extraction model to be trained respectively to obtain the sample audio features and speaker features corresponding to the sample audios;
[0012] Obtain a sample codebook feature corresponding to the sample audio feature from the audio codebook included in the audio feature extraction model to be trained;
[0013] Concatenate the sample codebook feature and the speaker feature, and input the concatenated feature obtained by concatenation into the audio decoding model in the audio feature extraction model to be trained to obtain the predicted audio feature of the sample audio;
[0014] Iteratively train the audio feature extraction model to be trained according to the difference between the predicted audio feature of the sample audio and the actual audio feature of the sample audio to obtain the pre-trained audio feature extraction model.
[0015] In one embodiment, the sample audio feature is obtained through the following steps:
[0016] Perform convolution processing, batch normalization processing, and activation processing on the initial sample of the sample audio in sequence through the audio encoding model in the audio feature extraction model to be trained to obtain the processed sample feature of the sample audio;
[0017] Perform convolution processing, batch normalization processing, and activation processing on the fused feature obtained by fusing the processed sample feature of the sample audio and the initial sample feature of the sample audio in sequence to obtain the encoded feature of the sample audio;
[0018] Perform dimensionality reduction processing on the encoded feature of the sample audio to obtain the dimensionality-reduced feature of the sample audio;
[0019] Input the dimensionality-reduced feature of the sample audio into a gated recurrent network to obtain the sample audio feature of the sample audio.
[0020] In one embodiment, the speaker feature is obtained through the following steps:
[0021] Perform convolution processing, batch normalization processing, and activation processing on the initial sample feature of the sample audio in sequence through the speaker encoding model in the audio feature extraction model to be trained to obtain the processed sample feature of the sample audio;
[0022] Perform convolution processing, batch normalization processing, and activation processing on the fused feature obtained by fusing the processed sample feature of the sample audio and the initial sample feature of each frame of the sample in sequence to obtain the encoded feature of the sample audio;
[0023] Perform dimensionality reduction processing on the encoded feature of the sample audio to obtain the dimensionality-reduced feature of the sample audio;
[0024] Perform mean processing on the dimensionality-reduced features of the sample audio to obtain the speaker features corresponding to the sample audio.
[0025] In one embodiment, before respectively inputting sample audios in different languages into an audio encoding model and a speaker encoding model of a to-be-trained audio feature extraction model to obtain the sample audio features and speaker features corresponding to the sample audios, it further includes:
[0026] Obtain initial audios in different languages;
[0027] Perform voice activation processing on the initial audios in each language to obtain the valid audios in the initial audios of each language;
[0028] According to the duration of the valid audio in each language, perform speed change processing and / or pitch change processing on the valid audio in each language respectively to obtain the sample audio.
[0029] In one embodiment, obtaining the distribution feature vector of the to-be-recognized audio according to the distribution of each codebook feature vector in the target codebook feature includes:
[0030] Obtain the histogram of the target codebook feature according to the quantity distribution of the codebook feature vectors in the target codebook feature;
[0031] Perform normalization processing on the histogram of the target codebook feature to obtain the distribution feature vector of the to-be-recognized audio.
[0032] In one embodiment, determining the language category of the to-be-recognized audio as the language category corresponding to the target distribution feature vector whose distance from the distribution feature vector of the sample audio satisfies a preset distance condition includes:
[0033] Screen out a preset number of target distribution feature vectors whose distances from the distribution feature vector of the to-be-recognized audio satisfy a preset first distance condition from the distribution feature vectors of each sample audio;
[0034] Screen out the language category with the largest number of language categories from the language categories corresponding to the target distribution feature vectors as the language category of the to-be-recognized audio.
[0035] In one embodiment, obtaining the target codebook feature corresponding to the audio feature of the to-be-recognized audio from the audio codebook in the pre-trained audio feature extraction model includes:
[0036] Screen out the target codebook feature vectors corresponding to each audio feature vector in the audio feature of the to-be-recognized audio from the codebook feature vectors in the audio codebook;
[0037] Combine each target codebook feature vector into the target codebook feature corresponding to the audio feature of the audio to be recognized.
[0038] In a second aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0039] Input the audio to be recognized into a pre-trained audio feature extraction model to obtain the audio feature of the audio to be recognized; the pre-trained audio feature extraction model is trained by sample audios in different languages;
[0040] Obtain the target codebook feature corresponding to the audio feature of the audio to be recognized from the audio codebook included in the pre-trained audio feature extraction model; the audio codebook includes codebook feature vectors in different languages;
[0041] Obtain the distribution feature vector of the audio to be recognized according to the distribution of each codebook feature vector in the target codebook feature;
[0042] Determine the language category corresponding to the target distribution feature vector whose distance from the distribution feature vector of the audio to be recognized in the distribution feature vectors of the sample audios satisfies a preset distance condition as the language category of the audio to be recognized.
[0043] In a third aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the following steps are implemented:
[0044] Input the audio to be recognized into a pre-trained audio feature extraction model to obtain the audio feature of the audio to be recognized; the pre-trained audio feature extraction model is trained by sample audios in different languages;
[0045] Obtain the target codebook feature corresponding to the audio feature of the audio to be recognized from the audio codebook included in the pre-trained audio feature extraction model; the audio codebook includes codebook feature vectors in different languages;
[0046] Obtain the distribution feature vector of the audio to be recognized according to the distribution of each codebook feature vector in the target codebook feature;
[0047] Determine the language category corresponding to the target distribution feature vector whose distance from the distribution feature vector of the audio to be recognized in the distribution feature vectors of the sample audios satisfies a preset distance condition as the language category of the audio to be recognized.
[0048] Fourthly, the present application also provides a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the following steps:
[0049] Input the audio to be recognized into a pre-trained audio feature extraction model to obtain the audio features of the audio to be recognized; the pre-trained audio feature extraction model is trained with sample audios in different languages;
[0050] Obtain, from the audio codebook included in the pre-trained audio feature extraction model, the target codebook features corresponding to the audio features of the audio to be recognized; the audio codebook includes codebook feature vectors in different languages;
[0051] Obtain the distribution feature vector of the audio to be recognized according to the distribution of each codebook feature vector in the target codebook features;
[0052] Determine, as the language category of the audio to be recognized, the language category corresponding to the target distribution feature vector whose distance from the distribution feature vector of the audio to be recognized in the distribution feature vectors of the sample audios meets a preset distance condition.
[0053] For the above language recognition method, computer device, storage medium and computer program product, the audio to be recognized is input into a pre-trained audio feature extraction model to obtain the audio features of the audio to be recognized; the pre-trained audio feature extraction model is trained with sample audios in different languages; the target codebook features corresponding to the audio features of the audio to be recognized are obtained from the audio codebook included in the pre-trained audio feature extraction model; the audio codebook includes codebook feature vectors in different languages; the distribution feature vector of the audio to be recognized is obtained according to the distribution of each codebook feature vector in the target codebook features; and the language category corresponding to the target distribution feature vector whose distance from the distribution feature vector of the audio to be recognized in the distribution feature vectors of the sample audios meets a preset distance condition is determined as the language category of the audio to be recognized. By using this method, unsupervised learning can be performed based on unlabeled sample audios in different languages to obtain a pre-trained audio feature extraction model, and then the target codebook features of the audio to be recognized can be obtained through the pre-trained audio feature extraction model, without predicting the language category of the audio to be recognized through the pre-trained audio feature extraction model; the language category of the audio to be recognized is determined according to the language category corresponding to the target distribution feature vector whose distance from the feature vector meets a preset distance condition, which has the advantage of relatively high language recognition accuracy, and can also improve the recognition ability for the audio to be recognized in a new language category without retraining the pre-trained audio feature extraction model. Description of the Drawings
[0054] Figure 1It is an application environment diagram of a language recognition method in an embodiment;
[0055] Figure 2 It is a schematic flowchart of a language recognition method in an embodiment;
[0056] Figure 3 It is a schematic flowchart of the training steps of a pre-trained audio feature extraction model in an embodiment;
[0057] Figure 4 It is a schematic structural diagram of an audio encoding model and an audio decoding model in an audio feature extraction model in an embodiment;
[0058] Figure 5 It is a schematic structural diagram of a speaker encoding model in an audio feature extraction model in an embodiment;
[0059] Figure 6 It is a schematic flowchart of a language recognition method in another embodiment;
[0060] Figure 7 It is a schematic diagram of the training of an audio feature extraction model in an embodiment;
[0061] Figure 8 It is an internal structural diagram of a computer device in an embodiment. Detailed implementation manners
[0062] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0063] The language recognition method provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 101 communicates with the server 102 through a network. The terminal 101 can collect audio data such as initial audio or audio to be recognized, and can also provide related services for language recognition; the server 102 can generally refer to a background system that provides related services for language recognition. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or can be placed in the cloud or other network servers. Among them, the terminal 101 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 102 can be implemented by an independent server or a server cluster composed of multiple servers.
[0064] In one embodiment, the terminal 101 inputs the acquired audio to be recognized into a pre-trained audio feature extraction model to obtain the audio features of the audio to be recognized; the pre-trained audio feature extraction model is trained with sample audios in different languages; from the audio codebook included in the pre-trained audio feature extraction model, the target codebook feature corresponding to the audio features of the audio to be recognized is obtained; the audio codebook includes codebook feature vectors in different languages; according to the distribution of each codebook feature vector in the target codebook feature, the distribution feature vector of the audio to be recognized is obtained; the language category corresponding to the target distribution feature vector whose distance from the distribution feature vector of the audio to be recognized in the distribution feature vectors of the sample audios satisfies the preset distance condition is determined as the language category of the audio to be recognized. Therefore, the execution subject of the above language recognition method can be the terminal 101.
[0065] In one embodiment, the above language recognition method can also be implemented based on the server 102 alone. For example, the server 102 can obtain the audio to be recognized from the background database and obtain the language category of the audio to be recognized by executing the above language recognition method.
[0066] In one embodiment, the above language recognition method can also be implemented based on the interaction between the terminal 101 and the server. For example, after the terminal 101 acquires the audio to be recognized, it sends the audio to be recognized to the server 102. The server 102 obtains the language category of the audio to be recognized by executing the above language recognition method, and then the server 102 can also send the language category of the audio to be recognized to the terminal 101 for display. For another example, the server can obtain the dry voice from the background database, send the audio to be recognized to the terminal 101, and the terminal 101 obtains the language category of the audio to be recognized by executing the above language recognition method. The terminal 101 can also upload the language category of the audio to be recognized to the server for storage.
[0067] As can be seen from the above, in this exemplary embodiment, the execution subject of the above language recognition method can be the above terminal 101 or the server 102, and it can also be applied to a system including the terminal 101 and the server 102 and implemented through the interaction between the terminal 101 and the server 102. The present disclosure does not limit this.
[0068] In one embodiment, as Figure 2 shown, a language recognition method is provided. Taking the method applied to the Figure 1 terminal as an example, the method includes the following steps:
[0069] Step S201: Input the audio to be recognized into a pre-trained audio feature extraction model to obtain the audio features of the audio to be recognized; the pre-trained audio feature extraction model is trained with sample audios in different languages.
[0070] Among them, the audio to be recognized refers to the audio data of the language category to be recognized. The audio to be recognized can be speech data and singing data.
[0071] Among them, the audio feature extraction model refers to a model used to extract the features of audio data (such as the audio to be recognized); the audio feature extraction model can be a standard variational autoencoder in deep learning, or a variational autoencoder based on vector quantization. The audio feature extraction model can also be a model composed of an audio encoding model, an audio decoding model, and a speaker encoding model. Of course, the audio feature extraction model can also be changed accordingly according to the feature extraction requirements of the audio to be recognized.
[0072] Among them, the audio features of the audio to be recognized refer to the information that can characterize the prominent sound properties such as the pitch, intonation, and rhythm changes of the audio to be recognized.
[0073] Among them, the sample audio refers to the audio data used to train the audio feature extraction model to be trained; the sample audio can be labeled audio data or unlabeled audio data.
[0074] Specifically, the terminal trains the audio feature extraction model to be trained with multiple sample audios in different languages to obtain a pre-trained audio feature extraction model; in practical applications, the multiple sample audios in different languages can include a large number of unlabeled sample audios and a small number of labeled sample audios.
[0075] An audio acquisition device can be deployed on the terminal. Furthermore, the terminal can collect the user's initial audio, or can also collect the user's initial audio through a separate audio acquisition device and send the initial audio to the terminal. Further, the terminal preprocesses the initial audio to obtain the audio to be recognized; then the audio to be recognized is input into the audio encoding model in the pre-trained audio feature extraction model, and the audio encoding model performs feature extraction processing on the audio to be recognized to obtain the audio features of the audio to be recognized. At the same time, the terminal inputs the audio to be recognized into the speaker encoding model in the pre-trained audio feature extraction model to obtain the speaker features of the audio to be recognized.
[0076] In practical applications, the terminal can convert the audio to be recognized into Mel spectrogram features; among them, the dimension of the Mel spectrogram features is [T, M], where T represents the number of frames of the audio to be recognized, and M represents the dimension of the Mel spectrogram, and M can be set to 80. The Mel spectrogram features are input into the audio encoding model in the pre-trained audio feature extraction model to obtain the audio features of the Mel spectrogram features as the audio features of the audio to be recognized.
[0077] Step S202: Obtain a target codebook feature corresponding to the audio feature of the audio to be recognized from the audio codebook included in the pre-trained audio feature extraction model; the audio codebook includes codebook feature vectors of different languages.
[0078] It should be noted that the audio codebook is a codebook shared by the encoder and decoder in the pre-trained audio feature extraction model. The audio codebook refers to a set in the form of a lookup table with codebook feature vectors of different languages. The target codebook feature refers to the feature information composed of multiple codebook feature vectors in the audio codebook whose similarity with the audio feature of the audio to be recognized meets a preset similarity condition. The codebook feature vector refers to the latent variable of audio in different languages.
[0079] Specifically, the audio feature of the audio to be recognized contains audio feature vectors of each frame of the audio to be recognized. The terminal performs similarity processing on each audio feature vector in the audio feature of the audio to be recognized and each codebook feature vector in the audio codebook to obtain the similarity between each audio feature vector in the audio feature and each codebook feature vector in the audio codebook; from each codebook feature vector in the audio codebook, filter out the codebook feature vectors whose similarity with each audio feature vector meets the preset similarity condition as the target codebook feature vectors corresponding to each audio feature vector, and all the target codebook feature vectors together serve as the target codebook feature corresponding to the audio feature of the audio to be recognized. In practical applications, the terminal can implement similarity processing through nearest neighbor search.
[0080] Step S203: Obtain the distribution feature vector of the audio to be recognized according to the distribution of each codebook feature vector in the target codebook feature.
[0081] Specifically, the target codebook feature corresponding to the audio to be recognized is composed of multiple codebook feature vectors. The terminal determines the distribution of each codebook feature vector corresponding to the audio to be recognized according to the number of each codebook feature vector in the target codebook feature; and then normalizes the distribution of each codebook feature vector to obtain the distribution feature vector of the audio to be recognized.
[0082] Step S204: Determine the language category corresponding to the target distribution feature vector whose distance from the distribution feature vector of the audio to be recognized in the distribution feature vector of the sample audio meets the preset distance condition as the language category of the audio to be recognized.
[0083] Among them, the preset distance condition refers to the determination condition set for the distance between the distribution feature vector of the sample audio and the distribution feature vector of the audio to be recognized. The language category refers to the language to which the audio data belongs; for example, Chinese, English, Japanese, etc. The distribution feature vector refers to the data obtained by processing the frequency distribution of the codebook feature vectors in the target codebook feature.
[0084] Specifically, during the process of training using sample audio, the terminal can store the language category of each sample audio. Then, when determining the language category of the audio to be recognized, it can query from the distribution feature vectors of the sample audio (for the convenience of distinguishing from the distribution feature vectors of other audios, it can be called sample distribution feature vectors) the sample distribution feature vectors whose distances from the distribution feature vectors of the audio to be recognized meet the preset distance condition as the target distribution feature vectors. Then, among the language categories of the target distribution feature vectors, the language categories that meet the preset language category quantity condition are used as the language category of the audio to be recognized.
[0085] In practical applications, the terminal can use a parameter-free classification method to determine the language category of the audio to be recognized. For example, the K Nearest Neighbor (KNN) method.
[0086] In the above language recognition method, the audio to be recognized is input into a pre-trained audio feature extraction model to obtain the audio features of the audio to be recognized; the pre-trained audio feature extraction model is trained with sample audios of different languages; from the audio codebook included in the pre-trained audio feature extraction model, the target codebook features corresponding to the audio features of the audio to be recognized are obtained; the audio codebook includes codebook feature vectors of different languages; according to the distribution of each codebook feature vector in the target codebook features, the distribution feature vectors of the audio to be recognized are obtained; the language category corresponding to the target distribution feature vector whose distance from the distribution feature vector of the sample audio to the distribution feature vector of the audio to be recognized meets the preset distance condition is determined as the language category of the audio to be recognized. By using this method, unsupervised learning can be performed based on unlabeled sample audios of different languages to obtain a pre-trained audio feature extraction model, and then the target codebook features of the audio to be recognized can be obtained through the pre-trained audio feature extraction model, without predicting the language category of the audio to be recognized through the pre-trained audio feature extraction model; the language category of the audio to be recognized is determined according to the language category corresponding to the target distribution feature vector whose distance from the feature vector meets the preset distance condition, which has the advantage of relatively high language recognition accuracy, and can also improve the recognition ability of the audio to be recognized in a new language category without retraining the pre-trained audio feature extraction model.
[0087] In one embodiment, as Figure 3 shown, the pre-trained audio feature extraction model in the above step S201 is trained in the following manner:
[0088] Step S301, input sample audios of different languages into the audio encoding model and the speaker encoding model in the audio feature extraction model to be trained respectively, and obtain the sample audio features and speaker features corresponding to the sample audios.
[0089] Among them, the audio encoding model refers to a model used to encode audio data (such as sample audio). This model can identify the characteristics of the audio data and perform encoding conversion. The speaker encoding model refers to a model used to encode the data of the speaker corresponding to the audio data (such as sample audio). This model can identify the voice characteristics of the speaker corresponding to the audio data and perform encoding conversion.
[0090] Among them, the sample audio feature refers to information that can characterize the prominent properties of the sample audio, such as pitch, intonation, and rhythm changes. The speaker feature refers to information that characterizes the prominent properties of the speaker, such as timbre, pitch, and vocalization frequency. Specifically, the audio feature extraction model to be trained includes an audio encoding model, a speaker encoding model, and an audio decoding model; the encoder in the audio encoding model processes the sample audio to obtain the sample audio feature corresponding to the sample audio. The speaker encoding model processes the sample audio to obtain the speaker feature corresponding to the sample audio.
[0091] It should be noted that the execution time of the speaker encoding model can be before the audio encoding model, can be executed simultaneously with the audio encoding model, or can be after the audio encoding model. The execution order of the audio encoding model and the speaker encoding model is not specifically limited here.
[0092] Step S302: Obtain the sample codebook feature corresponding to the sample audio feature from the audio codebook included in the audio feature extraction model to be trained.
[0093] Among them, the audio codebook in the audio feature extraction model to be trained is the same as the audio codebook in the pre-trained audio feature extraction model.
[0094] Specifically, the sample audio feature of the sample audio includes the sample audio feature vectors of each frame of the sample audio. The terminal processes to obtain the similarity between each sample audio feature vector in the sample audio feature and each codebook feature vector in the audio codebook; from each codebook feature vector in the audio codebook, the codebook feature vectors whose similarity to each sample audio feature vector meets the preset similarity condition are screened out as the sample codebook feature vectors corresponding to each sample audio feature vector, and all the sample codebook feature vectors together serve as the sample codebook feature corresponding to the sample audio feature.
[0095] In practical applications, the terminal can set the preset similarity condition to the codebook feature vector with the highest similarity. For example, assume that the sample audio features include sample audio feature vector A and sample audio feature vector B, and the audio codebook includes codebook feature vector 1, codebook feature vector 2, codebook feature vector 3, and codebook feature vector 4. If the similarities between sample audio feature vector A and codebook feature vectors 1, 2, 3, and 4 are 70, 82, 85, and 92 respectively, then codebook feature vector 4 is used as the sample codebook feature vector corresponding to sample audio feature vector A. If the similarities between sample audio feature vector B and codebook feature vectors 1, 2, 3, and 4 are 65, 77, 95, and 80 respectively, then codebook feature vector 3 is used as the sample codebook feature vector corresponding to sample audio feature vector B.
[0096] Step S303: Concatenate the sample codebook feature and the speaker feature, and input the concatenated feature obtained by concatenation into the audio decoding model in the audio feature extraction model to be trained, to obtain the predicted audio feature of the sample audio.
[0097] Among them, the audio decoding model refers to a model used to perform decoding processing on audio data (such as the concatenated feature in step S303 above). The predicted audio feature of the sample audio refers to the prediction result corresponding to the input sample audio output by the audio feature extraction model to be trained. The actual audio feature of the sample audio refers to the speech feature obtained by processing the sound signal of the sample audio. For example, according to the spectrogram of the sample audio, the sample audio is processed into Mel spectrogram features, then the Mel spectrogram features obtained after processing can be regarded as the actual audio features of the sample audio. When the sample audio is input into the audio feature extraction model and the Mel spectrogram features output by the audio feature extraction model are obtained, then the Mel spectrogram features output by the audio feature extraction model can be regarded as the predicted audio features of the sample audio.
[0098] Specifically, the terminal performs concatenation processing on the sample codebook feature and the speaker feature to obtain a concatenated feature. Then, the terminal processes the concatenated feature through the audio decoding model in the audio feature extraction model to be trained to obtain the predicted audio feature of the sample audio.
[0099] Step S304: According to the difference between the predicted audio feature of the sample audio and the actual audio feature of the sample audio, perform iterative training on the audio feature extraction model to be trained to obtain a pre-trained audio feature extraction model.
[0100] Specifically, the terminal obtains the actual audio feature of the sample audio, then calculates the mean square error between the predicted audio feature and the actual audio feature of the sample audio, and constructs a loss function based on the mean square error. The audio feature extraction model to be trained is iteratively trained through the loss function to obtain a pre-trained audio feature extraction model.
[0101] In this embodiment, the audio coding model and the speaker coding model in the audio feature extraction model to be trained are used to process sample audio in different languages, and the corresponding sample audio features and speaker features of the sample audio are obtained; the audio decoding model in the audio feature extraction model to be trained is used to process the concatenated features obtained by concatenating the sample audio features and the speaker features, and the predicted audio features of the sample audio are obtained; furthermore, according to the difference between the predicted audio features of the sample audio and the actual audio features of the sample audio, the audio feature extraction model to be trained is iteratively trained to obtain a pre-trained audio feature extraction model, which can extract more feature information in terms of audio and speaker in sample audio in different languages, and is beneficial to improving the recognition accuracy of language categories in subsequent steps.
[0102] In one embodiment, the sample audio features in step S301 are obtained through the following method: the initial sample features of the sample audio are sequentially subjected to convolution processing, batch normalization processing, and activation processing through the audio coding model in the audio feature extraction model to be trained, and the processed sample features of the sample audio are obtained; the convolution processing, batch normalization processing, and activation processing are sequentially performed on the fused features obtained by fusing the processed sample features of the sample audio and the initial sample features of the sample audio, and the encoded features of the sample audio are obtained; the dimensionality reduction processing is performed on the encoded features of the sample audio, and the dimensionality-reduced features of the sample audio are obtained; the dimensionality-reduced features of the sample audio are input into a gated recurrent network to obtain the sample audio features of the sample audio.
[0103] Among them, the activation processing can be implemented based on the Rectified Linear Unit (ReLU). The fused feature refers to the feature obtained by fusing the input feature and the feature output after the input feature is sequentially subjected to convolution processing, batch normalization processing, and activation processing.
[0104] Specifically, the terminal obtains sample audio in different languages; performs feature conversion processing on the sample audio to obtain the corresponding initial sample features of the sample audio; in practical applications, the terminal can convert the sample audio in different languages into corresponding Mel spectrogram features, and then use the audio coding model and the speaker coding model in the audio feature extraction model to be trained to process the Mel spectrogram features of each frame of audio in the sample audio in different languages respectively, and obtain the corresponding sample audio features and speaker features.
[0105] Further, the terminal sequentially performs convolution processing, batch normalization (Batchnorm) processing, and activation processing on the initial sample features through an audio encoding model to obtain processed sample features corresponding to the initial sample features; fuses the processed sample features with the initial sample features to obtain a fusion feature between the processed sample features and the initial sample features; and sequentially performs convolution processing, batch normalization processing, and activation processing on the fusion feature again to obtain encoded features. In subsequent steps, the terminal will repeatedly perform fusion processing on the features obtained by performing convolution processing, batch normalization processing, and activation processing on the input, and the features output after sequentially performing convolution processing, batch normalization processing, and activation processing, and then use the features obtained by the fusion processing as the input to perform convolution processing, batch normalization processing, and activation processing again until a preset loop termination condition is met. The terminal obtains the encoded features when the preset loop termination condition is met, inputs the encoded features into a linear layer (Linear) for dimensionality reduction processing to obtain corresponding dimensionality-reduced features; and inputs the dimensionality-reduced features into a gated recurrent network to obtain sample audio features of the sample audio. Among them, the gated recurrent network can be a standard GRU network (Gate Recurrent Unit), or a bidirectional gated recurrent unit (Bidirectional Gated Recurrent Unit, Bi-GRU), and of course it can also be other networks deformed from GRU.
[0106] In practical applications, the audio encoding model and the audio decoding model in the audio feature extraction model can have the same model structure. Figure 4 As shown in the structural schematic diagram of the audio encoding model and the audio decoding model in the audio feature extraction model, Figure 4 as shown, a convolutional unit with a kernel size of 3 and a channel number of 128 is constructed, denoted as Conv(3,128), and a batch processing unit and a ReLU-based activation unit are constructed. The convolutional unit, the batch processing unit, and the activation unit are used as a residual connection structure. Among them, the number of residual connection structures in the audio encoding model or the audio decoding model can be 5; taking the sample audio features as an example, first input the initial sample features of the sample audio features into 5 residual connection structures, output the encoded features, and then input the encoded features into a linear layer (Linear) for dimensionality reduction processing to obtain corresponding dimensionality-reduced features; input the dimensionality-reduced features into a Bi-GRU to obtain the sample audio features of the sample audio. Further, the processing method for the audio to be recognized is the same.
[0107] In this embodiment, through the audio encoding model in the audio feature extraction model to be trained, the initial sample features of the sample audio are sequentially subjected to convolution processing, batch normalization processing, and activation processing to obtain the processed sample features of the sample audio; the convolution processing, batch normalization processing, and activation processing are sequentially performed on the fusion features obtained by fusing the processed sample features of the sample audio and the initial sample features of the sample audio to obtain the encoded features of the sample audio; the dimensionality reduction processing is performed on the encoded features of the sample audio to obtain the dimensionality-reduced features of the sample audio; and then the dimensionality-reduced features of the sample audio are input into the gated recurrent network to obtain the sample audio features of the sample audio. By fusing the processed sample features of the sample audio and the initial sample features of the sample audio, the vanishing gradient of the audio encoding model can be avoided, and the audio encoding model in the audio feature extraction model can learn the sample audio features in the sample audio, thereby improving the accuracy of the trained audio feature extraction model.
[0108] In one embodiment, the speaker features in the above step S301 can be obtained through the following method: through the speaker encoding model in the audio feature extraction model to be trained, the initial sample features of the sample audio are sequentially subjected to convolution processing, batch normalization processing, and activation processing to obtain the processed features of the sample audio; the convolution processing, batch normalization processing, and activation processing are sequentially performed on the fusion features obtained by fusing the processed sample features of the sample audio and the initial sample features of the sample audio to obtain the encoded features of the sample audio; the dimensionality reduction processing is performed on the encoded features of the sample audio to obtain the dimensionality-reduced features of the sample audio; and the mean value processing is performed on the dimensionality-reduced features of the sample audio to obtain the speaker features corresponding to the sample audio.
[0109] Specifically, after the terminal obtains sample audio in different languages, it can perform feature transformation processing on the sample audio to obtain the initial sample features corresponding to the sample audio; the terminal sequentially performs convolution processing, batch normalization (Batchnorm) processing, and activation processing on the initial sample features through a speaker encoding model to obtain the processed sample features corresponding to the initial sample features; the processed sample features and the initial sample features are fused to obtain the fusion features between the processed sample features and the initial sample features; convolution processing, batch normalization processing, and activation processing are sequentially performed on the fusion features to obtain the encoded features. In subsequent steps, the terminal will repeatedly perform fusion processing on the features obtained by performing convolution processing, batch normalization processing, and activation processing on the input, and the features output after performing convolution processing, batch normalization processing, and activation processing, and then use the features obtained by the fusion processing as the input to perform convolution processing, batch normalization processing, and activation processing again until the preset loop termination condition is met. The terminal obtains the encoded features when the preset loop termination condition is met, inputs the encoded features into a linear layer (Linear) for dimensionality reduction processing to obtain the corresponding dimensionality reduction features; the dimensionality reduction features are averaged to obtain the speaker features corresponding to the sample audio.
[0110] In practical applications, Figure 5 FIG. is a schematic structural diagram of a speaker encoding model in an audio feature extraction model, as Figure 5 shown, a convolutional unit with a kernel size of 3 and a channel number of 128 is constructed, denoted as Conv(3,128), and a batch processing unit and a ReLU-based activation unit are constructed. The convolutional unit, batch processing unit, and activation unit are used as a residual connection structure. Among them, there can be 5 residual connection structures in the speaker encoding model; taking the sample audio features as an example, first input the initial sample features of the sample audio features into 5 residual connection structures, output the encoded features, and then input the encoded features into a linear layer (Linear) for dimensionality reduction processing to obtain the corresponding dimensionality reduction features; the dimensionality reduction features are averaged to obtain the speaker features corresponding to the sample audio. Further, the processing method for the audio to be recognized is the same.
[0111] In this embodiment, through the speaker encoding model in the audio feature extraction model to be trained, the initial sample features of the sample audio are subjected to convolution processing, batch normalization processing, and activation processing to obtain the processed features of the sample audio; the fusion features between the processed sample features of the sample audio and the initial sample features of the sample audio are subjected to convolution processing, batch normalization processing, and activation processing again to obtain the encoded features of the sample audio; then, the encoded features of the sample audio are subjected to dimensionality reduction processing to obtain the dimensionality-reduced features of the sample audio; the dimensionality-reduced features of the sample audio are subjected to mean processing to obtain the speaker features corresponding to the sample audio; by fusing the processed sample features of the sample audio and the initial sample features of the sample audio, the gradient disappearance of the speaker encoding model can be avoided, and by enabling the speaker encoding model in the audio feature extraction model to learn the speaker features in the sample audio, the accuracy of the pre-trained audio feature extraction model is improved.
[0112] In one embodiment, before step S201 of inputting the sample audio in different languages into the audio encoding model and the speaker encoding model in the audio feature extraction model to be trained to obtain the sample audio features and speaker features corresponding to the sample audio, it further includes: obtaining the initial audio in different languages; performing voice activation processing on the initial audio in each language to obtain the valid audio in the initial audio in each language; and respectively performing speed change processing and / or pitch change processing on the valid audio in each language according to the duration of the valid audio in each language to obtain the sample audio.
[0113] Among them, the valid audio refers to the audio segment with valid vocalization in the initial audio.
[0114] Specifically, the terminal can obtain the initial audio in different languages from multiple channels such as the Internet, singing software, and databases; through the Voice Activity Detection (VAD) technology, perform voice activation processing on the initial audio in each language to obtain the silent audio in the initial audio in each language, and by deleting the silent audio in the initial audio, obtain the valid audio in the initial audio; respectively perform speed change processing and / or pitch change processing on the valid audio in each language according to the duration of the valid audio in each language to obtain the processed audio in each language; the terminal uses the processed audio as the sample audio.
[0115] It should be noted that the durations of the valid audio in the initially obtained audio in different languages may vary. However, even if the duration of the valid audio in a certain initial audio is relatively short, the initial audio will not be processed to slow down. Instead, the total duration of the valid audio in each language is counted. For the languages with relatively short total durations of valid audio, the terminal can perform more times of speed change processing and pitch change processing to enhance the sample audio of the languages with relatively short total durations of valid audio. For example, if the valid audio in Chinese has been subjected to 5 times of speed change processing and pitch change, then the valid audio in Chinese is expanded by five times.
[0116] In this embodiment, by obtaining the initial audio in different languages; performing voice activation processing on the initial audio of each language to obtain the valid audio in the initial audio of each language; and performing speed change processing and / or pitch change processing on the valid audio of each language respectively according to the duration of the valid audio of each language to obtain the sample audio, so that the sample audio in different languages has a balanced duration, thereby effectively improving the training effect of the pre-trained audio feature extraction model, and thus improving the feature extraction ability of the pre-trained audio feature extraction model for the audio to be recognized.
[0117] In one embodiment, in the above step S203, according to the distribution of each codebook feature vector in the target codebook feature, the distribution feature vector of the audio to be recognized is obtained, which specifically includes the following content: according to the quantity distribution of the codebook feature vectors in the target codebook feature, the histogram of the target codebook feature is obtained; and the histogram of the target codebook feature is normalized to obtain the distribution feature vector of the audio to be recognized.
[0118] Among them, the histogram is used to represent the distribution of each codebook feature vector in the target codebook feature.
[0119] Specifically, the terminal detects the quantity of each codebook feature vector in the target codebook feature, and then the terminal generates the histogram of the target codebook feature according to the quantity of each codebook feature vector; and the histogram of the target codebook feature is normalized to obtain the distribution feature vector of the audio to be recognized. It should be noted that since each codebook feature vector in the target codebook feature is selected from the codebook feature vectors in different languages carried by the audio codebook, the quantity of each codebook feature vector in the target codebook feature can also be regarded as the number of times the codebook feature vectors in the audio codebook are selected; by converting the quantity distribution of the codebook feature vectors in the target codebook feature into a feature vector, the similarity degree between the distribution feature vector of the audio to be recognized and the distribution feature vectors of the sample audio in different languages can be analyzed, and then the language category of the audio to be recognized can be reasonably determined.
[0120] In this embodiment, according to the quantity distribution of the codebook feature vectors in the target codebook feature, a histogram of the target codebook feature is obtained; the histogram of the target codebook feature is normalized to obtain the distribution feature vector of the audio to be recognized, realizing the conversion of the codebook feature vectors in the target codebook feature into the distribution feature vector of the audio to be recognized, which is beneficial to performing the language recognition step based on the distribution feature vector of the audio to be recognized in the subsequent stage.
[0121] In one embodiment, for step S204 above, the language category corresponding to the target distribution feature vector whose distance from the distribution feature vector of the audio to be recognized in the distribution feature vectors of the sample audios meets the preset distance condition is determined as the language category of the audio to be recognized, and specifically includes the following content: from the distribution feature vectors of each sample audio, a preset number of target distribution feature vectors whose distance from the distribution feature vector of the audio to be recognized meets the preset first distance condition are screened out; from the language categories corresponding to the target distribution feature vectors, the language category with the largest number of language categories is selected as the language category of the audio to be recognized.
[0122] Among them, the first distance condition refers to the condition set for the distance from the distribution feature vector of the audio to be recognized. The first distance condition is used to screen out the target distribution feature vectors from the distribution feature vectors of each sample audio, and the first distance condition can be adaptively set according to the actual application scenario.
[0123] Specifically, the terminal inputs the sample audio into a pre-trained audio feature extraction model to obtain the sample audio feature corresponding to the sample audio, obtains the sample codebook feature corresponding to the sample audio feature from the audio codebook included in the pre-trained audio feature extraction model, obtains the distribution feature vector of the sample audio according to the distribution of each sample codebook feature vector in the sample codebook feature, and obtains the language category of each sample audio; the terminal detects the distance between the distribution feature vector of the audio to be recognized and the distribution feature vectors of each sample audio. The terminal can adopt the K-Nearest Neighbor (KNN) method. By traversing the distribution feature vectors of each sample audio, a preset number (denoted as K) of target distribution feature vectors whose distance from the distribution feature vector of the audio to be recognized meets the first distance condition are obtained. For example, K can be set to 16, and the first distance condition can also be set to the 16 sample audio distribution feature vectors that are closest to the distribution feature vector of the audio to be recognized; the language categories corresponding to each target distribution feature vector are obtained, and the language category with the largest number of language categories is used as the language category of the audio to be recognized.
[0124] In this embodiment, by screening a preset number of target distribution feature vectors from the distribution feature vectors of each sample audio, the distance between which and the distribution feature vector of the audio to be recognized satisfies a preset first distance condition; and screening out the language category with the largest number of language categories from the language categories corresponding to the target distribution feature vectors as the language category of the audio to be recognized, a reasonable judgment of the language category of the audio to be recognized is achieved.
[0125] In one embodiment, in step S202 above, obtaining the target codebook feature corresponding to the audio feature of the audio to be recognized from the audio codebook in the pre-trained audio feature extraction model specifically includes the following content: screening out the target codebook feature vectors corresponding to each audio feature vector in the audio feature of the audio to be recognized from the codebook feature vectors in the audio codebook; and combining each target codebook feature vector into the target codebook feature corresponding to the audio feature of the audio to be recognized.
[0126] Specifically, before the above step S202, the terminal pre-constructs an audio codebook (codebook) with dimensions [N, D], where the audio codebook can be regarded as a discrete latent space, N represents the size of the discrete latent space, that is, the size of the audio codebook, and D represents the dimension of each latent embedding vector in the latent space, that is, the dimension of each codebook feature vector in the audio codebook. Then the terminal performs a similarity process on each audio feature vector in the audio feature of the audio to be recognized and each codebook feature vector in the audio codebook to obtain the similarity between each audio feature vector in the audio feature and each codebook feature vector in the audio codebook; furthermore, the terminal can select, by means of nearest neighbor search, the codebook feature vectors whose similarity with each audio feature vector satisfies a preset similarity condition from the codebook feature vectors in the audio codebook as the target codebook feature vectors corresponding to each audio feature vector, and all the target codebook feature vectors are combined into the target codebook feature corresponding to the audio feature of the audio to be recognized. Mark the target codebook feature vector as Enc′(x), and Enc′(x) can be calculated in the following way:
[0127] Enc′(x) = e k , where k = argmin j ||Enc(x) - e j ||2
[0128] where Enc′(x) represents the target codebook feature vector corresponding to the x-th frame of audio in the audio to be recognized; Enc(x) represents the audio feature vector corresponding to the x-th frame of audio in the audio to be recognized; e j represents the j-th codebook feature vector in the audio codebook; Enc(x) - e j can be expressed as the audio feature vector Enc(x) and the j-th codebook feature vector e in the audio codebookj The similarity between; e k It represents the k-th codebook feature vector in the audio codebook whose similarity with the audio feature vector Enc(x) satisfies a preset similarity condition.
[0129] Mark the target codebook feature as Z. Combine all the target codebook feature vectors into the target codebook feature corresponding to the audio feature of the audio to be recognized. Then the target codebook feature can be represented in the following way: Z = {Enc′(1), Enc′(2), …, Enc′(x)}.
[0130] It should be noted that the method for the audio feature extraction model to be trained to obtain the sample codebook feature vector from the sample audio feature vector of the sample audio is the same as the calculation method of the target codebook feature vector Enc′(x).
[0131] In this embodiment, by screening from the codebook feature vectors in the audio codebook, the target codebook feature vectors corresponding to each audio feature vector in the audio feature of the audio to be recognized are obtained; combining each target codebook feature vector into the target codebook feature corresponding to the audio feature of the audio to be recognized. Compared with the traditional technology that directly outputs the distribution feature vector of the model to be recognized through the model, the audio feature extraction model to be trained in this application can perform unsupervised learning based on a large number of unlabeled sample audios in different languages, so that the pre-trained audio feature extraction model can still have good feature extraction ability for the audio to be recognized in a new language without learning the sample audio of the new language, thereby improving the recognition ability for the audio to be recognized in a new language category.
[0132] In one embodiment, as Figure 6 shown, another language recognition method is provided. Taking the application of this method to a terminal as an example for illustration, it includes the following steps:
[0133] Step S601, obtain the initial audios in different languages; perform voice activation processing on the initial audios of each language to obtain the valid audios in the initial audios of each language.
[0134] Step S602, respectively perform speed change processing and / or pitch change processing on the valid audios of each language according to the duration of the valid audios of each language to obtain sample audios.
[0135] Step S603, input the audio to be recognized into the pre-trained audio feature extraction model to obtain the audio feature of the audio to be recognized; wherein, the pre-trained audio feature extraction model is trained by sample audios in different languages.
[0136] Step S604: From the codebook feature vectors in the audio codebook, screen out the target codebook feature vectors corresponding to each audio feature vector in the audio features of the audio to be recognized; combine each target codebook feature vector into the target codebook feature corresponding to the audio features of the audio to be recognized.
[0137] Step S605: Obtain the histogram of the target codebook feature according to the quantity distribution of the codebook feature vectors in the target codebook feature; perform normalization processing on the histogram of the target codebook feature to obtain the distribution feature vector of the audio to be recognized.
[0138] Step S606: From the distribution feature vectors of each sample audio, screen out a preset number of target distribution feature vectors whose distances from the distribution feature vector of the audio to be recognized satisfy a preset first distance condition.
[0139] Step S607: From the language categories corresponding to the target distribution feature vectors, screen out the language category with the largest number of language categories as the language category of the audio to be recognized.
[0140] The above language identification method can achieve the following beneficial effects: It can perform unsupervised learning based on unlabeled sample audios of different languages to obtain a pre-trained audio feature extraction model, and then obtain the target codebook feature of the audio to be recognized through the pre-trained audio feature extraction model, without predicting the language category of the audio to be recognized through the pre-trained audio feature extraction model; determine the language category of the audio to be recognized according to the language category corresponding to the target distribution feature vector whose distance from the feature vector satisfies the preset distance condition, which has the advantage of high language identification accuracy, and can also improve the identification ability of the audio to be recognized in a new language category without retraining the pre-trained audio feature extraction model.
[0141] To more clearly illustrate the language identification method provided by the embodiments of the present disclosure, the above language identification method will be specifically described below with a specific embodiment. Another language identification method is provided, which can be applied to Figure 1 the terminal in
[0142] (1) Data preparation: The terminal obtains initial audio in different languages from multiple sources such as the Internet, singing software, and databases; performs voice activity detection (VAD) processing on the initial audio of each language to obtain the silent audio in the initial audio of each language, and obtains the valid audio in the initial audio by deleting the silent audio in the initial audio; according to the duration of the valid audio of each language, performs speed change processing and pitch change processing on the valid audio of each language to obtain the processed audio of each language; the terminal uses the processed audio as the sample audio to make the durations of the audio in each language in the sample audio more balanced.
[0143] (2) Model training: Figure 7 Schematic diagram for the training of the audio feature extraction model, as Figure 7 shown, the terminal constructs the audio feature extraction model to be trained based on the vector quantization variational autoencoder, and constructs an audio codebook with a dimension of [N, D] in the audio feature extraction model to be trained for sharing by the audio encoding model and the audio decoding model in the audio feature extraction model to be trained. Among them, the audio codebook can be regarded as a discrete latent space, N represents the size of the discrete latent space, that is, the size of the audio codebook, and D represents the dimension of each latent embedding vector in the latent space, that is, the dimension of each codebook feature vector in the audio codebook.
[0144] The terminal converts the sample audio into Mel-spectrum features. Among them, the dimension of the Mel-spectrum features is [T, M], where T represents the number of frames of the audio to be recognized, and M represents the dimension of the Mel-spectrum, and M can be set to 80. The terminal inputs the Mel-spectrum features into the audio encoding model in the audio feature extraction model to be trained. There are 5 residual connection structures in the audio encoding model. The first residual connection structure performs convolution processing, batch normalization (Batchnorm) processing, and activation processing on the Mel-spectrum features to obtain the processed sample features corresponding to the Mel-spectrum features, and fuses the processed sample features with the Mel-spectrum features to obtain fused features; the fused features are used as the input of the next residual connection structure, and continue to perform convolution processing, batch normalization processing, and activation processing, fuse the output of the residual connection structure with the input of this residual connection structure, and use the fused features obtained by the fusion as the input of the next residual connection structure, and finally obtain the sample audio features of the Mel-spectrum features. According to the codebook feature vectors in the audio codebook, perform a nearest neighbor search process on the sample audio feature vectors in the sample audio features to obtain the sample codebook feature vectors corresponding to each sample audio feature vector, and all the sample codebook feature vectors together serve as the sample codebook features corresponding to the sample audio features; splice the sample codebook features with the speaker features, and input the processed spliced features into the audio decoding model in the audio feature extraction model to be trained to obtain the predicted Mel-spectrum as the predicted audio features.
[0145] Calculate the mean square error between the predicted Mel-spectrum of the sample audio and the actual Mel-spectrum features of the sample audio to obtain the minimum mean square error (MSE Loss) loss function; according to the minimum mean square error loss function, perform gradient backpropagation to update the model parameters of the audio feature extraction model to be trained, with a learning rate of 0.001 and an Adam optimizer; when the loss value converges, obtain the pre-trained audio feature extraction model.
[0146] (3) Language identification: The terminal obtains the audio to be recognized, inputs the audio to be recognized into the pre-trained audio feature extraction model to obtain the audio features of the audio to be recognized; obtains the target codebook features corresponding to the audio features of the audio to be recognized from the audio codebook; mark the target codebook feature vector as Enc′(x), which can be calculated in the following way:
[0147] Enc′(x) = e k , where k = argmin j ||Enc(x) - e j ||2
[0148] Among them, Enc′(x) represents the target codebook feature vector corresponding to the x-th frame of audio in the audio to be recognized; Enc(x) represents the audio feature vector corresponding to the x-th frame of audio in the audio to be recognized; ej Denote the j-th codebook feature vector in the audio codebook; Enc(x) - e j Can be expressed as the similarity between the audio feature vector Enc(x) and the j-th codebook feature vector e in the audio codebook j ; e k Denote the k-th codebook feature vector in the audio codebook whose similarity with the audio feature vector Enc(x) satisfies the preset similarity condition. The terminal generates a histogram of the target codebook feature according to the number of each codebook feature vector in the target codebook feature; performs normalization processing on the histogram of the target codebook feature to obtain the distribution feature vector of the audio to be recognized. By traversing the distribution feature vectors of each sample audio, query and obtain 16 distribution feature vectors of the sample audio with the closest distance to the distribution feature vector of the audio to be recognized as the target distribution feature vector; take the language category with the largest number of language categories among the language categories of these 16 target distribution feature vectors as the language category of the audio to be recognized.
[0149] In this embodiment, (1) it is possible to perform unsupervised learning based on a large amount of unlabeled speech data and singing data to obtain a pre-trained audio feature extraction model independent of language; (2) for newly added languages, it is also possible to dynamically increase the recognition ability of the pre-trained audio feature extraction model for the audio to be recognized in the newly added languages without retraining the pre-trained audio feature extraction model. Applied in scenarios such as speech recognition, singing recognition, and work classification, it is beneficial to improve the accuracy of language recognition for audio data.
[0150] It should be understood that although each step in the flowcharts involved in the above-described embodiments is shown in sequence according to the indication of the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages, and these steps or stages do not necessarily need to be executed at the same time, but can be executed at different times, and the execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0151] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 8As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a language recognition method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0152] Those skilled in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0153] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.
[0154] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0155] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0156] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0157] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0158] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0159] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A language identification method, characterized in that, The method includes: Inputting the audio to be recognized into a pre-trained audio feature extraction model to obtain the audio features of the audio to be recognized; the pre-trained audio feature extraction model is trained by sample audios in different languages; Obtaining, from the audio codebook included in the pre-trained audio feature extraction model, a target codebook feature corresponding to the audio features of the audio to be recognized; the audio codebook includes codebook feature vectors in different languages; the target codebook feature is used to represent the feature information composed of multiple codebook feature vectors in the audio codebook whose similarity with the audio features of the audio to be recognized satisfies a preset similarity condition; Obtaining a distribution feature vector of the audio to be recognized according to the distribution of each codebook feature vector in the target codebook feature; the distribution feature vector is used to represent the data obtained after processing the frequency distribution of the codebook feature vectors in the target codebook feature; the codebook feature vector is used to represent the latent variables of audios in different languages; Determining the language category corresponding to the target distribution feature vector whose distance from the distribution feature vector of the audio to be recognized in the distribution feature vectors of the sample audios satisfies a preset distance condition as the language category of the audio to be recognized; the preset distance condition is used to represent the determination condition set for the distance between the distribution feature vector of the sample audio and the distribution feature vector of the audio to be recognized.
2. The method according to claim 1, characterized in that, The pre-trained audio feature extraction model is trained in the following manner: Inputting sample audios in different languages into an audio encoding model and a speaker encoding model in the audio feature extraction model to be trained respectively to obtain sample audio features and speaker features corresponding to the sample audios; Obtaining, from the audio codebook included in the audio feature extraction model to be trained, a sample codebook feature corresponding to the sample audio features; Concatenating the sample codebook feature and the speaker feature, and inputting the concatenated feature into an audio decoding model in the audio feature extraction model to be trained to obtain predicted audio features of the sample audio; Iteratively training the audio feature extraction model to be trained according to the difference between the predicted audio features of the sample audio and the actual audio features of the sample audio to obtain the pre-trained audio feature extraction model.
3. The method according to claim 2, characterized in that, The sample audio features are processed in the following manner: Through the audio encoding model in the audio feature extraction model to be trained, performing convolution processing, batch normalization processing, and activation processing on the initial sample features of the sample audio in sequence to obtain the processed sample features of the sample audio; Performing convolution processing, batch normalization processing, and activation processing on the fusion features obtained after fusing the processed sample features of the sample audio and the initial sample features of the sample audio in sequence to obtain the encoded features of the sample audio; Performing dimensionality reduction processing on the encoded features of the sample audio to obtain the dimensionality-reduced features of the sample audio; Inputting the dimensionality-reduced features of the sample audio into a gated recurrent network to obtain the sample audio features of the sample audio.
4. The method according to claim 2, characterized in that The speaker feature is obtained through the following processing: Through the speaker encoding model in the audio feature extraction model to be trained, the initial sample features of the sample audio are sequentially subjected to convolution processing, batch normalization processing, and activation processing to obtain the processed features of the sample audio; The fused features obtained by fusing the processed sample features of the sample audio and the initial sample features of the sample audio are sequentially subjected to convolution processing, batch normalization processing, and activation processing to obtain the encoded features of the sample audio; The encoded features of the sample audio are subjected to dimensionality reduction processing to obtain the dimensionality-reduced features of the sample audio; The dimensionality-reduced features of the sample audio are subjected to mean processing to obtain the speaker feature corresponding to the sample audio.
5. The method according to claim 2, characterized in that, Before inputting the sample audio in different languages into the audio encoding model and the speaker encoding model in the audio feature extraction model to be trained to obtain the sample audio features and speaker features corresponding to the sample audio, it further includes: Obtaining initial audio in different languages; Performing voice activation processing on the initial audio of each language to obtain the valid audio in the initial audio of each language; According to the duration of the valid audio of each language, the valid audio of each language is respectively subjected to variable speed processing and / or pitch shifting processing to obtain the sample audio.
6. The method according to claim 1, wherein The obtaining the distribution feature vector of the audio to be recognized according to the distribution of each codebook feature vector in the target codebook feature includes: According to the quantity distribution of the codebook feature vectors in the target codebook feature, obtaining the histogram of the target codebook feature; Normalizing the histogram of the target codebook feature to obtain the distribution feature vector of the audio to be recognized.
7. The method according to claim 1, characterized in that, The determining the language category of the audio to be recognized as the language category corresponding to the target distribution feature vector whose distance from the distribution feature vector of the sample audio satisfies the preset distance condition includes: Screening out a preset number of target distribution feature vectors whose distance from the distribution feature vector of the audio to be recognized satisfies the preset first distance condition from the distribution feature vectors of each sample audio; Screening out the language category with the largest number of language categories from the language categories corresponding to the target distribution feature vectors as the language category of the audio to be recognized.
8. The method according to claim 1, wherein The obtaining the target codebook feature corresponding to the audio feature of the audio to be recognized from the audio codebook in the pre-trained audio feature extraction model includes: Screening out the target codebook feature vectors corresponding to each audio feature vector in the audio feature of the audio to be recognized from the codebook feature vectors in the audio codebook; Combining each target codebook feature vector into the target codebook feature corresponding to the audio feature of the audio to be recognized.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 8.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Language feature extraction model training method and device, equipment and storage medium
CN113160795A
Language recognition method and device, equipment and storage medium
CN113327584A