Training method, recognition method and device for Hausa voiceprint recognition model

Through the transfer learning method, the Hausa voiceprint recognition model is initially trained based on the English audio samples, and then the Hausa audio samples are trained again, which solves the problem of low training accuracy of the Hausa voiceprint recognition model and improves the recognition accuracy.

CN114694634BActive Publication Date: 2025-06-27DMAI (GUANGZHOU) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011564085.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-25
Publication Date
2025-06-27
Estimated Expiration
2040-12-25

AI Technical Summary

Technical Problem

As a small language, Hausa language can collect less sample data, resulting in lower training accuracy of the voiceprint recognition model, which in turn affects the accuracy of the recognition.

Method used

Through the transfer learning method, the Hausa voiceprint recognition model is first trained based on the English audio samples, and the initial parameters are determined, and the initial model is then trained again based on the Hausa audio samples, and the parameters are adjusted to improve the accuracy of the model.

Benefits of technology

Through this method, the problem of insufficient audio samples of Hausa language can be effectively avoided, the accuracy of the voiceprint recognition model can be improved, and the accuracy of the recognition can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114694634B_ABST
    Figure CN114694634B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of voiceprint recognition, and specifically to a training method, recognition method and device for a Hausa voiceprint recognition model. The training method includes obtaining a first frequency domain feature and a first voiceprint feature of an English audio sample, as well as a second frequency domain feature and a second voiceprint feature of a Hausa audio sample; training a Hausa voiceprint recognition model based on the first frequency domain feature and the first voiceprint feature to determine initial parameters of the Hausa voiceprint recognition model, and obtaining an initial Hausa voiceprint recognition model; training the initial Hausa voiceprint recognition model based on the second frequency domain feature and the second voiceprint feature, adjusting the initial parameters of the initial Hausa voiceprint recognition model, and determining a target Hausa voiceprint recognition model. By means of transfer learning, the problem of insufficient Hausa audio samples can be avoided, and the accuracy of the trained Hausa voiceprint recognition model can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voiceprint recognition, and particularly to a training method, a recognition method and a device for a Hausa voiceprint recognition model. Background Art

[0002] Speech recognition is the process of converting human voice signals into text, and is one of the important technologies in the field of artificial intelligence perception. With the development of deep learning technology, both the accuracy and speed of speech recognition have made long-term progress. Nowadays, speech recognition technology has penetrated into many applications in our daily life, such as smart speakers, shopping guide robots and other products. However, most of the existing speech recognition researches only focus on the languages with the largest number of users, such as English and Chinese, which results in the application of speech recognition being limited to relatively developed regions and cities.

[0003] There are 6,809 languages in the world, and most of them are minority languages with a small number of users. The research on speech recognition for minority languages is the key bridge to narrow the communication gap between people speaking different languages. Among them, Hausa belongs to the Chadic branch of the Afro-Asiatic language family and is one of the three most important languages in Africa. For a voiceprint recognition model, its training generally requires thousands of hours of audio. As a minority language, there is less sample data that can be collected for Hausa. Due to the lack of sample data, the accuracy of the trained voiceprint recognition model is relatively low, and thus the accuracy of voiceprint recognition is relatively low. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a training method, a recognition method and a device for a Hausa voiceprint recognition model to solve the problem of relatively low accuracy of voiceprint recognition.

[0005] According to a first aspect, an embodiment of the present invention provides a training method for a Hausa voiceprint recognition model, including:

[0006] Obtaining a first frequency domain feature and a first voiceprint feature of an English audio sample, and a second frequency domain feature and a second voiceprint feature of a Hausa audio sample;

[0007] Training a Hausa voiceprint recognition model based on the first frequency domain feature and the first voiceprint feature to determine initial parameters of the Hausa voiceprint recognition model, and obtaining an initial Hausa voiceprint recognition model;

[0008] Training the initial Hausa voiceprint recognition model based on the second frequency domain feature and the second voiceprint feature, adjusting the initial parameters of the initial Hausa voiceprint recognition model, and determining a target Hausa voiceprint recognition model, wherein the output of the target Hausa voiceprint recognition model is the voiceprint feature of a speaker.

[0009] The training method of the Hausa voiceprint recognition model provided by the embodiments of the present invention. Since the number of Hausa audio samples is small and Hausa is relatively similar to English, when training the Hausa voiceprint recognition model, first train the Hausa voiceprint recognition model based on English audio samples, and use the obtained parameters as the initial parameters of the Hausa voiceprint recognition model; then retrain the initial Hausa voice model based on Hausa audio samples, and further fine-tune the initial parameters. That is, through the method of transfer learning, the problem of insufficient Hausa audio samples can be avoided, and the accuracy of the trained Hausa voiceprint recognition model can be ensured.

[0010] Combined with the first aspect, in the first implementation manner of the first aspect, training the Hausa voiceprint recognition model based on the first frequency domain feature and the first voiceprint feature to determine the initial parameters of the Hausa voiceprint recognition model and obtain the initial Hausa voiceprint recognition model includes:

[0011] Input the first frequency domain feature into the Hausa voiceprint recognition model to obtain the first predicted voiceprint feature;

[0012] Based on the error between the first voiceprint feature and the first predicted voiceprint feature, adjust the parameters of the Hausa voiceprint recognition model to determine the initial Hausa voiceprint recognition model.

[0013] The training method of the Hausa voiceprint recognition model provided by the embodiments of the present invention processes the first frequency domain feature corresponding to the English audio sample through the Hausa voiceprint recognition model to obtain the first predicted voiceprint feature, and then compares the error between the voiceprint feature predicted by the network and the voiceprint feature corresponding to the audio sample to adjust the model parameters, which can ensure the accuracy of the initial parameters of the determined initial Hausa voiceprint recognition model.

[0014] Combined with the first implementation manner of the first aspect, in the second implementation manner of the first aspect, the inputting the first frequency domain feature into the Hausa voiceprint recognition model to obtain the first predicted voiceprint feature includes:

[0015] Use the first network model in the Hausa voiceprint recognition model to process the first frequency domain feature to obtain frame-level speaker information;

[0016] Use the second network model in the Hausa voiceprint recognition model to cluster the frame-level speaker information to obtain sentence-level speaker information, and determine the first predicted voiceprint feature.

[0017] The training method of the Hausa voiceprint recognition model provided by the embodiment of the present invention sets two network models in the Hausa voiceprint recognition model. First, the speaker information at the frame level is obtained, and then the second network model is used to perform clustering analysis on the output of the first network model to determine the first predicted voiceprint feature; that is, the first predicted voiceprint feature is obtained through clustering, which can ensure the accuracy of the first predicted voiceprint feature and improve the efficiency of model training.

[0018] Combined with the first implementation manner of the first aspect, in the third implementation manner of the first aspect, adjusting the parameters of the Hausa voiceprint recognition model based on the error between the first voiceprint feature and the first predicted voiceprint feature to determine the initial parameters of the initial Hausa voiceprint recognition model includes:

[0019] Calculating a loss function using the first voiceprint feature and the first predicted voiceprint feature;

[0020] Based on the calculation result of the loss function, adjusting the parameters of the Hausa voiceprint recognition model to determine the initial parameters of the initial Hausa voiceprint recognition model.

[0021] Combined with the first aspect, in the fourth implementation manner of the first aspect, training the initial Hausa voiceprint recognition model based on the Hausa audio sample and the second voiceprint feature, and adjusting the initial parameters of the initial Hausa voiceprint recognition model to determine the target Hausa voiceprint recognition model includes:

[0022] Inputting the second frequency domain feature into the initial Hausa voiceprint recognition model to obtain a second predicted voiceprint feature;

[0023] Based on the error between the second voiceprint feature and the second predicted voiceprint feature, adjusting the initial parameters of the initial Hausa voiceprint recognition model to determine the target Hausa voiceprint recognition model.

[0024] The training method of the Hausa voiceprint recognition model provided by the embodiment of the present invention, on the basis of determining the initial parameters, further fine-tunes the initial parameters using the Hausa audio sample. On the one hand, it can ensure the accuracy of the target Hausa voiceprint recognition model, and on the other hand, it can improve the efficiency of model training.

[0025] Combined with the first aspect, in the fifth implementation manner of the first aspect, obtaining the first frequency domain feature of the English audio sample and obtaining the second frequency domain feature of the Hausa audio sample includes:

[0026] Dividing the English audio sample and the Hausa audio sample into silent segments and non-silent segments respectively;

[0027] Perform Fourier transform processing on the English audio sample of the non-silent segment and the Hausa audio sample of the non-silent segment respectively to obtain the first frequency domain feature and the second frequency domain feature.

[0028] Before processing the frequency domain features, the training method of the Hausa voiceprint recognition model provided by the embodiments of the present invention first removes the silent segments in the audio samples, which can reduce the amount of data processing and improve the training efficiency.

[0029] Combined with the first aspect, or any one of the first to fifth implementation manners of the first aspect, in the sixth implementation manner of the first aspect, it further includes:

[0030] Obtain intra-class data and inter-class data, where the intra-class data is the audio data of the same speaker, and the inter-class data is the audio data of different speakers;

[0031] Extract the frequency domain features of the intra-class data and the inter-class data;

[0032] Input the extracted frequency domain features into the target Hausa voiceprint recognition model to determine the voiceprint features corresponding to each intra-class data and the voiceprint features corresponding to each inter-class data;

[0033] Based on the similarity of the voiceprint features corresponding to each intra-class data and the similarity of the voiceprint features corresponding to each inter-class data, determine the voiceprint recognition threshold.

[0034] After the target Hausa voiceprint recognition model is determined, the training method of the Hausa voiceprint recognition model provided by the embodiments of the present invention determines the voiceprint recognition threshold by using a large amount of intra-class data and inter-class data, which can ensure the accuracy of the determined voiceprint recognition threshold and improve the accuracy of subsequent voiceprint recognition using this model.

[0035] According to the second aspect, the embodiments of the present invention further provide a Hausa voiceprint recognition method, including:

[0036] Obtain the audio to be recognized;

[0037] Extract the frequency domain features of the audio to be recognized;

[0038] Input the extracted frequency domain features into the target Hausa voiceprint recognition model to obtain the target voiceprint feature, where the target Hausa voiceprint recognition model is trained according to the training method of the Hausa voiceprint recognition model described in the first aspect of the present invention or any one of the implementation manners of the first aspect;

[0039] Based on the target voiceprint feature, the voiceprint features to be matched in the voiceprint feature library, and the voiceprint recognition threshold, determine the speaker corresponding to the audio to be recognized.

[0040] The Hausa voiceprint recognition method provided by the embodiment of the present invention can ensure the accuracy of recognition by recognizing the audio to be recognized based on an accurate target voiceprint recognition model.

[0041] According to a third aspect, the embodiment of the present invention further provides a training device for a Hausa voiceprint recognition model, including:

[0042] A first acquisition module, configured to acquire a first frequency-domain feature and a first voiceprint feature of an English audio sample, as well as a second frequency-domain feature and a second voiceprint feature of a Hausa audio sample;

[0043] A first extraction module, configured to extract the first frequency-domain feature of the English audio sample and the second frequency-domain feature of the Hausa audio sample;

[0044] A first training module, configured to train a Hausa voiceprint recognition model based on the first frequency-domain feature and the first voiceprint feature, determine initial parameters of the Hausa voiceprint recognition model, and obtain an initial Hausa voiceprint recognition model;

[0045] A second training module, configured to train the initial Hausa voiceprint recognition model based on the second frequency-domain feature and the second voiceprint feature, adjust the initial parameters of the initial Hausa voiceprint recognition model, and determine a target Hausa voiceprint recognition model, where the output of the target Hausa voiceprint recognition model is the voiceprint feature of the speaker.

[0046] For the training device of the Hausa voiceprint recognition model provided by the embodiment of the present invention, since the number of Hausa audio samples is small and Hausa and English are relatively similar, when training the Hausa voiceprint recognition model, first train the Hausa voiceprint recognition model based on English audio samples, and use the trained parameters as the initial parameters of the Hausa voiceprint recognition model; then retrain the initial Hausa voice model based on Hausa audio samples, and further fine-tune the initial parameters. That is, through the method of transfer learning, the problem of insufficient Hausa audio samples can be avoided, and the accuracy of the trained Hausa voiceprint recognition model can be ensured.

[0047] According to a fourth aspect, the embodiment of the present invention further provides a Hausa voiceprint recognition device, including:

[0048] A second acquisition module, configured to acquire the audio to be recognized;

[0049] A second extraction module, configured to extract the frequency-domain feature of the audio to be recognized;

[0050] An identification module, configured to input the extracted frequency-domain features into a target Hausa voiceprint recognition model to obtain target voiceprint features, where the target Hausa voiceprint recognition model is trained according to the training method of the Hausa voiceprint recognition model described in the first aspect of the present invention or any implementation manner of the first aspect;

[0051] A determination module, configured to determine the speaker corresponding to the audio to be recognized based on the target voiceprint features, the voiceprint features to be matched in the voiceprint feature library, and the voiceprint recognition threshold.

[0052] The Hausa voiceprint recognition device provided by the embodiments of the present invention can ensure the accuracy of recognition by recognizing the audio to be recognized based on an accurate target voiceprint recognition model.

[0053] According to a fifth aspect, an embodiment of the present invention provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the training method of the Hausa voiceprint recognition model described in the first aspect or any implementation manner of the first aspect, or execute the Hausa voiceprint recognition method described in the second aspect.

[0054] According to a sixth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores computer instructions for causing the computer to execute the training method of the Hausa voiceprint recognition model described in the first aspect or any implementation manner of the first aspect, or execute the Hausa voiceprint recognition method described in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1 It is a flowchart of the training method of the Hausa voiceprint recognition model according to an embodiment of the present invention;

[0057] Figure 2 It is a flowchart of the training method of the Hausa voiceprint recognition model according to an embodiment of the present invention;

[0058] Figure 3 It is a flowchart of the training method of the Hausa voiceprint recognition model according to an embodiment of the present invention;

[0059] Figure 4 is a flowchart of a Hausa voiceprint recognition method according to an embodiment of the present invention;

[0060] Figure 5 is a structural block diagram of a training device for a Hausa voiceprint recognition model according to an embodiment of the present invention;

[0061] Figure 6 is a structural block diagram of a training device for a Hausa voiceprint recognition model according to an embodiment of the present invention;

[0062] Figure 7 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0064] According to an embodiment of the present invention, an embodiment of a training method for a Hausa voiceprint recognition model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0065] In this embodiment, a training method for a Hausa voiceprint recognition model is provided, which can be used in an electronic device such as a computer, a mobile terminal, etc. Figure 1 is a flowchart of a training method for a Hausa voiceprint recognition model according to an embodiment of the present invention, as Figure 1 shown, the process includes the following steps:

[0066] S11, obtain the first frequency domain feature and the first voiceprint feature of the English audio sample, and the second frequency domain feature and the second voiceprint feature of the Hausa audio sample.

[0067] Since Hausa is relatively similar to English, the Hausa voiceprint recognition model can be trained using English audio data; and then its parameters can be fine-tuned using Hausa audio data.

[0068] The first frequency-domain feature and the first voiceprint feature of the English audio sample, as well as the second frequency-domain feature and the second voiceprint feature of the Hausa audio sample, can be obtained by the electronic device from the outside world or stored in the electronic device, etc. There is no restriction on the way the electronic device obtains the above features here.

[0069] For example, there are databases of two languages stored in the electronic device, namely the English database and the Hausa database. Among them, the English database includes the first frequency-domain feature and the first voiceprint feature of each English audio sample; the Hausa database includes the second frequency-domain feature and the second voiceprint feature of each Hausa audio sample. It can also be considered that each English audio sample in the English database corresponds to the first frequency-domain feature and the first voiceprint feature; each Hausa audio sample in the Hausa database corresponds to the second frequency-domain feature and the second voiceprint feature.

[0070] S12. Train the Hausa voiceprint recognition model based on the first frequency-domain feature and the first voiceprint feature, determine the initial parameters of the Hausa voiceprint recognition model, and obtain the initial Hausa voiceprint recognition model.

[0071] When the electronic device trains the Hausa voiceprint model, it can first set each parameter in the Hausa voiceprint model according to the empirical value; then use the first time-domain feature and the first voiceprint feature to train the Hausa voiceprint recognition model, determine the initial parameters of the Hausa voiceprint recognition model, and obtain the initial Hausa voiceprint recognition model.

[0072] The input of the Hausa voiceprint recognition model is the frequency-domain feature, and the output is the voiceprint feature. That is, after the electronic device inputs the first frequency-domain feature into the Hausa voiceprint recognition model, the predicted voiceprint feature is output. The electronic device then uses the predicted voiceprint feature and the first voiceprint feature to update the model parameters. After multiple trainings and corresponding parameter updates, the initial parameters can be determined and the initial Hausa voiceprint recognition model can be obtained.

[0073] There is no restriction on the specific model structure of the Hausa voiceprint recognition model here. As long as its input is the frequency-domain feature and the output is the voiceprint feature, the specific model structure can be set accordingly according to the actual situation.

[0074] S13. Train the initial Hausa voiceprint recognition model based on the second frequency-domain feature and the second voiceprint feature, adjust the initial parameters of the initial Hausa voiceprint recognition model, and determine the target Hausa voiceprint recognition model.

[0075] Among them, the output of the target Hausa voiceprint recognition model is the voiceprint feature of the speaker.

[0076] Similarly to S12 above, after obtaining the initial Hausa voiceprint recognition model in S12 above, the electronic device then uses the second frequency domain feature and the second voiceprint feature corresponding to the Hausa audio sample to train the initial Hausa voiceprint recognition model, adjusts the initial parameters in the initial Hausa voiceprint recognition model, and finally determines the target Hausa voiceprint recognition model.

[0077] Specifically, the Hausa voiceprint recognition model is trained in batches based on the English audio sample and the Hausa audio sample. First, the first training is performed based on the English audio sample to obtain the initial Hausa voiceprint recognition model; secondly, the second training is performed based on the Hausa audio sample on the basis of the initial Hausa voiceprint recognition model to obtain the target Hausa voiceprint recognition model.

[0078] In the training method of the Hausa voiceprint recognition model provided in this embodiment, since the number of Hausa audio samples is small and Hausa and English are relatively similar, when training the Hausa voiceprint recognition model, first train the Hausa voiceprint recognition model based on the English audio sample, and use the trained parameters as the initial parameters of the Hausa voiceprint recognition model; then train the initial Hausa voice model again based on the Hausa audio sample, and then fine-tune the initial parameters. That is, through the method of transfer learning, the problem of insufficient Hausa audio samples can be avoided, and the accuracy of the trained Hausa voiceprint recognition model can be ensured.

[0079] In this embodiment, a training method of a Hausa voiceprint recognition model is provided, which can be used in electronic devices such as computers and mobile terminals. Figure 2 It is a flowchart of the training method of the Hausa voiceprint recognition model according to the embodiment of the present invention, as Figure 2 shown, and the process includes the following steps:

[0080] S21, obtain the first frequency domain feature and the first voiceprint feature of the English audio sample, and the second frequency domain feature and the second voiceprint feature of the Hausa audio sample.

[0081] For details, please refer to Figure 1 S11 of the embodiment shown, which will not be elaborated here.

[0082] S22, train the Hausa voiceprint recognition model based on the first frequency domain feature and the first voiceprint feature, determine the initial parameters of the Hausa voiceprint recognition model, and obtain the initial Hausa voiceprint recognition model.

[0083] Wherein, the Hausa voiceprint recognition model includes a first network model and a second network model, and the input of the second network model is connected to the output of the first network model. The second network model is used to cluster the output of the first network model to obtain the predicted voiceprint feature.

[0084] Specifically, the above S22 may include the following steps:

[0085] S221: Input the first frequency-domain feature into the Hausa voiceprint recognition model to obtain the first predicted voiceprint feature.

[0086] The electronic device inputs the first frequency-domain feature corresponding to the English audio sample into the Hausa voiceprint recognition model and outputs the first predicted voiceprint feature.

[0087] As an optional implementation manner of this embodiment, the above S221 may include the following steps:

[0088] (1) Process the first frequency-domain feature using the first network model in the Hausa voiceprint recognition model to obtain speaker information at the frame level.

[0089] The first network model may be a mobilenet network module. The electronic device outputs the first frequency-domain feature into the mobilenet network module to obtain speaker information at the frame level.

[0090] For example, the input of the first network model is 200 frames and the output is 20 frames. When using the first network to process 200 frames of data, a sliding window can be used to extract features from the 200 frames of data and output speaker information at the frame level.

[0091] It should be noted that the first network model is not limited to the above mobilenet network module and may also be other network modules. For example, networks such as mobilenetv1, mobilenetv3, resnet, and vgg can be replaced. Among them, the mobilenetv2 network is mainly used because it has more stable performance than the other two mobilenet networks; it has better performance than vgg and is faster to train and smaller in size than the resnet network with basically no significant loss in performance.

[0092] There is no specific limitation on which network module is used as the first network model, and specific settings can be made according to the actual situation.

[0093] (2) Cluster the speaker information at the frame level using the second network model in the Hausa voiceprint recognition model to obtain speaker information at the sentence level and determine the first predicted voiceprint feature.

[0094] The second network model may be a glvad network module. The electronic device inputs the speaker information at the frame level output by the first network model into the second network model, clusters the speaker information at the frame level using the second network model to obtain speaker information at the sentence level, and determines the speaker information at the sentence level as the first predicted voiceprint feature.

[0095] Among them, the second network model mentioned above is not limited to the glvad network module described above, and can also be amsoftmax, ge2e, etc. No limitation is made here, as long as it is ensured that the second network model can perform clustering processing on the output of the first network model.

[0096] S222. Based on the error between the first voiceprint feature and the first predicted voiceprint feature, adjust the parameters of the Hausa voiceprint recognition model to determine the initial Hausa voiceprint recognition model.

[0097] After the electronic device uses the Hausa voiceprint recognition model to predict the first predicted voiceprint feature, it compares the error between the predicted first predicted voiceprint feature and the first voiceprint feature corresponding to the English audio sample, so as to adjust the parameters of the Hausa voiceprint recognition model and determine the initial Hausa voiceprint recognition model.

[0098] In an alternative implementation manner of this embodiment, the above S222 may include the following steps:

[0099] (1) Calculate the loss function using the first voiceprint feature and the first predicted voiceprint feature.

[0100] The electronic device calculates the loss function using the first voiceprint feature and the first predicted voiceprint feature to obtain the corresponding loss function value.

[0101] (2) Based on the calculation result of the loss function, adjust the parameters of the Hausa voiceprint recognition model to determine the initial parameters of the initial Hausa voiceprint recognition model.

[0102] After the electronic device calculates the loss function value, it adjusts the parameters of the Hausa voiceprint recognition model to determine the initial parameters, and accordingly, the initial Hausa voiceprint recognition model can be determined.

[0103] S23. Train the initial Hausa voiceprint recognition model based on the second frequency domain feature and the second voiceprint feature, adjust the initial parameters of the initial Hausa voiceprint recognition model, and determine the target Hausa voiceprint recognition model.

[0104] Among them, the output of the target Hausa voiceprint recognition model is the voiceprint feature of the speaker.

[0105] Specifically, the above S23 may include the following steps:

[0106] S231. Input the second frequency domain feature into the initial Hausa voiceprint recognition model to obtain the second predicted voiceprint feature.

[0107] The initial Hausa voiceprint recognition model obtained by the electronic device in S22 above is obtained by training the Hausa voiceprint recognition model using a large number of English audio samples. Then, in the way of transfer learning, the parameters of the initial Hausa voiceprint recognition model are fine-tuned using Hausa audio samples. Among them, the parameter fine-tuning here can be to learn and adjust the parameters of the first network model and the second network model respectively by setting different learning rates.

[0108] Among them, the specific way for the electronic device to input the second frequency domain feature corresponding to the Hausa audio sample into the initial Hausa voiceprint recognition model to obtain the second predicted voiceprint feature is similar to the implementation manner of obtaining the first predicted voiceprint feature in S22 above. For specific details, please refer to the detailed description of S22 above, which will not be elaborated here.

[0109] S232. Based on the error between the second voiceprint feature and the second predicted voiceprint feature, adjust the initial parameters of the initial Hausa voiceprint recognition model to determine the target Hausa voiceprint recognition model.

[0110] After the electronic device obtains the second predicted voiceprint feature, the way to adjust the initial parameters using the second voiceprint feature and the error of the second predicted voiceprint feature is similar to the way to determine the initial parameters in S22 above. For detailed information, please refer to the relevant description of S22 above, which will not be elaborated here.

[0111] The training method of the Hausa voiceprint recognition model provided in this embodiment processes the first frequency domain feature corresponding to the English audio sample through the Hausa voiceprint recognition model to obtain the first predicted voiceprint feature, and then compares the error between the voiceprint feature predicted by the network and the voiceprint feature corresponding to the audio sample to adjust the model parameters, which can ensure the accuracy of the initial parameters of the determined initial Hausa voiceprint recognition model; in addition, on the basis of determining the initial parameters, the initial parameters are further fine-tuned using Hausa audio samples, which can ensure the accuracy of the target Hausa voiceprint recognition model on the one hand and improve the efficiency of model training on the other hand.

[0112] In this embodiment, a training method of a Hausa voiceprint recognition model is provided, which can be used in electronic devices such as computers and mobile terminals. Figure 3 It is a flowchart of the training method of the Hausa voiceprint recognition model according to the embodiment of the present invention, as Figure 3 shown, and the process includes the following steps:

[0113] S31. Obtain the first frequency domain feature and the first voiceprint feature of the English audio sample, and the second frequency domain feature and the second voiceprint feature of the Hausa audio sample.

[0114] Specifically, the above S31 may include the following steps:

[0115] S311. Divide the English audio samples and Hausa audio samples into silent segments and non-silent segments respectively.

[0116] Among them, the processing methods of the frequency domain features of the English audio samples and Hausa audio samples are the same. The processing of the English audio samples will be described in detail below as an example.

[0117] After the electronic device obtains the English audio samples, after processing such as denoising, endpoint detection, and feature extraction of the English audio samples, the corresponding first frequency domain features are obtained. Specifically, the electronic device uses the logmmse algorithm to enhance the denoising of the English audio samples, and then uses the webrtc_vad technology to divide the enhanced speech into silent segments and non-silent segments.

[0118] S312. Perform Fourier transform processing on the non-silent segments of the English audio samples and the non-silent segments of the Hausa audio samples respectively to obtain the first frequency domain features and the second frequency domain features.

[0119] The electronic device removes the audio data of the silent segments and only processes the audio data of the non-silent segments. Specifically, the electronic device performs Fourier transform and normalization processing on the non-silent segments of the English audio samples to obtain the corresponding first frequency domain features.

[0120] S32. Train the Hausa voiceprint recognition model based on the first frequency domain features and the first voiceprint features, determine the initial parameters of the Hausa voiceprint recognition model, and obtain the initial Hausa voiceprint recognition model.

[0121] For details, please refer to Figure 2 S22 of the illustrated embodiment, which will not be elaborated here.

[0122] S33. Train the initial Hausa voiceprint recognition model based on the second frequency domain features and the second voiceprint features, adjust the initial parameters of the initial Hausa voiceprint recognition model, and determine the target Hausa voiceprint recognition model.

[0123] Among them, the output of the target Hausa voiceprint recognition model is the voiceprint feature of the speaker.

[0124] For details, please refer to Figure 2 S23 of the illustrated embodiment, which will not be elaborated here.

[0125] S34. Obtain intra-class data and inter-class data.

[0126] Among them, the intra-class data is the audio data of the same speaker, and the inter-class data is the audio data of different speakers.

[0127] An electronic device can obtain audio data belonging to different speakers. After obtaining the audio data of different speakers, the obtained audio data is divided according to the speakers to obtain the audio data belonging to each speaker.

[0128] For example, for speaker 1, there are N1 audio data;

[0129] For speaker 2, there are N2 audio data;

[0130] For speaker m, there are Nm audio data.

[0131] Furthermore, the N1 audio data of speaker 1 can be called intra-class data; the N1 audio data of speaker 1 and the N2 audio data of speaker 2 can be called inter-class data.

[0132] S35, Extract the frequency domain features of the intra-class data and the inter-class data.

[0133] After the electronic device obtains the intra-class data and the inter-class data, it extracts the frequency domain features of each audio data respectively. Among them, the extraction method of the frequency domain features can refer to the description in S31 above and will not be elaborated here.

[0134] S36, Input the extracted frequency domain features into the target Hausa voiceprint recognition model to determine the voiceprint features corresponding to each intra-class data and the voiceprint features corresponding to each inter-class data.

[0135] The electronic device inputs the extracted frequency domain features into the target Hausa voiceprint recognition model to obtain the voiceprint features corresponding to the intra-class data and the voiceprint features corresponding to the inter-class data.

[0136] S37, Determine the voiceprint recognition threshold based on the similarity of the voiceprint features corresponding to each intra-class data and the similarity of the voiceprint features corresponding to each inter-class data.

[0137] After obtaining the voiceprint features of the intra-class data and the voiceprint features of the inter-class data, calculate the similarity of the Hausa voiceprint features of the intra-class and the inter-class. Among them, the similarity between different sentences of the same speaker is as high as possible, and the similarity between different speakers is as low as possible.

[0138] After the electronic device calculates the similarity of the voiceprint features corresponding to each intra-class data and the similarity of the voiceprint features corresponding to each inter-class data, it can determine the voiceprint recognition threshold based on the above goals of the similarity of the intra-class data and the inter-class data.

[0139] For example, the similarity value at the minimum average error probability can be taken as the threshold for speaker recognition. When the rates are equal, the common value is called the equal error rate. This value indicates that the proportion of false acceptances is equal to the proportion of false rejections. The lower the equal error rate value, the higher the accuracy of the biometric system. The electronic device can set different equal error rates according to actual needs to obtain voiceprint recognition thresholds that meet different requirements.

[0140] In the training method of the Hausa voiceprint recognition model provided in this embodiment, before processing the frequency domain features, the silent segments in the audio samples are removed first, which can reduce the amount of data processing and improve the training efficiency. In addition, after the target Hausa voiceprint recognition model is determined, when determining the voiceprint recognition threshold using a large amount of within-class data and between-class data, the accuracy of the determined voiceprint recognition threshold can be guaranteed, and the accuracy of subsequent voiceprint recognition using this model can be improved.

[0141] According to an embodiment of the present invention, there is provided an embodiment of a Hausa voiceprint recognition method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0142] In this embodiment, a Hausa voiceprint recognition method is provided, which can be used in electronic devices such as computers and mobile terminals. Figure 4 is a flowchart of the Hausa voiceprint recognition method according to an embodiment of the present invention, as Figure 4 shown, the process includes the following steps:

[0143] S41, Obtain the audio to be recognized.

[0144] The audio to be recognized can be collected by the electronic device in real time, or obtained by the electronic device from the outside world, etc. No limitation is made on the acquisition method of the audio to be recognized here.

[0145] S42, Extract the frequency domain features of the audio to be recognized.

[0146] After the electronic device obtains the audio to be recognized, it can use Figure 3 the method of S31 in the shown embodiment to extract the frequency domain features of the audio to be recognized.

[0147] S43, Input the extracted frequency domain features into the target Hausa voiceprint recognition model to obtain the target voiceprint features.

[0148] Among them, the target Hausa voiceprint recognition model is trained according to the training method of the Hausa voiceprint recognition model described in the above embodiment.

[0149] The electronic device inputs the frequency-domain features into the target Hausa voiceprint recognition model to obtain the target voiceprint features. Regarding the target Hausa voiceprint recognition model, please refer to the relevant descriptions in the above embodiments and will not be elaborated here.

[0150] In some alternative embodiments of this embodiment, after the electronic device obtains the target voiceprint features, it can segment them to obtain multiple target voiceprint sub-features; then perform equalization and smoothing processing on the multiple target voiceprint sub-features to obtain the processed voiceprint features. Subsequently, using the processed voiceprint features for speaker recognition can improve the short speech recognition effect, making the short speech effect approximate that of long speech.

[0151] S44. Based on the target voiceprint features, the to-be-matched voiceprint features in the voiceprint feature library, and the voiceprint recognition threshold, determine the speaker corresponding to the to-be-recognized audio.

[0152] After the electronic device obtains the target voiceprint features, it can calculate the similarity between the target voiceprint features and the to-be-matched voiceprint features in the voiceprint feature library respectively; after calculating the similarity, compare it with the voiceprint recognition threshold, so as to determine whether the target voiceprint features and the to-be-matched voiceprint features come from the same speaker.

[0153] For example, the to-be-matched voiceprint features of different speakers are stored in the electronic device. When it is determined that the target voiceprint features match the to-be-matched voiceprint features, it can be determined that the to-be-recognized audio matches the speaker in the voiceprint feature library; further, by searching for the speaker information corresponding to the to-be-matched voiceprint features, the speaker information corresponding to the to-be-recognized audio can be determined.

[0154] The Hausa voiceprint recognition method provided in this embodiment can ensure the recognition accuracy by recognizing the to-be-recognized audio based on an accurate target voiceprint recognition model.

[0155] In this embodiment, a training device for a Hausa voiceprint recognition model and a Hausa voiceprint recognition device are also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be elaborated again. As used below, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0156] This embodiment provides a training device for a Hausa voiceprint recognition model, as Figure 5 shown, including:

[0157] The first acquisition module 51 is used to acquire the first frequency-domain features and the first voiceprint features of English audio samples, and the second frequency-domain features and the second voiceprint features of Hausa audio samples;

[0158] A first extraction module 52, configured to extract a first frequency-domain feature of the English audio sample and a second frequency-domain feature of the Hausa audio sample;

[0159] A first training module 53, configured to train a Hausa voiceprint recognition model based on the first frequency-domain feature and the first voiceprint feature, determine initial parameters of the Hausa voiceprint recognition model, and obtain an initial Hausa voiceprint recognition model;

[0160] A second training module 54, configured to train the initial Hausa voiceprint recognition model based on the second frequency-domain feature and the second voiceprint feature, adjust the initial parameters of the initial Hausa voiceprint recognition model, and determine a target Hausa voiceprint recognition model, where an output of the target Hausa voiceprint recognition model is a voiceprint feature of a speaker.

[0161] In the training device for the Hausa voiceprint recognition model provided in this embodiment, since the number of Hausa audio samples is small and Hausa and English are relatively similar, when training the Hausa voiceprint recognition model, the Hausa voiceprint recognition model is first trained based on English audio samples, and the obtained parameters are used as the initial parameters of the Hausa voiceprint recognition model; then the initial Hausa voice model is retrained based on the Hausa audio samples, and thus the initial parameters are finely tuned. That is, the method of transfer learning can not only avoid the problem of insufficient Hausa audio samples, but also ensure the accuracy of the trained Hausa voiceprint recognition model.

[0162] This embodiment provides a Hausa voiceprint recognition device, as Figure 6 shown, including:

[0163] A second acquisition module 61, configured to acquire an audio to be recognized;

[0164] A second extraction module 62, configured to extract a frequency-domain feature of the audio to be recognized;

[0165] A recognition module 63, configured to input the extracted frequency-domain feature into a target Hausa voiceprint recognition model to obtain a target voiceprint feature, where the target Hausa voiceprint recognition model is trained according to the training method of the Hausa voiceprint recognition model in the first aspect of the present invention or any implementation manner of the first aspect;

[0166] A determination module 64, configured to determine a speaker corresponding to the audio to be recognized based on the target voiceprint feature, a voiceprint feature to be matched in a voiceprint feature library, and a voiceprint recognition threshold.

[0167] The Hausa voiceprint recognition device provided in this embodiment can ensure the accuracy of recognition by recognizing the audio to be recognized based on an accurate target voiceprint recognition model.

[0168] The training device of the Hausa voiceprint recognition model or the Hausa voiceprint recognition device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0169] The further function descriptions of the above-mentioned respective modules are the same as those in the corresponding embodiments above, and will not be elaborated here.

[0170] The embodiment of the present invention further provides an electronic device having the above-mentioned Figure 5 shown training device of the Hausa voiceprint recognition model, or Figure 6 shown Hausa voiceprint recognition device.

[0171] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an electronic device provided by an alternative embodiment of the present invention. As Figure 7 shown, the electronic device may include: at least one processor 71, such as a CPU (Central Processing Unit), at least one communication interface 73, a memory 74, and at least one communication bus 72. Among them, the communication bus 72 is used to realize the connection and communication between these components. Among them, the communication interface 73 may include a display screen (Display), a keyboard (Keyboard), and optionally the communication interface 73 may further include a standard wired interface and a wireless interface. The memory 74 may be a high-speed RAM memory (Random Access Memory, volatile random access memory), or a non-volatile memory, such as at least one disk memory. Optionally, the memory 74 may further be at least one storage device located far from the aforementioned processor 71. Among them, the processor 71 may be combined with Figure 5 or Figure 6 the device described, the memory 74 stores an application program, and the processor 71 calls the program code stored in the memory 74 to execute any of the above method steps.

[0172] Among them, the communication bus 72 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus 72 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7It is only represented by a thick line, but it does not mean that there is only one bus or one type of bus.

[0173] Among them, the memory 74 may include a volatile memory (English: volatile memory), such as a random access memory (English: random-access memory, abbreviation: RAM); the memory may also include a non-volatile memory (English: non-volatile memory), such as a flash memory (English: flash memory), a hard disk (English: hard disk drive, abbreviation: HDD) or a solid-state drive (English: solid-state drive, abbreviation: SSD); the memory 74 may also include a combination of the above types of memories.

[0174] Among them, the processor 71 may be a central processing unit (English: central processing unit, abbreviation: CPU), a network processor (English: network processor, abbreviation: NP) or a combination of a CPU and an NP.

[0175] Among them, the processor 71 may further include a hardware chip. The above hardware chip may be an application-specific integrated circuit (English: application-specific integrated circuit, abbreviation: ASIC), a programmable logic device (English: programmable logic device, abbreviation: PLD) or a combination thereof. The above PLD may be a complex programmable logic device (English: complex programmable logic device, abbreviation: CPLD), a field-programmable gate array (English: field-programmable gate array, abbreviation: FPGA), a generic array logic (English: generic array logic, abbreviation: GAL) or any combination thereof.

[0176] Optionally, the memory 74 is also used to store program instructions. The processor 71 may call the program instructions to implement, as shown in the embodiments of the present application Figures 1 to 3 the training method of the Hausa voiceprint recognition model shown in the embodiments, or Figure 4 the Hausa voiceprint recognition method shown in the embodiments.

[0177] An embodiment of the present invention also provides a non-transitory computer storage medium. The computer storage medium stores computer-executable instructions, and these computer-executable instructions can execute the training method of the Hausa voiceprint recognition model or the Hausa voiceprint recognition method in any of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (abbreviation: HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.

[0178] Although the embodiments of the present invention are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations fall within the scope defined by the appended claims.

Claims

1. A training method for a Hausa voiceprint recognition model, characterized in that, Including: Obtaining the first frequency-domain feature and the first voiceprint feature of the English audio sample, as well as the second frequency-domain feature and the second voiceprint feature of the Hausa audio sample; Training the Hausa voiceprint recognition model based on the first frequency-domain feature and the first voiceprint feature, determining the initial parameters of the Hausa voiceprint recognition model, and obtaining an initial Hausa voiceprint recognition model; Training the initial Hausa voiceprint recognition model based on the second frequency-domain feature and the second voiceprint feature, adjusting the initial parameters of the initial Hausa voiceprint recognition model, and determining a target Hausa voiceprint recognition model, where the output of the target Hausa voiceprint recognition model is the voiceprint feature of the speaker; Obtaining intra-class data and inter-class data, where the intra-class data is the audio data of the same speaker, and the inter-class data is the audio data of different speakers; extracting the frequency-domain features of the intra-class data and the inter-class data; Inputting the extracted frequency-domain features into the target Hausa voiceprint recognition model, determining the voiceprint features corresponding to each piece of the intra-class data and the voiceprint features corresponding to each piece of the inter-class data; determining a voiceprint recognition threshold based on the similarity of the voiceprint features corresponding to each piece of the intra-class data and the similarity of the voiceprint features corresponding to each piece of the inter-class data.

2. The training method according to claim 1, wherein The training of the Hausa voiceprint recognition model based on the first frequency-domain feature and the first voiceprint feature, determining the initial parameters of the Hausa voiceprint recognition model, and obtaining an initial Hausa voiceprint recognition model includes: Inputting the first frequency-domain feature into the Hausa voiceprint recognition model to obtain a first predicted voiceprint feature; Adjusting the parameters of the Hausa voiceprint recognition model based on the error between the first voiceprint feature and the first predicted voiceprint feature, and determining the initial Hausa voiceprint recognition model.

3. The training method according to claim 2, wherein The inputting the first frequency-domain feature into the Hausa voiceprint recognition model to obtain a first predicted voiceprint feature includes: Processing the first frequency-domain feature using a first network model in the Hausa voiceprint recognition model to obtain speaker information at the frame level; Clustering the speaker information at the frame level using a second network model in the Hausa voiceprint recognition model to obtain speaker information at the sentence level, and determining the first predicted voiceprint feature.

4. The training method according to claim 2, characterized in that The adjusting the parameters of the Hausa voiceprint recognition model based on the error between the first voiceprint feature and the first predicted voiceprint feature, and determining the initial parameters of the initial Hausa voiceprint recognition model includes: Calculating a loss function using the first voiceprint feature and the first predicted voiceprint feature; Adjusting the parameters of the Hausa voiceprint recognition model based on the calculation result of the loss function, and determining the initial parameters of the initial Hausa voiceprint recognition model.

5. The training method according to claim 1, characterized in that The training of the initial Hausa voiceprint recognition model based on the Hausa audio sample and the second voiceprint feature, adjusting the initial parameters of the initial Hausa voiceprint recognition model, and determining a target Hausa voiceprint recognition model includes: Input the second frequency-domain feature into the initial Hausa voiceprint recognition model to obtain a second predicted voiceprint feature; Based on the error between the second voiceprint feature and the second predicted voiceprint feature, adjust the initial parameters of the initial Hausa voiceprint recognition model to determine the target Hausa voiceprint recognition model.

6. The training method according to claim 1, wherein The obtaining of the first frequency-domain feature of the English audio sample and the obtaining of the second frequency-domain feature of the Hausa audio sample include: Divide the English audio sample and the Hausa audio sample into silent segments and non-silent segments respectively; Perform Fourier transform processing on the non-silent segment of the English audio sample and the non-silent segment of the Hausa audio sample respectively to obtain the first frequency-domain feature and the second frequency-domain feature.

7. A Hausa voiceprint recognition method, characterized in that, Include: Obtain the audio to be recognized; Extract the frequency-domain feature of the audio to be recognized; Input the extracted frequency-domain feature into the target Hausa voiceprint recognition model to obtain a target voiceprint feature, where the target Hausa voiceprint recognition model is trained according to the training method of the Hausa voiceprint recognition model described in any one of claims 1-6; Based on the target voiceprint feature, the voiceprint feature to be matched in the voiceprint feature library, and the voiceprint recognition threshold, determine the speaker corresponding to the audio to be recognized.

8. A training device for a Hausa voiceprint recognition model, characterized in that, Include: A first acquisition module, configured to acquire the first frequency-domain feature and the first voiceprint feature of the English audio sample, and the second frequency-domain feature and the second voiceprint feature of the Hausa audio sample; A first extraction module, configured to extract the first frequency-domain feature of the English audio sample and the second frequency-domain feature of the Hausa audio sample; A first training module, configured to train the Hausa voiceprint recognition model based on the first frequency-domain feature and the first voiceprint feature, determine the initial parameters of the Hausa voiceprint recognition model, and obtain an initial Hausa voiceprint recognition model; A second training module, configured to train the initial Hausa voiceprint recognition model based on the second frequency-domain feature and the second voiceprint feature, adjust the initial parameters of the initial Hausa voiceprint recognition model, and determine a target Hausa voiceprint recognition model, where the output of the target Hausa voiceprint recognition model is the voiceprint feature of the speaker; A threshold determination module, configured to acquire intra-class data and inter-class data, where the intra-class data is audio data of the same speaker, and the inter-class data is audio data of different speakers; extract the frequency-domain features of the intra-class data and the inter-class data; Input the extracted frequency-domain features into the target Hausa voiceprint recognition model, determine the voiceprint features corresponding to each piece of the intra-class data and the voiceprint features corresponding to each piece of the inter-class data; based on the similarity of the voiceprint features corresponding to each piece of the intra-class data and the similarity of the voiceprint features corresponding to each piece of the inter-class data, determine the voiceprint recognition threshold.

9. A Hausa voiceprint recognition device, characterized in that, Include: A second acquisition module, configured to acquire the audio to be recognized; A second extraction module, configured to extract the frequency-domain feature of the audio to be recognized; An identification module, configured to input the extracted frequency-domain features into a target Hausa voiceprint recognition model to obtain target voiceprint features, where the target Hausa voiceprint recognition model is trained according to the training method of the Hausa voiceprint recognition model described in any one of claims 1-6; A determination module, configured to determine the speaker corresponding to the audio to be recognized based on the target voiceprint features, the voiceprint features to be matched in the voiceprint feature library, and the voiceprint recognition threshold.

10. An electronic device, characterized in that, Comprising: A memory and a processor, where the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the training method of the Hausa voiceprint recognition model described in any one of claims 1-6, or the Hausa voiceprint recognition method described in claim 7.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to execute the training method of the Hausa voiceprint recognition model described in any one of claims 1-6, or the Hausa voiceprint recognition method described in claim 7.

Citation Information

Patent Citations

  • Method and system for training voiceprint recognition model

    CN107610709A

  • End-to-end architecture Lhasa dialect voice recognition method based on Tibetan components

    CN109949796A

  • Speaker confirmation method and device based on neural network, equipment and storage medium

    CN110415708A