A method, device, and storage medium for training a voiceprint recognition model
By performing reverse operation and random splicing of audio data, new audio data is generated, and the problem of insufficient number and diversity of audio data in the existing technology is solved, and the recognition effect and anti-interference of the voiceprint recognition model are improved.
Patent Information
- Application Number
- CN202111582909.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-22
AI Technical Summary
The existing voiceprint recognition model training methods need to maintain the timing of the audio, limiting the number and diversity of audio data obtained, resulting in insufficient recognition effect and anti-interference.
By performing audio reverse operation and random splicing operations on some audio data in the audio training set, new audio data is generated, and these data are added to the training set, and audio features are extracted to train the voiceprint recognition model.
The number and diversity of audio data have been increased, and the recognition effect and anti-interference of the voiceprint recognition model have been improved.
Smart Images

Figure CN114420136B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voiceprint recognition, and particularly to a method, device and storage medium for training a voiceprint recognition model. Background Art
[0002] In the field of voiceprint recognition, there are generally differences between the actual application scenarios and the recording environments of standard data sets. In order to reduce the impact of such differences on the recognition effect, when training a voiceprint recognition model, various data augmentation methods such as adding noise, adding reverberation, accelerating / slowing down the speed, and spectral enhancement are used to process audio data. Through the above data augmentation methods, the quality of the audio collected in the standard recording environment is made close to the quality of the audio collected in the actual application scenario. However, the existing data augmentation methods need to maintain the timeliness of the audio, which limits the quantity and diversity of the audio data that can be obtained. Summary of the Invention
[0003] In order to overcome the above technical problems, the present invention proposes a method for training a voiceprint recognition model, and the technical solution of the method is as follows:
[0004] S1, obtaining an audio training set;
[0005] S2, performing an audio reverse operation on at least part of the audio data in the audio training set to obtain reverse audio data, and adding the reverse audio data as audio data to the audio training set;
[0006] S3, extracting audio features of all audio data in the audio training set added with the reverse audio data;
[0007] S4, using the audio features of the extracted audio data to train a pre-constructed voiceprint recognition model;
[0008] Wherein, the output of the voiceprint recognition model is an embedded feature sequence of the audio data.
[0009] Further, the audio reverse operation includes: completely reversing the sampling points of the audio data in time.
[0010] Further, the completely reversing the sampling points of the audio data in time includes:
[0011] Calculating the number of sampling points of the audio data and the values of each sampling point, and then using the center point as the axis of symmetry to exchange the values corresponding to two symmetric sampling points to generate reverse audio data.
[0012] Further, before performing the audio reverse operation on at least part of the audio data in the audio training set, a random splicing operation is also included on at least part of the audio data in the audio training set.
[0013] Further, the audio data includes speaker information. The random splicing operation specifically cuts the audio data into segments according to a preset time length to obtain cut segments of the audio data, randomly splices the cut segments of the audio data with the same speaker information to obtain spliced audio data, and combines the audio data and the spliced audio data.
[0014] Further, the audio data includes speaker information. Embedded feature sequences of two different audio data are extracted through a trained voiceprint recognition model, and the similarity score between the two embedded feature sequences is calculated. When the speaker information of the two different audio data is the same, the similarity score is higher than a preset first threshold; when the speaker information of the two different audio data is different, the similarity score is lower than a preset second threshold; wherein, the preset first threshold is not less than the preset second threshold.
[0015] Further, the audio feature of the audio data is specifically 80-dimensional Fbank feature, and cepstral mean normalization is performed on the 80-dimensional Fbank feature.
[0016] Further, before the step S3, data augmentation operations are performed on at least part of the audio data in the audio training set obtained in the step S2, and the data augmentation operations include at least one of the following: adding noise, adding reverberation, changing speed, spectral enhancement;
[0017] Voice activity detection is performed on all the audio data in the audio training set after the data augmentation operation, and the silent segments of the audio data are removed.
[0018] The present invention also provides an apparatus for training a voiceprint recognition model. The apparatus for training a voiceprint recognition model stores computer instructions; the computer instructions cause the apparatus for training a voiceprint recognition model to execute the method for training a voiceprint recognition model as described in any one of the above.
[0019] The present invention also provides a computer-readable storage medium. The computer-readable storage medium stores computer instructions, and the computer instructions cause a computer to execute the method for training a voiceprint recognition model as described in any one of the above.
[0020] The beneficial effects brought by the technical solution provided by the present invention are:
[0021] The method and apparatus for training a voiceprint recognition model of the present invention can increase the quantity and diversity of audio data, and improve the recognition effect and anti-interference ability (i.e., robustness) of the voiceprint recognition model. Description of the Drawings
[0022] Figure 1Flowchart of a method for training a voiceprint recognition model according to an embodiment of the present invention;
[0023] Figure 2 Flowchart of a method for training a voiceprint recognition model according to an embodiment of the present invention;
[0024] Figure 3 Schematic structural diagram of a device for training a voiceprint recognition model according to an embodiment of the present invention. Detailed implementation manners
[0025] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention.
[0026] Embodiment 1:
[0027] As Figure 1 shown is a flowchart of a method for training a voiceprint recognition model according to an embodiment of the present invention, showing the specific steps of the method, including:
[0028] S1. Obtain an audio training set;
[0029] S2. Perform an audio reverse operation on at least part of the audio data in the audio training set to obtain reverse audio data, and add the reverse audio data as audio data to the audio training set;
[0030] S3. Extract audio features of all audio data in the audio training set added with the reverse audio data;
[0031] S4. Use the audio features of the extracted audio data to train a pre-constructed voiceprint recognition model;
[0032] Wherein, the output of the voiceprint recognition model is an embedded feature sequence of the audio data.
[0033] Specifically, the audio reverse operation includes: completely reversing the sampling points of the audio data in time.
[0034] Specifically, the completely reversing the sampling points of the audio data in time includes:
[0035] Calculating the number of sampling points of the audio data and the values of each sampling point, and then using the center point as the axis of symmetry to exchange the corresponding values of two symmetric sampling points to generate reverse audio data.
[0036] Specifically, before performing audio reverse operation on at least part of the audio data in the audio training set, a random splicing operation is also performed on at least part of the audio data in the audio training set.
[0037] Specifically, the audio data contains speaker information. The random splicing operation specifically cuts the audio data into segments according to a preset time length to obtain cut segments of the audio data, randomly splices the cut segments of the audio data with the same speaker information to obtain spliced audio data, and combines the audio data and the spliced audio data.
[0038] Specifically, the audio data contains speaker information. The embedded feature sequences of two different audio data are extracted through a trained voiceprint recognition model, and the similarity score between the two embedded feature sequences is calculated. When the speaker information of the two different audio data is the same, the similarity score is higher than a preset first threshold; when the speaker information of the two different audio data is different, the similarity score is lower than a preset second threshold; wherein, the preset first threshold is not less than the preset second threshold.
[0039] Specifically, the audio feature of the audio data is specifically 80-dimensional Fbank feature, and cepstral mean normalization is performed on the 80-dimensional Fbank feature.
[0040] Specifically, before the step S3, a data augmentation operation is performed on at least part of the audio data in the audio training set obtained in the step S2. The data augmentation operation includes at least one of the following: adding noise, adding reverberation, changing speed, spectral augmentation;
[0041] Perform voice activity detection on all the audio data in the audio training set after the data augmentation operation, and remove the silent segments of the audio data.
[0042] Embodiment 2:
[0043] As Figure 2 shown is a flowchart of training a voiceprint recognition model according to an embodiment of the present invention, showing the specific implementation steps of training the voiceprint recognition model, including:
[0044] In step S201, an audio data set is constructed and divided into a training set and a test set.
[0045] In a possible implementation, audio data is collected through a recording pen, a microphone, WeChat, telephone recording, and / or voice synthesis, etc., and the speaker information of the audio data is labeled to construct an audio data set. The audio data set is divided into a training set and a test set by using a random splitting method or a K-fold splitting method.
[0046] In step S202, a random splicing operation is performed on all the audio data in the training set.
[0047] In a possible implementation, all audio data in the training set can be first cut according to a preset time length to generate cut segments of the audio data; then the cut segments with the same speaker information are randomly spliced to generate spliced audio data, and the number of spliced audio data is the same as the number of audio data with the same speaker information; finally, the audio data with the same speaker information and the spliced audio data are combined to obtain new audio data, and the speaker information of the new audio data is the same as that of the audio data, and the audio data in the training set is replaced with the new audio data obtained after the above combination. In other embodiments, other random splicing methods can be considered to splice the audio data.
[0048] In another possible implementation, the new audio data obtained after the above combination can be added to the training set to increase the number of audio data in the training set.
[0049] By randomly splicing the audio data of the same speaker, the combination of different speech segments of the same speaker can be realized, and the data for training the voiceprint recognition model is enhanced.
[0050] In another possible implementation, the random splicing operation of this step can be not performed, and step S203 can be directly executed.
[0051] In step S203, an audio reverse operation is performed on all audio data in the training set;
[0052] In a possible implementation, an audio reverse operation is performed on all audio data in the training set, that is, the sampling points of the audio data are completely reversed in time. Exemplarily, completely reversing the sampling points of the audio data in time can specifically include: calculating the number of sampling points of each audio data and the values of each sampling point, and then using the center point as the axis of symmetry to swap the values corresponding to two symmetric sampling points to generate reverse audio data. Among them, the speaker information of the reverse audio data is the same as that of the audio data. The reverse audio data is added to the training set to increase the number of audio data in the training set. For voiceprint recognition, obtaining reverse data by changing the time sequence as described above is equivalent to adding a new audio data of the same speaker. Thereby, the data for training the voiceprint recognition model is enhanced, and the recognition effect and anti-interference ability of the voiceprint recognition model are improved.
[0053] In step S204, a data augmentation operation is performed on all audio in the training set.
[0054] In a possible implementation, other data augmentation operations at least include one of the following: adding noise, adding reverberation, changing speed, and spectral augmentation. Of course, other types of data augmentation operations can also be performed on the audio. Perform the operation of step S205 on the data obtained after performing the data augmentation operation. It should be noted that the audio data after the data augmentation operation can also be added to the training set and used together with the original data in the training set as the data in the training set, which can expand the quantity of the audio data in the training set.
[0055] In step S205, extract the audio features of all the audio data in the training set and the test set.
[0056] In a possible implementation, first perform voice activity detection (VAD) on all the audio data in the training set and the test set to remove the silent segments of the audio data; then extract the 80-dimensional Fbank features of the audio data, and perform cepstral mean normalization (CMN) on the 80-dimensional Fbank features as the audio features of the audio data.
[0057] In step 206, use the training set and the test set respectively to train and test the pre-constructed speaker recognition model to obtain the trained speaker recognition model.
[0058] In a possible implementation, the speaker recognition model is implemented using a residual network (ResNet). The audio features of the audio data are sliced into 200 frames as the input. The number of network layers of the residual network is 34 layers. The convolution of each convolutional layer of the residual network uses one-dimensional convolution, and an SE module is added. Batch normalization is performed on the output of the convolutional layer. The output of the last convolutional layer is input into the attention pooling layer to output the embedding feature sequence (Embedding) of the audio data; the optimizer of the residual network selects the AdamW optimization algorithm, the learning rate strategy selects cyclic learning rates (CyclicLR), and the AAM-Softmax and cross-entropy loss functions are used for the classification of the embedding feature sequence of the audio data; the training set and the test set are used to train and test the speaker recognition model respectively, and the trained speaker recognition model is obtained after multiple rounds of training and testing.
[0059] Extract the embedded feature sequences of two different audio data using the trained voiceprint recognition model, calculate the similarity score of the two embedded feature sequences. When the speaker information of the two different audio data is the same, the similarity score is higher than a preset first threshold; when the speaker information of the two different audio data is different, the similarity score is lower than a preset second threshold; wherein, the preset first threshold is greater than the preset second threshold, and the similarity score is calculated using cosine similarity.
[0060] It should be noted that the number of frames for splitting the above audio features is 200 frames, the number of network layers of the residual network is 34 layers, the optimizer of the residual network is the AdamW optimization algorithm, the learning strategy is the cyclic learning rate, and the loss function is the AAM-Softmax and cross-entropy loss function. The similarity score is calculated using cosine similarity, and other implementation methods can be adopted, and the present invention does not make specific limitations.
[0061] After obtaining the trained voiceprint recognition model, preferably, the trained voiceprint recognition model can be used for voiceprint verification and voiceprint identification, and can also be used in other usage scenarios of voiceprint recognition applications.
[0062] The trained voiceprint recognition model obtained by the method of the embodiment of the present invention can be used for voiceprint verification or voiceprint identification, as shown in steps S207 and S208.
[0063] A preferred embodiment of the present invention utilizes the reverse of the audio in the time domain and the random splicing of segments of the same speaker, which can increase the quantity and diversity of the voiceprint model training data during the voiceprint recognition process, and at the same time weaken the influence of the time sequence on voiceprint recognition, improving the recognition effect and robustness of the system.
[0064] In step S207, the trained voiceprint recognition model is used for voiceprint verification.
[0065] Voiceprint verification, also known as speaker verification, whose English is Speaker Verification, refers to determining whether the speakers corresponding to two audio data are the same person. The steps of voiceprint verification include: first, obtain the two audio data to be verified, respectively extract the audio features corresponding to the audio data, then input the audio features into the trained voiceprint recognition model to obtain the embedded feature sequences corresponding to the audio data, and finally use cosine similarity to calculate the similarity score of the two embedded feature sequences. When the similarity score is higher than the preset threshold, it is determined to be the same person, otherwise, it is determined not to be the same person.
[0066] In step S208, the trained voiceprint recognition model is used for voiceprint identification.
[0067] Voiceprint recognition, also known as speaker recognition, is called Speaker Recognition / Identification in English, which refers to determining which speaker a piece of audio belongs to. The steps of voiceprint recognition include: First, establish a basic audio database, which contains one or more pieces of audio data. The audio data contains speaker information and corresponding embedded feature sequences. Then, obtain the audio data to be recognized, extract the audio features of the audio data to be recognized, input the audio features of the audio data to be recognized into the voiceprint recognition model trained by the method of the embodiments of the present invention to obtain the embedded feature sequence of the audio data to be recognized. Then, calculate the similarity scores between the embedded feature sequence of the audio data to be recognized and the embedded feature sequences of the audio data in the basic audio database respectively, and screen out the audio data in the basic audio database whose similarity scores are greater than the preset threshold. If no audio data is screened out, it means that there is no speaker information corresponding to the audio data to be recognized in the basic audio database; on the contrary, confirm the speaker information corresponding to the audio data to be recognized from the screened-out audio data.
[0068] Embodiment 3:
[0069] The present invention also provides a device for training a voiceprint recognition model, as Figure 3 shown. The device includes a processor 301, a memory 302, a bus 303, and a computer program stored in the memory 302 and executable on the processor 301. The processor 301 includes one or more processing cores. The memory 302 is connected to the processor 301 through the bus 303. The memory 302 stores program instructions. When the processor executes the computer program, it implements the steps in the method embodiments of the present invention described above.
[0070] Further, as an executable solution, the device for training a voiceprint recognition model can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The system / electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above-described composition structure of the system / electronic device is only an example of the system / electronic device, and does not constitute a limitation on the system / electronic device. It may include more or fewer components than the above, or combine some components, or different components. For example, the system / electronic device may further include input / output devices, network access devices, a bus, etc. The embodiments of the present invention do not make any limitations in this regard.
[0071] Further, as an executable solution, the so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the system / electronic device, and connects various parts of the entire system / electronic device using various interfaces and circuits.
[0072] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the system / electronic device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0073] Embodiment 4:
[0074] The present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method in the above embodiments of the present invention are realized.
[0075] When the modules / units integrated in a system / electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.
[0076] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in terms of form and details without departing from the spirit and scope of the present invention defined by the appended claims, and all such changes are within the protection scope of the present invention.
Claims
1. A method for training a voiceprint recognition model, characterized in that, it includes: S1, obtaining an audio training set; S2, performing an audio reverse operation on at least part of the audio data in the audio training set to obtain reverse audio data, and adding the reverse audio data as audio data to the audio training set; wherein, the audio reverse operation includes: calculating the number of sampling points of the audio data and the values of each sampling point, and then using the center point as the axis of symmetry to swap the values corresponding to two symmetric sampling points to generate reverse audio data; S3, extracting the audio features of all audio data in the audio training set added with the reverse audio data; S4, using the audio features of the extracted audio data to train a pre-constructed voiceprint recognition model; wherein, the output of the voiceprint recognition model is the embedded feature sequence of the audio data.
2. The method according to claim 1, characterized in that, before performing the audio reverse operation on at least part of the audio data in the audio training set, a random splicing operation is also included on at least part of the audio data in the audio training set.
3. The method according to claim 2, characterized in that, the audio data contains speaker information, and the random splicing operation is specifically to cut the audio data according to a preset time length to obtain cut segments of the audio data, randomly splice the cut segments of the audio data with the same speaker information to obtain spliced audio data, and merge the audio data and the spliced audio data.
4. The method according to claim 1, characterized in that, the audio data contains speaker information, the embedded feature sequences of two different audio data are extracted through a trained voiceprint recognition model, the similarity score of the two embedded feature sequences is calculated, when the speaker information of the two different audio data is the same, the similarity score is higher than a preset first threshold; when the speaker information of the two different audio data is different, the similarity score is lower than a preset second threshold; wherein, the preset first threshold is not less than the preset second threshold.
5. The method according to claim 1, characterized in that, the audio features of the audio data are specifically 80-dimensional Fbank features, and cepstral mean normalization is performed on the 80-dimensional Fbank features.
6. The method according to claim 1, characterized in that, before the step S3, a data augmentation operation is performed on at least part of the audio data in the audio training set obtained in step S2, and the data augmentation operation includes at least one of the following: adding noise, adding reverberation, changing speed, spectral enhancement; Performing voice activity detection on all audio data in the audio training set after the data augmentation operation, and removing the silent segments of the audio data.
7. An apparatus for training a voiceprint recognition model, characterized in that, it includes a memory and a processor, the memory stores at least one segment of program, and the at least one segment of program is executed by the processor to implement the voiceprint recognition model training method according to any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, At least one program is stored in the storage medium, and the at least one program is executed by a processor to implement the voiceprint recognition model training method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Training method and device of voiceprint recognition model, electronic equipment and storage medium
CN109801636A
Voiceprint recognition system and method
CN110060692A