A method and device for training a voiceprint recognition model and voiceprint recognition
By using a voiceprint recognition model with multiple rounds of iterative training and an innovative objective function, the problem of low voiceprint recognition accuracy in offline speaker separation was solved, and high-accuracy voiceprint recognition was achieved in offline scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-22
- Publication Date
- 2026-03-27
AI Technical Summary
Existing voiceprint recognition technologies cannot accurately identify voiceprint features in offline speaker separation scenarios, resulting in low recognition accuracy.
By training a voiceprint recognition model, utilizing multi-round iterative training and an innovative objective function, combined with data processing methods and multi-round sampling techniques, the robustness and accuracy of the model are improved, making it suitable for offline speaker separation scenarios.
It improves the accuracy of voiceprint recognition for offline speaker separation, and can accurately distinguish the voiceprint features of different objects over a longer time span, making it suitable for speech transcription assistance in multi-person scenarios.
Smart Images

Figure CN116110404B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voiceprint recognition technology, and more specifically, to a method and apparatus for training a voiceprint recognition model and for voiceprint recognition. Background Technology
[0002] Speaker diarization typically refers to segmenting an audio clip into different speakers to obtain the voiceprint features of each speaker.
[0003] Currently, thanks to the development of voiceprint recognition or voiceprint verification technologies, academia and industry have conducted continuous and in-depth research on voiceprint feature extraction. Abundant reference materials exist for speaker separation voiceprint feature extraction modules, and many speaker separation systems even directly reuse voiceprint feature extraction modules from voiceprint recognition. Voiceprint recognition or voiceprint identification systems are typically used in application scenarios with a long time span, such as when a user registers a set of voiceprint information and does not need to register again for several months or even years, allowing for voiceprint recognition or identification. However, because a speaker's voice can change due to recording equipment, physical development, health conditions, and even emotional changes, it cannot be applied to offline speaker voiceprint recognition, making it impossible to obtain accurate voiceprint recognition results in offline speaker audio scenarios.
[0004] Therefore, how to provide a technical solution for voiceprint recognition with high accuracy has become an urgent technical problem to be solved. Summary of the Invention
[0005] The purpose of some embodiments of this application is to provide a method and apparatus for training a voiceprint recognition model and for voiceprint recognition. The technical solutions of the embodiments of this application can improve the accuracy of voiceprint recognition of the speaker and have good versatility.
[0006] In a first aspect, some embodiments of this application provide a method for training a voiceprint recognition model, comprising: acquiring a training sample dataset, wherein the training sample dataset includes: multiple audio samples of multiple objects; cyclically executing the following process until at least the (i+1)th voiceprint similarity satisfies a preset condition, and using the (i+1)th voiceprint recognition model as a target voiceprint recognition model; acquiring an (i+1)th training sample set based on the (i)th voiceprint similarity and the training sample dataset, wherein the (i)th voiceprint similarity is obtained based on the (i)th voiceprint recognition model and the audio of multiple speaking objects; training the (i+1)th voiceprint recognition model using the (i+1)th training sample set to obtain the (i+1)th voiceprint recognition model; and acquiring the (i+1)th voiceprint similarity among the multiple speaking objects based on the (i+1)th voiceprint recognition model and the audio of the multiple speaking objects.
[0007] Some embodiments of this application train voiceprint recognition models at different training stages using training sample datasets until a target voiceprint recognition model that meets preset conditions is obtained. These embodiments can obtain robust models, providing a model foundation for subsequent voiceprint recognition, thereby improving the accuracy of speaker voiceprint recognition and demonstrating good versatility.
[0008] In some embodiments, when i=1, the i-th voiceprint recognition model is obtained by randomly selecting a first training sample set from the training sample dataset to train the initial voiceprint recognition model, thereby obtaining the first voiceprint recognition model.
[0009] Some embodiments of this application obtain a first voiceprint recognition model by randomly selecting a first training sample set to train the initial voiceprint recognition model, which can achieve effective training of the model.
[0010] In some embodiments, the objective function of the i-th voiceprint recognition model is related to the (i+1)-th hyperparameter, the mean similarity of samples in the (i+1)-th training sample set, and the extreme value of sample similarity, wherein the first hyperparameter of the initial voiceprint recognition model and the i-th voiceprint recognition model are different.
[0011] Some embodiments of this application design the objective function of the i-th voiceprint recognition model so that the objective function is related to multiple parameters, which can achieve accurate training of the model and improve the training effect of the model.
[0012] In some embodiments, the mean sample similarity is related to the mean voiceprint feature similarity of similar samples and the mean voiceprint feature similarity of dissimilar samples, wherein the similar samples represent audio samples of the same object, and the dissimilar samples represent audio samples of different objects.
[0013] Some embodiments of this application can accurately obtain model-related parameters by acquiring the mean sample similarity, thereby improving the effectiveness of the model.
[0014] In some embodiments, the extreme value of sample similarity is related to the i-th adjustment hyperparameter, the minimum value of voiceprint feature similarity of the same type of samples, and the maximum value of voiceprint feature similarity of the different type of samples.
[0015] Some embodiments of this application obtain the extreme values of sample similarity through multiple parameters, ensuring the effectiveness of the model.
[0016] In some embodiments, obtaining the (i+1)th training sample set based on the i-th voiceprint similarity and the training sample dataset includes: randomly selecting a first proportion value; obtaining a first sample corresponding to the first proportion value based on the i-th voiceprint similarity, and selecting a second sample corresponding to a second proportion value from the training sample dataset, wherein the sum of the first proportion value and the second proportion value is 1, and the first sample and the second sample constitute the (i+1)th training sample set.
[0017] Some embodiments of this application can improve the robustness and accuracy of the trained model by selecting a training sample set.
[0018] In some embodiments, the at least i+1th voiceprint similarity satisfies a preset condition, including: the accuracy of the i+1th voiceprint similarity is not less than a preset value; or, both the i-th voiceprint similarity and the i+1th voiceprint similarity fall within a preset range.
[0019] Some embodiments of this application use different methods as preset conditions for terminating model training, which can quickly obtain a target voiceprint recognition model with good robustness and high accuracy.
[0020] Secondly, some embodiments of this application provide a voiceprint recognition method, including: acquiring an audio to be recognized; inputting the audio to be recognized into a target voiceprint recognition model obtained by the method described in any embodiment of the first aspect, to obtain a voiceprint recognition result, wherein the voiceprint recognition result is used to confirm the number of target objects contained in the audio.
[0021] Some embodiments of this application can achieve accurate identification of the audio to be identified through the target voiceprint recognition model, confirming whether it belongs to the target object, with high accuracy.
[0022] Thirdly, some embodiments of this application provide an apparatus for training a voiceprint recognition model, comprising: a sample acquisition module for acquiring a training sample dataset, wherein the training sample dataset includes: multiple audio samples of multiple objects; and a model training module for: cyclically executing the following process until at least the (i+1)th voiceprint similarity satisfies a preset condition, and using the (i+1)th voiceprint recognition model as the target voiceprint recognition model: acquiring an (i+1)th training sample set based on the (i)th voiceprint similarity and the training sample dataset, wherein the (i)th voiceprint similarity is obtained based on the (i)th voiceprint recognition model and the audio of multiple speaking objects; training the (i+1)th voiceprint recognition model using the (i+1)th training sample set to obtain the (i+1)th voiceprint recognition model; and acquiring the (i+1)th voiceprint similarity among the multiple speaking objects based on the (i+1)th voiceprint recognition model and the audio of the multiple speaking objects.
[0023] Fourthly, some embodiments of this application provide a voiceprint recognition apparatus, comprising: an acquisition module for acquiring audio to be recognized; and an identification module for inputting the audio to be recognized into a target voiceprint recognition model obtained by the method described in any embodiment of the first aspect, to obtain a voiceprint recognition result, wherein the voiceprint recognition result is used to confirm the number of target objects contained in the audio.
[0024] Thirdly, some embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the method described in any embodiment of the first aspect.
[0025] Fourthly, some embodiments of this application provide an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, can implement the method as described in any embodiment of the first aspect.
[0026] Fifthly, some embodiments of this application provide a computer program product, the computer program product including a computer program, wherein the computer program, when executed by a processor, can implement the method described in any embodiment of the first aspect. Attached Figure Description
[0027] To more clearly illustrate the technical solutions of some embodiments of this application, the accompanying drawings used in some embodiments of this application will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A system diagram for voiceprint recognition is provided for some embodiments of this application;
[0029] Figure 2 Flowchart of a method for training a voiceprint recognition model provided for some embodiments of this application;
[0030] Figure 3 A flowchart of a voiceprint recognition method provided for some embodiments of this application;
[0031] Figure 4 Block diagram of an apparatus for training a voiceprint recognition model provided for some embodiments of this application;
[0032] Figure 5 Block diagrams of a voiceprint recognition device provided for some embodiments of this application;
[0033] Figure 6A schematic diagram of an electronic device provided for some embodiments of this application. Detailed Implementation
[0034] The technical solutions of some embodiments of this application will now be described with reference to the accompanying drawings.
[0035] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0036] In related technologies, speaker digitization typically refers to the technique of segmenting an audio segment according to different speaker IDs. It generally includes speech detection (i.e., speech endpoint detection, or VoiceActivity Detection, or VAD for short, used to detect whether an audio segment belongs to a specific speaker), speech segmentation (or speaker conversion detection, which initially cuts long audio segments into shorter segments that minimize the number of speakers), speaker feature extraction (also known as speaker embedding, which extracts speaker voiceprint features from the audio), clustering analysis (clustering the feature vectors representing speaker segments), and secondary segmentation (further segmenting and adjusting audio segments containing multiple speakers). This speaker digitization technique is commonly used in speech transcription assistance work in multi-person scenarios such as meetings and interviews. Based on real-time requirements, it can be divided into online speaker digitization and offline speaker digitization. Online speaker digitization requires real-time speaker segmentation results, while offline speaker digitization can provide the results some time after the meeting or interview has ended. However, regardless of whether it is offline or online, voiceprint feature extraction is an indispensable and crucial core component in the speaker separation process. Because the application scenarios of voiceprint recognition or identification differ somewhat from speaker separation, especially offline speaker separation, there are currently no technical solutions to apply voiceprint recognition or identification to offline speaker separation to improve the accuracy of voiceprint recognition for different individuals.
[0037] The applicant's research revealed that, on the one hand, voiceprint recognition or identification tasks (such as personalized product recommendations based on voiceprint recognition results, or determining whether a user is a target user based on voiceprint recognition results and then performing authentication operations) generally require the system to provide identification results within a very short time (usually within hundreds or even tens of milliseconds). Offline speaker separation, however, has a certain tolerance for real-time performance, which allows voiceprint feature models suitable for offline speaker separation to have a larger parameter scale and thus greater modeling capabilities.
[0038] On the other hand, voiceprint recognition or identification systems are often used in applications with a long time span (for example, a user can register a voiceprint and then use the voiceprint recognition or identification function for months or even years without needing to register again). However, human voices change with recording equipment, physical development, health conditions, and even emotional changes. This requires that the voiceprint features extracted by voiceprint recognition or identification systems be robust enough, not only to distinguish between different speakers but also to classify the timbre features of the same person into one category over a long time span. Speaker segmentation, on the other hand, is usually performed in relatively short time spans (e.g., tens of minutes or hours), and it is generally possible to ensure that the audio segments to be segmented are recorded using the same device. Therefore, the voiceprint feature extraction model can focus more on distinguishing between different speakers.
[0039] In view of this, some embodiments of this application propose a voiceprint extraction method, which can mainly complete the offline speaker separation task. By inputting audio into the target voiceprint recognition model of the training number, the voiceprint features in the voiceprint recognition result are obtained, thereby enabling the differentiation of offline speakers and improving the accuracy of offline speaker separation.
[0040] The following is in conjunction with the appendix Figure 1 The present application provides an exemplary embodiment of the overall structure of a voiceprint recognition system.
[0041] like Figure 1As shown, some embodiments of this application provide a voiceprint recognition system. The voiceprint recognition system includes a terminal 100 and a voiceprint recognition server 200. The terminal 100 is used to collect and / or store audio to be recognized. The audio to be recognized can be recorded in real-time by the terminal 100 during a meeting and stored in the terminal 100 after the meeting ends. Alternatively, the terminal 100 can perform recording and storage simultaneously. The terminal 100 sends the audio to be recognized to the voiceprint recognition server 200. The voiceprint recognition server 200 uses an internally deployed target voiceprint recognition model to extract features from the audio to be recognized, obtains the voiceprint recognition result, and sends it to the terminal 100. The terminal 100 can determine the number of speakers (as a specific example of a target object) contained in the audio to be recognized based on the voiceprint recognition result.
[0042] In some embodiments of this application, the target voiceprint recognition model is obtained by iteratively training an initial voiceprint recognition model and pre-deployed to the voiceprint recognition server 200. Alternatively, in other embodiments of this application, if the terminal 100 can also deploy the target voiceprint recognition model, then the voiceprint recognition server 200 may not be required. The embodiments of this application are not limited to these.
[0043] It should be noted that before extracting the audio to be recognized, the initial voiceprint recognition model needs to be iteratively trained to obtain the target voiceprint extraction model.
[0044] The following is in conjunction with the appendix Figure 2 The implementation process of the method for training a voiceprint recognition model provided by some embodiments of this application is illustrated by way of example.
[0045] Please see the appendix Figure 2 , Figure 2 A flowchart of a method for training a voiceprint recognition model is provided for some embodiments of this application. The method for training a voiceprint recognition model includes:
[0046] S210, Obtain the training sample dataset, wherein the training sample dataset includes: multiple audio samples of multiple objects.
[0047] For example, in some embodiments of this application, unlike traditional voiceprint recognition tasks, speaker separation tasks often require audio segmentation for voiceprint analysis. Since conversations often involve alternating short phrases of one or two words, the audio segments used for voiceprint feature analysis are short, typically between 0.5 and 2 seconds. Therefore, in the data preprocessing stage, it is necessary to simulate real-world scenarios by segmenting the audio, randomly dividing the original audio (after removing silence) into segments of random lengths between 0.5 and 2.0 seconds. Furthermore, the audio volume needs to be normalized and uniformly sampled as 16kHz WAV format audio. Using a frame length of 25 milliseconds and a step size of 10 milliseconds, 128-dimensional MFCC (Mel Frequency Cepstrum Coefficient) features are extracted and stored as a data file for later use. Following the above data preprocessing method, multiple samples from multiple different speakers are collected to form a training sample dataset. For example, each training batch samples 16 different speakers, and each speaker randomly samples 16 samples, meaning that each training batch samples a total of 256 audio samples for training.
[0048] In some embodiments of this application, before executing S220, the method for training the voiceprint recognition model further includes: when i=1, obtaining the i-th voiceprint recognition model by randomly selecting a first training sample set from the training sample dataset to train the initial voiceprint recognition model and obtain the first voiceprint recognition model.
[0049] For example, in some embodiments of this application, during the first round of training, speaker sampling is carried out by random sampling in the early stage of training, that is, 256 audio samples are randomly selected from the training sample dataset to train the initial voiceprint recognition model for the first time, so as to obtain the first voiceprint recognition model.
[0050] S220, the following process is executed repeatedly until at least the (i+1)th voiceprint similarity meets the preset condition, and the (i+1)th voiceprint recognition model is used as the target voiceprint recognition model:
[0051] S221, based on the i-th voiceprint similarity and the training sample dataset, obtain the (i+1)-th training sample set, wherein the i-th voiceprint similarity is obtained based on the i-th voiceprint recognition model and the audio of multiple speaking objects;
[0052] S222, the i-th voiceprint recognition model is trained using the (i+1)-th training sample set to obtain the (i+1)-th voiceprint recognition model;
[0053] S223, based on the (i+1)th voiceprint recognition model and the audio of the multiple speaking objects, obtain the (i+1)th voiceprint similarity among the multiple speaking objects.
[0054] For example, in some embodiments of this application, the training model from the previous round is used as the model basis for the next round of training, thereby achieving iterative training of the model and finally obtaining a target voiceprint recognition model that meets preset conditions.
[0055] In some embodiments of this application, S221 may include: randomly selecting a first proportion value; obtaining a first sample corresponding to the first proportion value based on the i-th voiceprint similarity, and selecting a second sample corresponding to a second proportion value from the training sample dataset, wherein the sum of the first proportion value and the second proportion value is 1, and the first sample and the second sample constitute the (i+1)-th training sample set.
[0056] For example, in some embodiments of this application, taking i=1 as an example, multiple speaking objects' audio is input into a first voiceprint recognition model to obtain the voiceprint features of each speaking object in the output multiple speaking objects. Then, the similarity between the voiceprint features of each speaking object output by the model and the voiceprint features corresponding to the audio of multiple speaking objects is calculated to obtain a first voiceprint similarity, wherein the first voiceprint similarity includes the voiceprint similarity of each speaking object. When selecting a second training sample set, probabilistic hard sampling is performed, that is, training samples are selected with a probability fluctuation within a certain range, for example, a random number is taken between 0.7 and 0.9 (as a specific example of the first proportion value). For example, if the random number is 0.75, then the voiceprint features of the speaking objects with higher similarity in the first voiceprint similarity are taken as hard samples, which account for 75% of the second training sample set. Then, the remaining 25% (that is, the second proportion value) of samples (as a specific example of the second sample) are randomly selected from the training sample dataset and added to the second training sample set. The similarity can be calculated using a similarity algorithm, such as the cosine similarity algorithm.
[0057] It should be noted that in some embodiments of this application, the number of audio samples included in the i-th training sample set can be set. For example, the i-th training sample set may contain 256 audio samples during each training iteration. Furthermore, for audio samples of different lengths within the same batch, zeros need to be padded to the end of the audio features to the length of the longest sample in that batch (ensuring there are 256 audio samples). It is understood that in subsequent iterative training processes, the selection method of the i-th training sample set and the calculation method of similarity are the same as described above, and will not be repeated here to avoid repetition.
[0058] In some embodiments of this application, the at least i+1th voiceprint similarity in S220 satisfies a preset condition, including: the accuracy of the i+1th voiceprint similarity is not less than a preset value; or, both the i-th voiceprint similarity and the i+1th voiceprint similarity fall within a preset range.
[0059] For example, in some embodiments of this application, training can end when the accuracy of the (i+1)th voiceprint similarity calculated by the similarity algorithm is not less than a preset value. The preset value can be 95%, 90%, or 85%, etc., and can be set according to actual conditions. In other embodiments of this application, during the loop, training can end when the voiceprint similarity of the i-th voiceprint recognition model and the (i+1)-th voiceprint recognition model in two consecutive training rounds is stable within a preset range. The preset range can be 0.5 to 1, or 0.6 to 1, etc., and can be set according to actual conditions.
[0060] In some embodiments of this application, the objective function of the i-th voiceprint recognition model is related to the i-th hyperparameter, the mean sample similarity of the i-th training sample set, and the extreme sample similarity, wherein the first hyperparameter of the initial voiceprint recognition model and the i-th voiceprint recognition model are different.
[0061] In speaker separation scenarios, only short time spans (typically tens of minutes to several hours) of audio data need to be processed. This significantly reduces the impact of changes in the speaker's physical condition over long periods (such as developmental changes, aging, or illness), which could lead to fluctuations in the speaker's vocal characteristics. Therefore, the objective function designed in this application, compared to the objective functions in traditional voiceprint recognition tasks, focuses more on distinguishing differences between different speakers. For example, in some embodiments of this application, by using hyperparameters, the mean sample similarity, and the extreme value of sample similarity as key parameters, an objective function that emphasizes distinguishing differences between different speakers is obtained.
[0062] In some embodiments of this application, the method for obtaining the objective function of the i-th voiceprint recognition model includes: multiplying the i-th hyperparameter by the mean of sample similarity to obtain a first result; multiplying the difference between 1 and the i-th hyperparameter by the extreme value of sample similarity to obtain a second result; and adding the first result and the second result to obtain the objective function of the i-th voiceprint recognition model.
[0063] Specifically, the objective function of the i-th voiceprint recognition model is formulated as follows: Loss = a i *S avg +(1-a i )*S mm Among them, S avg S is the mean of sample similarity. mm This represents the extreme value of sample similarity. i ∈[0,1] represents the i-th hyperparameter, used to adjust the proportion of the mean and extreme values of sample similarity to the loss. For example, in the first round of model training, a1 is set to 1, allowing the model to converge stably and quickly. From the second round onwards, the value is biased towards 0, thus making the feature vectors extracted by the model more discriminative. The specific value of a1 can be set according to the actual situation. iThe value of is not specifically limited in this embodiment of the application.
[0064] In some embodiments of this application, the mean sample similarity is related to the mean voiceprint feature similarity of similar samples and the mean voiceprint feature similarity of dissimilar samples, wherein the similar samples represent audio samples of the same object, and the dissimilar samples represent audio samples of different objects.
[0065] For example, in some embodiments of this application, the mean sample similarity is the difference between the mean voiceprint feature similarity of samples of the same type and the mean voiceprint feature similarity of samples of different types.
[0066] Specifically, the formula for the mean sample similarity is: S avg =S out-avg -S in-avg ;
[0067] Among them, S in-avg It is the mean of pairwise similarity of voiceprint features among similar samples (audio samples belonging to the same speaker), while S out-avg This represents the average similarity between dissimilar samples (i.e., the average similarity of voiceprint features among dissimilar samples). The similarity here is calculated using cosine similarity. Thus, the higher the similarity between samples of the same class and the lower the similarity between dissimilar samples, the higher S becomes. avg The smaller the value, the better.
[0068] In some embodiments of this application, the extreme value of sample similarity is related to the i-th adjustment hyperparameter, the minimum value of voiceprint feature similarity of the same type of samples, and the maximum value of voiceprint feature similarity of the different type of samples.
[0069] For example, in some embodiments of this application, the extreme value of sample similarity is related to the exponential function e. The extreme value of sample similarity is the difference between the result of the first exponential function and the result of the second exponential function. The result of the first exponential function is a function with base e and exponent S. out-max The sum of the i-th adjusted hyperparameter. The result of the first exponential function is a function with base e and exponent S. in-min .
[0070] Specifically, the formula for the extreme value of sample similarity is: S mm =exp(S out-avg -b i )-exp(S in-avg );
[0071] Where exp(x)=e x And b i∈(0,1) is a hyperparameter that adjusts the proportions of discriminative power (low dissimilarity) and convergent power (high similarity among similar classes) in the loss. Since the slope of exp(x) monotonically increases with x, i.e., exp(x+b+0.1)-exp(x+b)>exp(x+0.1)-exp(x), the extreme value of dissimilarity S can be optimized. out-max The reduced benefit is greater than the extreme value S of similarity among similar classes. in-min The increased gains lead the model to focus more on distinguishing outlier samples.
[0072] Some embodiments of this application use the aforementioned Loss as the loss function during model training and optimize and iterate the model using gradient descent. Ultimately, this results in the voiceprint feature vectors obtained by the model having the highest possible cosine similarity among similar classes and the lowest possible cosine similarity among dissimilar classes. This can improve the robustness and accuracy of the trained target voiceprint recognition model, thereby improving the accuracy of offline speaker separation.
[0073] In some embodiments of this application, the structures of the initial voiceprint recognition model and the i-th voiceprint recognition model include: a first reshape layer, a convolutional neural network layer, a second reshape layer, and a three-layer stacked GRU (Gate Recurrent Unit) structure.
[0074] It should be noted that the above model structure benefits from the tolerance of inference latency due to offline speaker separation. The role of each layer in the structure is as follows:
[0075] Layer 0: The first reshape layer (changing the input dimension without altering the data), reshaping the input features into [256, L0, 128, 1]. Here, L0 represents the time dimension of the longest audio segment in the current batch. For example, for a 1-second audio segment, L0 = 1000 / 10 = 100, where 1000 represents 1000 milliseconds, and 10 represents the step size of 10 milliseconds when extracting MFCCs. Furthermore, "256" refers to sampling 256 audio samples each time.
[0076] Layer 1: CNN (Convolutional Neural Network) layer, with a convolutional size of 5*5, a stride of 2, and 64 output channels. It uses ReLU as the activation function to initially extract local audio features. Calculating the features of the previous layer will give the output features of this layer as [256, L0 / 2, 64, 64].
[0077] Layer 2: The second reshape layer reshapes the output features of the previous layer into a feature vector of [256, L0 / 2, 4096].
[0078] Layer 3: A 3-layer stacked GRU (Gate Recurrent Unit, a variant of Recurrent Neural Network, RNN) structure with 512 hidden neuron parameters. The last step output feature [256, 256] of the last layer is taken to extract the accumulated temporal features of the entire audio segment, and this is used as the voiceprint feature representing the audio.
[0079] As can be seen from the above embodiments, the voiceprint recognition model of this application includes an innovative objective function, which is beneficial for multi-round model fine-tuning training. The first round helps the model converge quickly and stably, while the second to nth rounds focus on improving the effectiveness of the model in extracting voiceprints, while also emphasizing the differentiation of audio from different speakers, making it more suitable for offline speaker separation scenarios. The design of the network model structure uses a large-scale parameter scale, enabling the model to have a higher fitting ability. The embodiments of this application also employ data processing methods, multi-round sampling methods, and multi-round iterative fine-tuning training methods to improve the robustness of the trained model.
[0080] The following is in conjunction with the appendix Figure 3 The present application provides an exemplary description of the specific process of voiceprint recognition performed by the voiceprint recognition server 200, as illustrated in some embodiments of this application.
[0081] Please see the appendix Figure 3 , Figure 3 The present application provides a flowchart of a voiceprint recognition method according to some embodiments. The voiceprint recognition method includes: S310, acquiring audio to be recognized; S320, inputting the audio to be recognized into a target voiceprint recognition model to obtain a voiceprint recognition result, wherein the voiceprint recognition result is used to confirm the number of target objects contained in the audio.
[0082] For example, in some embodiments of this application, the target voiceprint recognition model is through Figure 2 The method embodiments described above train the target voiceprint recognition model, which can be pre-deployed in the voiceprint recognition server 200. In a multi-person meeting scenario, the meeting process is recorded to obtain audio containing multiple participants to be recognized. Inputting the audio to be recognized into the target voiceprint recognition model yields the voiceprint recognition result. The voiceprint recognition result includes the number of participants (as a specific example of the target object) in the audio to be recognized and the audio characteristics of each participant.
[0083] As can be seen from the above embodiments of this application, this application is suitable for voiceprint extraction in offline speaker separation scenarios and can improve the accuracy of speaker separation.
[0084] Please refer to Figure 4 , Figure 4The diagram shows a block diagram of an apparatus for training a voiceprint recognition model according to some embodiments of this application. It should be understood that the apparatus for training a voiceprint recognition model corresponds to the method embodiments described above and is capable of performing the various steps involved in the method embodiments described above. The specific functions of the apparatus for training a voiceprint recognition model can be found in the description above. To avoid repetition, detailed descriptions are appropriately omitted here.
[0085] Figure 4 The apparatus for training a voiceprint recognition model includes at least one software functional module that can be stored in a memory or embedded in the apparatus in the form of software or firmware. The apparatus includes: a sample acquisition module 410 for acquiring a training sample dataset, wherein the training sample dataset includes multiple audio samples from multiple objects; and a model training module 420 for: cyclically executing the following process until at least the (i+1)th voiceprint similarity meets a preset condition, and using the (i+1)th voiceprint recognition model as the target voiceprint recognition model: acquiring an (i+1)th training sample set based on the (i)th voiceprint similarity and the training sample dataset, wherein the (i)th voiceprint similarity is obtained based on the (i)th voiceprint recognition model and the audio of multiple speaking objects; training the (i+1)th voiceprint recognition model using the (i+1)th training sample set to obtain the (i+1)th voiceprint recognition model; and acquiring the (i+1)th voiceprint similarity between the multiple speaking objects based on the (i+1)th voiceprint recognition model and the audio of the multiple speaking objects.
[0086] In some embodiments of this application, the apparatus for training a voiceprint recognition model includes: a first training module (not shown in the figure), used to randomly select a first training sample set from the training sample dataset when i=1 to train the initial voiceprint recognition model and obtain a first voiceprint recognition model.
[0087] In some embodiments of this application, the objective function of the i-th voiceprint recognition model is related to the i-th hyperparameter, the mean sample similarity of the i-th training sample set, and the extreme sample similarity, wherein the first hyperparameter of the initial voiceprint recognition model and the i-th voiceprint recognition model are different.
[0088] In some embodiments of this application, the mean sample similarity is related to the mean voiceprint feature similarity of similar samples and the mean voiceprint feature similarity of dissimilar samples, wherein the similar samples represent audio samples of the same object, and the dissimilar samples represent audio samples of different objects.
[0089] In some embodiments of this application, the extreme value of sample similarity is related to the i-th adjustment hyperparameter, the minimum value of voiceprint feature similarity of the same type of samples, and the maximum value of voiceprint feature similarity of the different type of samples.
[0090] In some embodiments of this application, the model training module 420 is configured to: randomly select a first proportion value; obtain a first sample corresponding to the first proportion value based on the i-th voiceprint similarity, and select a second sample corresponding to a second proportion value from the training sample dataset, wherein the sum of the first proportion value and the second proportion value is 1, and the first sample and the second sample constitute the (i+1)-th training sample set.
[0091] In some embodiments of this application, the model training module 420 is configured to: ensure that the accuracy of the (i+1)th voiceprint similarity is not less than a preset value; or that both the i-th voiceprint similarity and the (i+1)th voiceprint similarity fall within a preset range.
[0092] Please refer to Figure 5 , Figure 5 The diagram shows a block diagram of a voiceprint recognition device provided in some embodiments of this application. It should be understood that the voiceprint recognition device corresponds to the method embodiments described above and is capable of performing the various steps involved in the method embodiments described above. The specific functions of the voiceprint recognition device can be found in the description above, and detailed descriptions are omitted here to avoid repetition.
[0093] Figure 5 The voiceprint recognition device includes at least one software function module that can be stored in a memory or embedded in the voiceprint recognition device in the form of software or firmware. The voiceprint recognition device includes: an acquisition module 510 for acquiring audio to be recognized; and an identification module 520 for inputting the audio to be recognized into a target voiceprint recognition model to obtain a voiceprint recognition result, wherein the voiceprint recognition result is used to confirm the number of target objects contained in the audio.
[0094] Some embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can perform the operation of any of the methods corresponding to the methods provided in the above embodiments.
[0095] Some embodiments of this application also provide a computer program product, which includes a computer program, wherein when the computer program is executed by a processor, it can implement the operation of any of the methods corresponding to the above embodiments provided in the above embodiments.
[0096] like Figure 6 As shown, some embodiments of this application provide an electronic device 600, which includes a memory 610, a processor 620, and a computer program stored in the memory 610 and executable on the processor 620. When the processor 620 reads the program from the memory 610 via a bus 630 and executes the program, it can implement the methods of any of the above embodiments.
[0097] Processor 620 can process digital signals and can include various computing architectures. For example, it can be a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements multiple instruction set combinations. In some examples, processor 620 can be a microprocessor.
[0098] The memory 610 can be used to store instructions executed by the processor 620 or data related to the execution of instructions. These instructions and / or data may include code for implementing some or all of the functions of one or more modules described in the embodiments of this application. The processor 620 of this disclosure embodiment can be used to execute the instructions in the memory 610 to implement the methods shown above. The memory 610 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memories well known to those skilled in the art.
[0099] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0100] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0101] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. A method for training a voiceprint recognition model, characterized in that, include: Obtain a training sample dataset, wherein the training sample dataset includes: multiple audio samples of multiple objects; The following process is repeated until at least the (i+1)th voiceprint similarity meets the preset condition, and the (i+1)th voiceprint recognition model is used as the target voiceprint recognition model: Based on the i-th voiceprint similarity and the training sample dataset, an (i+1)-th training sample set is obtained, wherein the i-th voiceprint similarity is obtained based on the i-th voiceprint recognition model and the audio of multiple speaking objects; the (i+1)-th training sample set includes: samples with a first proportion selected from samples with high similarity in the i-th voiceprint similarity and samples with a second proportion selected from the training sample dataset; the objective function of the i-th voiceprint recognition model is related to the i-th hyperparameter, the mean of sample similarity in the i-th training sample set, and the extreme value of sample similarity; the value of the i-th hyperparameter in each round of training ranges from 1 to close to 0; The i-th voiceprint recognition model is trained using the (i+1)-th training sample set to obtain the (i+1)-th voiceprint recognition model. Based on the (i+1)th voiceprint recognition model and the audio of the multiple speaking objects, the (i+1)th voiceprint similarity among the multiple speaking objects is obtained.
2. The method as described in claim 1, characterized in that, When i=1, the i-th voiceprint recognition model is obtained by the following method: The first training sample set is randomly selected from the training sample dataset to train the initial voiceprint recognition model, thereby obtaining the first voiceprint recognition model.
3. The method as described in claim 1 or 2, characterized in that, The first hyperparameter of the initial voiceprint recognition model is different from that of the i-th voiceprint recognition model.
4. The method as described in claim 3, characterized in that, The mean similarity of the samples is related to the mean similarity of voiceprint features of samples of the same type and the mean similarity of voiceprint features of samples of different types. The samples of the same type represent audio samples of the same object, and the samples of different types represent audio samples of different objects.
5. The method as described in claim 4, characterized in that, The extreme value of sample similarity is related to the i-th adjustment hyperparameter, the minimum similarity of voiceprint features of the same type of samples, and the maximum similarity of voiceprint features of the different type of samples.
6. The method as described in claim 1 or 2, characterized in that, The step of obtaining the (i+1)th training sample set based on the i-th voiceprint similarity and the training sample dataset includes: Randomly select the first percentage value; Based on the i-th voiceprint similarity, a first sample corresponding to the first proportion value is obtained, and a second sample corresponding to the second proportion value is selected from the training sample dataset, wherein the sum of the first proportion value and the second proportion value is 1, and the first sample and the second sample constitute the (i+1)-th training sample set.
7. The method as described in claim 1 or 2, characterized in that, The at least i+1th voiceprint similarity satisfies the preset conditions, including: The accuracy of the (i+1)th voiceprint similarity is not less than a preset value; or, Both the i-th voiceprint similarity and the (i+1)-th voiceprint similarity fall within a preset range.
8. A method for voiceprint recognition, characterized in that, include: Obtain the audio to be recognized; The audio to be identified is input into the target voiceprint recognition model obtained by the method according to any one of claims 1-7 to obtain a voiceprint recognition result, wherein the voiceprint recognition result is used to confirm the number of target objects contained in the audio.
9. An apparatus for training a voiceprint recognition model, characterized in that, The apparatus is used to perform the method as described in claim 1, comprising: The sample acquisition module is used to acquire a training sample dataset, wherein the training sample dataset includes: multiple audio samples of multiple objects; The model training module is used for: The following process is repeated until at least the (i+1)th voiceprint similarity meets the preset condition, and the (i+1)th voiceprint recognition model is used as the target voiceprint recognition model: Based on the i-th voiceprint similarity and the training sample dataset, the (i+1)-th training sample set is obtained, wherein the i-th voiceprint similarity is obtained based on the i-th voiceprint recognition model and the audio of multiple speaking objects; The i-th voiceprint recognition model is trained using the (i+1)-th training sample set to obtain the (i+1)-th voiceprint recognition model. Based on the (i+1)th voiceprint recognition model and the audio of the multiple speaking objects, the (i+1)th voiceprint similarity among the multiple speaking objects is obtained.
10. A voiceprint recognition device, characterized in that, include: The acquisition module is used to acquire the audio to be recognized; The recognition module is used to input the audio to be recognized into a target voiceprint recognition model obtained by the method of any one of claims 1-7, and obtain a voiceprint recognition result, wherein the voiceprint recognition result is used to confirm the number of target objects contained in the audio.
Citation Information
Patent Citations
Voiceprint recognition model training method and device and related equipment
CN112820299A
Open set identification method and device, electronic equipment, medium and program product
CN115222970A