A method for annotating voiceprint data based on a voiceprint model
Through the voiceprint data labeling method based on the voiceprint model, the initial model is trained using the marked data to identify and group annotate the unmarked data, which solves the problem of inefficient manual labeling and improves the efficiency and accuracy of the labeling.
Patent Information
- Application Number
- CN202111357126.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-16
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-11-16
AI Technical Summary
In deep learning-based voiceprint recognition technology, manual labeling of audio data is inefficient, which makes the data labeling process time-consuming and difficult to ensure accuracy.
A voiceprint data labeling method based on voiceprint model is proposed. By using the marked first voiceprint data to train the initial model, voiceprint recognition is performed on the unmarked second voiceprint data, voiceprint features are obtained, and identity information is grouped according to the characteristics.
The efficiency and accuracy of data annotation are improved, and the model is trained in a semi-supervised manner, taking into account the labeling speed and accuracy.
Smart Images

Figure CN114242077B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio recognition, and more specifically, to a method for annotating voiceprint data based on a voiceprint model. Background Art
[0002] Deep learning is the core technology in the field of artificial intelligence today. With the application and popularization of technologies based on deep learning, among which voiceprint recognition based on deep learning is one of the applications. Nowadays, voiceprint recognition based on deep learning has achieved rapid development and wide application. In voiceprint recognition based on deep learning, for the training of a voiceprint recognition model, a large amount of data and correct labels are particularly important. In related technologies, audio data is usually annotated by manual annotation. However, due to the extremely large amount of audio data to be annotated, the method of manual annotation is inefficient. Summary of the Invention
[0003] In view of the above problems, the present application proposes a method for annotating voiceprint data based on a voiceprint model.
[0004] An embodiment of the present application provides a method for annotating voiceprint data based on a voiceprint model. The method includes: obtaining a plurality of voiceprint data, where the plurality of voiceprint data includes a plurality of first voiceprint data already annotated with identity information, and a plurality of second voiceprint data not annotated with identity information; training an initial model based on the first voiceprint data to obtain an initial voiceprint recognition model; performing voiceprint recognition on the plurality of second voiceprint data based on the initial voiceprint recognition model to obtain the voiceprint feature corresponding to each second voiceprint data; based on the voiceprint feature corresponding to each second voiceprint data, obtaining multiple groups of voiceprint data existing in the plurality of second voiceprint data, and the second voiceprint data other than the multiple groups of voiceprint data in the plurality of second voiceprint data as other voiceprint data, where each second voiceprint data in each group of the multiple groups of voiceprint data belongs to the same user, and each group of voiceprint data includes at least two second voiceprint data; labeling the identity information corresponding to each group of voiceprint data in the multiple groups of voiceprint data as different identity information, and labeling each voiceprint data in the other voiceprint data as different identity information, where the identity information corresponding to the second voiceprint data in the same group of voiceprint data in the multiple groups of voiceprint data is the same, and the identity information corresponding to each group of voiceprint data is different from the identity information corresponding to each voiceprint data in the other voiceprint data.
[0005] The solution provided by this application trains an initial voiceprint recognition model by using multiple first voiceprint data marked with identity information to train an initial model, and uses this model to perform voiceprint recognition on multiple second voiceprint data without marked identity information, obtain corresponding voiceprint features, and group all the second voiceprint data according to the voiceprint features to obtain multiple groups of voiceprint data and other voiceprint data, and then label the multiple groups of voiceprint data and other voiceprint data with different identity information, which not only improves the efficiency of data annotation; moreover, using the manually marked voiceprint data to train the initial voiceprint recognition model to annotate the unmarked voiceprint data can also improve the accuracy of data annotation. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0007] Figure 1 FIG. shows a schematic flowchart of a method for annotating voiceprint data provided by an embodiment of this application.
[0008] Figure 2 FIG. shows a schematic flowchart of a method for annotating voiceprint data provided by another embodiment of this application.
[0009] Figure 3 FIG. shows a schematic flowchart of step S250 in a method for annotating voiceprint data provided by another embodiment of this application.
[0010] Figure 4 FIG. shows another schematic flowchart of a method for annotating voiceprint data provided by another embodiment of this application.
[0011] Figure 5 FIG. shows a schematic flowchart of a method for annotating voiceprint data provided by another embodiment of this application.
[0012] Figure 6 FIG. shows a schematic flowchart of a method for annotating voiceprint data provided by yet another embodiment of this application.
[0013] Figure 7 FIG. shows a schematic flowchart of a method for annotating voiceprint data provided by yet another embodiment of this application.
[0014] Figure 8 FIG. shows a structural block diagram of an apparatus for annotating voiceprint data provided by this application.
[0015] Figure 9A structural block diagram of a computer device provided by the present application is shown.
[0016] Figure 10 A structural block diagram of a computer-readable storage medium provided by the present application is shown. DETAILED DESCRIPTION
[0017] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0018] At present, with the application and promotion of deep learning technology, the field of artificial intelligence has also developed rapidly. Now deep learning technology is increasingly valued and applied in various fields. For example, in the field of voiceprint recognition, the neural network model can be trained through deep learning to extract voiceprint features from audio data, thereby achieving the purpose of voiceprint recognition.
[0019] In voiceprint recognition technology based on deep learning, in order to identify voiceprint data through a neural network model, it is first necessary to improve the accuracy of the neural network model, which requires training the model through deep learning with a large amount of audio data with correct labels. To train the neural network model with these audio data, one way is to correctly annotate the audio data through manual annotation, which has high accuracy but is often inefficient; another way is to automatically annotate through the model, which is fast but has low accuracy.
[0020] In response to the above problems, the inventors have proposed the voiceprint data labeling method, device, computer equipment and storage medium provided in the embodiments of the present application, which trains the model in a semi-supervised manner and labels the voiceprint data, thereby taking into account both the efficiency and accuracy of data labeling. The specific audio and video synchronization method will be described in detail in the subsequent embodiments.
[0021] The following will describe in detail the method for labeling voiceprint data based on a voiceprint model provided in an embodiment of the present application in conjunction with the accompanying drawings.
[0022] See also Figure 1 , Figure 1 A flow chart of a method for labeling voiceprint data based on a voiceprint model provided by an embodiment of the present application is shown. Figure 1 The process shown is described in detail, and the method for annotating voiceprint data based on the voiceprint model may specifically include the following steps:
[0023] Step S110: Acquire a plurality of voiceprint data, wherein the plurality of voiceprint data includes a plurality of first voiceprint data marked with identity information, and a plurality of second voiceprint data not marked with identity information.
[0024] In an embodiment of the present application, when training a voiceprint recognition model based on deep learning, multiple voiceprint data can be obtained first. After annotating the corresponding identity information for the voiceprint data, the annotated voiceprint data can be obtained, and then the annotated voiceprint data can be used for model training. These voiceprint data can include multiple voiceprint data that have been annotated with identity information, which are used as the first voiceprint data to train the voiceprint recognition model, and can also include multiple voiceprint data that have not been annotated with identity information, which are used as the second voiceprint data to annotate their identity information through the voiceprint recognition model. Through these multiple voiceprint data including the first voiceprint data and the second voiceprint data, semi-supervised annotation of voiceprint data can be achieved. Among them, the multiple first voiceprint data with annotated identity information can be obtained by the user manually annotating some of the voiceprint data after obtaining the original multiple voiceprint data.
[0025] In some embodiments, the acquisition source of the voiceprint data can be audio data obtained from an audio-video application software, or audio data independently uploaded by the user through a mobile terminal, or audio data actively acquired by an audio acquisition device, or audio data obtained from a server, which is not limited herein. These voiceprint data can contain parameters that can characterize the voice personality characteristics of the data subject, such as information at various levels such as spectrum, pitch, and tone.
[0026] In some embodiments, the identity information annotated for the voiceprint data can be the identity identification number (Identity document, ID) corresponding to the user (i.e., the speaker) who generated the voiceprint data, or a label code representing a person's identity, etc.
[0027] Step S120: Based on the first voiceprint data, train the initial model to obtain an initial voiceprint recognition model.
[0028] In the embodiments of the present application, after obtaining a plurality of voiceprint data, the initial model can be trained based on the first voiceprint data among the plurality of voiceprint data, so as to obtain an initial recognition model, and the second voiceprint data without identity information marked can be recognized by the obtained initial voiceprint recognition model. The initial model can first preprocess the input first voiceprint data with labels, that is, recognize the voiceprint features of the plurality of first voiceprint data, and then the discriminator classifies the voiceprint features output by the initial model; and the comparison error can be obtained by comparing the classification result with the label of the voiceprint data, and then the relevant parameters of the initial model can be adjusted according to the comparison error; then the initial model with adjusted parameters is used to recognize the voiceprint features of the first voiceprint data again until the final comparison error is less than the preset error, and the relevant parameters of the final initial model can be used as determined values. At this time, this initial model can be used as the initial voiceprint recognition model to recognize the second voiceprint data without identity information. That is to say, using the identity information marked by the first voiceprint data to supervise the training process in model training, so as to train the voiceprint recognition model.
[0029] In some embodiments, the initial model can be a machine learning model or a deep learning model, such as an hourglass network, an autoencoder network, etc.; for another example, to facilitate the deployment of the voiceprint recognition model on electronic devices such as mobile terminals, the initial model can be a model based on the Unet network, which is not limited here.
[0030] S130: Recognize the voiceprint of the plurality of second voiceprint data based on the initial voiceprint recognition model to obtain the voiceprint features corresponding to each second voiceprint data.
[0031] In the embodiments of the present application, after an initial voiceprint recognition model is trained through first voiceprint data, voiceprint recognition can be performed on multiple second voiceprint data without identity information marked based on this model to obtain the voiceprint features corresponding to each second voiceprint data, so as to perform subsequent grouping and annotation operations on the voiceprint data through the voiceprint features. Since the above initial voiceprint recognition model is trained based on voiceprint data accurately marked with identity information, it has a certain degree of accuracy. At this time, the voiceprint features recognized through this initial voiceprint recognition model can be used to annotate the voiceprint data without identity information, which can be used to obtain a voiceprint annotation with a relatively high accuracy rate. Voiceprint recognition is a type of biometric recognition technology, which refers to a technology for discriminating the identity of a speaker through voice. Since the slight differences in the vocal organs of each person will cause changes in the vocal airflow, resulting in differences in timbre and tone quality, the voiceprint maps of any two people are different. Therefore, the voiceprint features in the voiceprint data can be extracted through a voiceprint recognition model to determine the identity of the speaker. Among them, the voiceprint features in the voiceprint data can be the frequency values and their trends of formants, and this feature has strong specificity. The voiceprint features can also select features such as duration, intensity, and waveform, but due to poor stability, they can be used as references.
[0032] In some embodiments, at this time, when the initial voiceprint recognition model performs voiceprint recognition on the second voiceprint data, it can be to input the second voiceprint data into the initial voiceprint recognition model, and the initial voiceprint recognition model can recognize the input second voiceprint data to obtain the voiceprint features corresponding to the second voiceprint data. Since the initial voiceprint recognition model is a model obtained through the previous supervised learning training, the recognition accuracy rate of this model is relatively high, and the judgment made by it can be used as the basis for subsequent annotation of the voiceprint data. Since the annotation of the second voiceprint data is based on the voiceprint recognition by the second voiceprint recognition model trained through the already annotated first voiceprint data, and subsequent automatic annotation is performed based on the features of the voiceprint recognition, the annotation of the second voiceprint data can be understood as semi-supervised data annotation, thereby improving the accuracy of data annotation.
[0033] S140: Based on the voiceprint features corresponding to each of the second voiceprint data, obtain multiple groups of voiceprint data existing in the multiple second voiceprint data, and the second voiceprint data in the multiple second voiceprint data other than the multiple groups of voiceprint data as other voiceprint data. Each second voiceprint data in each group of the multiple groups of voiceprint data belongs to the same user, and each group of voiceprint data includes at least two second voiceprint data.
[0034] In an embodiment of the present application, after obtaining the voiceprint features corresponding to each second voiceprint data, based on these voiceprint features, the second voiceprint data with matching voiceprint features can be divided into the same group, thereby obtaining multiple groups of voiceprint data existing in the multiple second voiceprint data, and some individual voiceprint data that cannot be divided into the same group as other voiceprint data. Then, all unlabeled second voiceprint data are divided into multiple groups of voiceprint data and other voice data according to their respective voiceprint features. After the second voiceprint data is grouped, each group of voiceprint data in the multiple groups of voiceprint data includes two or more second voiceprint data, and the voiceprint features corresponding to each voiceprint data in each group of voiceprint data match. That is to say, for each group of voiceprint data, the second voiceprint data in this group all belong to the same user; other voiceprint data are second voiceprint data whose voiceprint features do not match the voiceprint features of all other second voiceprint data except itself. That is to say, each second voiceprint data in other voiceprint data belongs to a different user, and there is only one second voiceprint data for these users. Thus, according to the characteristics of the voiceprint features, the second voiceprint data is divided into multiple groups of voiceprint data and other voiceprint data, which can be used to label the identities of these different groups of voiceprint data and other voiceprint data.
[0035] Step S150: Label the identity information corresponding to each group of voiceprint data in the multiple groups of voiceprint data with different identity information, and label each voiceprint data in the other voiceprint data with different identity information, where the identity information corresponding to the second voiceprint data in the same group of voiceprint data in the multiple groups of voiceprint data is the same, and the identity information corresponding to each group of voiceprint data is different from the identity information corresponding to each voiceprint data in the other voiceprint data.
[0036] In an embodiment of the present application, after grouping all the second voiceprint data according to their respective voiceprint features, each voiceprint data in the other voiceprint data can be regarded as a group of voiceprint data. Different from the multiple groups of voiceprint data containing multiple second voiceprint data, at this time, the voiceprint features corresponding to the second voiceprint data in different groups are different from each other, while the voiceprint features of the second voiceprint data within the same group are the same. On this basis, label the identity information corresponding to each group of voiceprint data with different identity information, where in the multiple groups of voiceprint data containing multiple second voiceprint data, the identity label corresponding to each second voiceprint data in each group of voiceprint data is the same as the identity label corresponding to the group where it is located.
[0037] Understandably, the above-mentioned multiple sets of voiceprint data can be understood as at least two second voiceprint data corresponding to different users, and the other voiceprint data can be understood as one second voiceprint data corresponding to different users. Moreover, the multiple users corresponding to the multiple sets of voiceprint data are different from the multiple users corresponding to the other voiceprint data. Based on this, different identity information can be marked for each set of voiceprint data in the multiple sets of voiceprint data and each second voiceprint data in the other voiceprint data. However, since all the second voiceprint data in each set of voiceprint data belong to the same user, the identity information corresponding to all the second voiceprint data in each set of voiceprint data is the same.
[0038] It should be noted that the specific identity information marked for the multiple sets of voiceprint data and the other voiceprint data may not be limited. It only needs to ensure that different identity information is marked for each set of voiceprint data in the multiple sets of voiceprint data and each second voiceprint data in the other voiceprint data, and the identity information marked for the second voiceprint data in each set of voiceprint data is the same.
[0039] The voiceprint data annotation method provided by the embodiment of the present application trains an initial voiceprint recognition model by using multiple first voiceprint data that have been marked with identity information on an initial model, and performs voiceprint recognition on multiple second voiceprint data that have not been marked with identity information through this model to obtain corresponding voiceprint features, and groups all the second voiceprint data according to the voiceprint features to obtain multiple sets of voiceprint data and other voiceprint data, and then marks the multiple sets of voiceprint data and the other voiceprint data with different identity information. Since the annotation of the second voiceprint data is based on the voiceprint recognition by the second voiceprint recognition model trained by the marked first voiceprint data, and subsequent automatic annotation is performed based on the features of the voiceprint recognition, the annotation of the second voiceprint data can be understood as semi-supervised data annotation, thus improving both the efficiency and the accuracy of data annotation.
[0040] Please refer to Figure 2 , Figure 2 which shows a schematic flowchart of a voiceprint data annotation method based on a voiceprint model provided by another embodiment of the present application. The following will elaborate in detail on Figure 2 the process shown. The voiceprint data annotation method based on the voiceprint model may specifically include the following steps:
[0041] Step S210: Obtain multiple voiceprint data, where the multiple voiceprint data include multiple first voiceprint data that have been marked with identity information and multiple second voiceprint data that have not been marked with identity information.
[0042] Step S220: Train an initial model based on the first voiceprint data to obtain an initial voiceprint recognition model.
[0043] Step S230: Perform voiceprint recognition on the multiple second voiceprint data based on the initial voiceprint recognition model to obtain the voiceprint features corresponding to each second voiceprint data.
[0044] Step S240: Based on the voiceprint features corresponding to each second voiceprint data, obtain the similarity between every two second voiceprint data among the multiple second voiceprint data.
[0045] In the embodiment of the present application, after obtaining the voiceprint features corresponding to each second voiceprint data, the similarity between every two second voiceprint data among the multiple second voiceprint data can be obtained, so as to quantify the similarity degree between the voiceprint features corresponding to any two second voiceprint data through the magnitude of the similarity value. It can be understood that the higher the similarity value corresponding to two second voiceprint data, the more similar the voiceprint features corresponding to these two second voiceprint data.
[0046] In some embodiments, to obtain the voiceprint features corresponding to every two second voiceprint data, each second voiceprint data can be marked with a different number for distinction. After obtaining the similarity between every two second voiceprint data, the magnitude of this similarity is associated with the numbers marked for these two second voiceprint data, so as to quickly find the similarity between each second voiceprint data and another second voiceprint data.
[0047] In some embodiments, the above voiceprint features can be feature vectors, and the Euclidean distance or cosine distance between the feature vectors can be obtained, so as to obtain the similarity between the voiceprint features. The specific manner of obtaining the similarity between the voiceprint features may not be limited.
[0048] Step S250: Based on the similarity between every two second voiceprint data among the multiple second voiceprint data, obtain the second voiceprint data belonging to the same user from the multiple second voiceprint data, and use the second voiceprint data belonging to the same user as a group of voiceprint data to obtain multiple groups of voiceprint data.
[0049] In the embodiment of the present application, if the similarity between two second voiceprint data is relatively high, these two second voiceprint data can be regarded as the voiceprint data belonging to the same user. Thus, after obtaining the similarity between any two second voiceprint data among the multiple second voiceprint data, based on the similarity between different voiceprint data, all the second voiceprint data belonging to the same user can be determined, and all these second voiceprint data belonging to the same user are divided into a group of voiceprint data. Thus, multiple groups of voiceprint data belonging to different users can be obtained from the multiple second voiceprint data.
[0050] In some embodiments, the multiple second voiceprint data may further include second voiceprint data with relatively low similarity to all the remaining second voiceprint data, which means that another second voiceprint data belonging to the same user cannot be found among these multiple second voiceprint data. At this time, this second voiceprint data can be taken as a separate group, that is, this group of voiceprint data contains only one second voiceprint data, and the user to whom this second voiceprint data belongs is different from the users to whom the other multiple second voiceprint data belong.
[0051] In some embodiments, such as Figure 3 shown, based on the similarity between every two of the multiple second voiceprint data, obtaining the second voiceprint data belonging to the same user from the multiple second voiceprint data may further include the following steps:
[0052] Step S251: Determine whether the similarity between every two of the second voiceprint data is greater than a preset threshold.
[0053] In this embodiment, after obtaining the similarity between every two of the multiple second voiceprint data, it can be determined whether the similarity between every two of the second voiceprint data is greater than a preset threshold. The preset threshold is a value obtained based on the accuracy of the initial voiceprint recognition model trained by multiple first voiceprint data. The value range of the preset threshold can be the same as the value range of the similarity, so as to determine whether the two second voiceprint data corresponding to the similarity belong to the same user based on the size relationship between the similarity and the preset threshold.
[0054] In some embodiments, the preset threshold can be modified by user operation. After the user manually checks the second voiceprint data and the determined belonging users, the preset threshold can be appropriately lowered or increased to adjust the final output judgment.
[0055] In some embodiments, if the similarity between any two second voiceprint data is less than or equal to the preset threshold, it can be determined that these two second voiceprint data do not belong to the same user.
[0056] Step S252: If the similarity between any two target voiceprint data is greater than the preset threshold, determine the two target voiceprint data as the second voiceprint data belonging to the same user, and the target voiceprint data is any one of the multiple second voiceprint data.
[0057] In the embodiments of the present application, if the similarity between any two of the multiple second voiceprint data is greater than the preset threshold, it means that the overlapping degree of the voiceprint features of these two second voiceprint data is relatively high. Then, these two second voiceprint data can be used as two target voiceprint data, and these two target voiceprint data can be determined as the second voiceprint data belonging to the same user.
[0058] In some embodiments, after determining any two target voiceprint data with a similarity greater than a preset threshold as the second voiceprint data belonging to the same user, these two target voiceprint data can be divided into a group of voiceprint data. After comparing the similarity of any other two pieces of second voiceprint data with the preset threshold, all the second voiceprint data determined to belong to the same user are divided into the same group of voiceprint data, so as to label the identity information of each second voiceprint data in each group of voiceprint data.
[0059] Step S260: Label the identity information corresponding to each group of voiceprint data in the multiple groups of voiceprint data as different identity information, and label each voiceprint data in the other voiceprint data as different identity information, where the identity information corresponding to the second voiceprint data in the same group of voiceprint data in the multiple groups of voiceprint data is the same, and the identity information corresponding to each group of voiceprint data is different from the identity information corresponding to each voiceprint data in the other voiceprint data.
[0060] In the embodiments of the present application, steps S210, S220, S230, and S260 can refer to the content of other embodiments and will not be elaborated herein.
[0061] The method for labeling voiceprint data provided by the embodiments of the present application details the process of obtaining the above-mentioned multiple groups of voiceprint data existing in the multiple second voiceprint data based on the voiceprint features corresponding to the second voiceprint data, and realizes that after voiceprint recognition using the second voiceprint recognition model trained based on the labeled first voiceprint data, based on the similarity between the recognized voiceprint features, the voiceprint features of the same user are determined, which not only improves the efficiency of data labeling but also improves the accuracy of data labeling.
[0062] Please refer to Figure 4 , Figure 4 which shows a schematic flowchart of a method for labeling voiceprint data based on a voiceprint model provided by another embodiment of the present application. The following will elaborate in detail on Figure 4 the process shown. The method for labeling voiceprint data based on a voiceprint model may specifically include the following steps:
[0063] Step S301: Obtain multiple voiceprint data, where the multiple voiceprint data include multiple first voiceprint data with labeled identity information and multiple second voiceprint data without labeled identity information.
[0064] Step S302: Train an initial model based on the first voiceprint data to obtain an initial voiceprint recognition model.
[0065] Step S303: Perform voiceprint recognition on the multiple second voiceprint data based on the initial voiceprint recognition model to obtain the voiceprint features corresponding to each second voiceprint data.
[0066] Step S304: Based on the voiceprint features corresponding to each of the second voiceprint data, obtain the similarity between every two of the multiple second voiceprint data.
[0067] Step S305: Determine whether the similarity between every two of the second voiceprint data is greater than a preset threshold.
[0068] Step S306: If the similarity between any two target voiceprint data is greater than the preset threshold, then determine the two target voiceprint data as the second voiceprint data belonging to the same user, and use the second voiceprint data belonging to the same user as a group of voiceprint data to obtain multiple groups of voiceprint data, where the target voiceprint data is any one of the multiple second voiceprint data.
[0069] Step S307: Randomly extract a preset number of groups of voiceprint data from the multiple groups of voiceprint data.
[0070] In the embodiment of the present application, the voiceprint features corresponding to multiple second voiceprint data are obtained through the initial voiceprint recognition model, and then the multiple second voiceprint data are divided into multiple groups of voiceprint data belonging to different users. However, since the initial voiceprint recognition model does not extract the voiceprint features of the voiceprint data completely accurately, in order to improve the accuracy of grouping the multiple second voiceprint data, a preset number of groups of voiceprint data can be randomly extracted from the multiple groups of voiceprint data to determine whether the second voiceprint data included in each group of the extracted voiceprint data belongs to the same user through manual verification.
[0071] In some embodiments, in step S307, randomly extracting a preset number of groups of voiceprint data from the multiple groups of voiceprint data may include:
[0072] Obtain the number of second voiceprint data in each group of voiceprint data; according to the number of the second voiceprint data, adjust the weight of each group of voiceprint data, and the weight is positively correlated with the number of the second voiceprint data. Specifically, when randomly extracting a preset number of groups of voiceprint data from the multiple groups of voiceprint data, the weight of each group of voiceprint data can be increased or decreased according to the number of second voiceprint data in each group of voiceprint data to change the probability of each group of voiceprint data being randomly extracted, that is, the more the number of second voiceprint data included in a group of voiceprint data, the greater the probability of this group of voiceprint data being randomly extracted. Since the more the number of second voiceprint data included in a group of voiceprint data, the greater the probability that there are second voiceprint data that do not belong to the same user, the weight of this group of voiceprint data can be increased, so that this group of voiceprint data has a greater probability of being randomly extracted.
[0073] Step S308: Obtain the test result of the user's test on the voiceprint data of the preset number of groups. The test result is used to characterize whether the second voiceprint data in each group of voiceprint data belongs to the same user and whether the second voiceprint data in different groups of voiceprint data does not belong to the same user.
[0074] In the embodiment of the present application, the user can test the voiceprint features corresponding to the second voiceprint data in each randomly selected group of voiceprint data, and determine whether the voiceprint features corresponding to the second voiceprint data in each group of voiceprint data are the same. As the test result, it can be used to determine whether the second voiceprint data in each group of voiceprint data belongs to the same user. If the user tests that the voiceprint features corresponding to the second voiceprint data in each randomly selected group of voiceprint data are the same, it can be determined that the classification of this group is correct, and the second voiceprint data in this group of voiceprint data all belong to the same user; correspondingly, if the user tests that there is a voiceprint feature corresponding to a certain second voiceprint data in a randomly selected group of voiceprint data that is different from the voiceprint features corresponding to other second voiceprint data, it means that the second voiceprint data in this group of voiceprint data does not belong to the same user.
[0075] In some embodiments, the test result of the user's test on the voiceprint data of the preset number of groups can include not only testing whether the second voiceprint data in each group of voiceprint data all belong to the same user, but also testing whether the second voiceprint data belonging to different groups of voiceprint data among multiple groups of voiceprint data belong to the same user. Therefore, after randomly selecting the voiceprint data of the preset number of groups, the user can not only test whether the voiceprint features corresponding to all the second voiceprint data in the same group of voiceprint data belong to the same user, but also test whether the voiceprint features corresponding to the second voiceprint data between different groups of voiceprint data do not belong to the same user.
[0076] Step S309: Adjust the preset threshold according to the test result.
[0077] In the embodiment of the present application, after judging the test result of the randomly selected voiceprint data, the preset threshold can be adjusted according to the meaning represented by the test result, that is, if the test result shows that the grouping situation of the second voiceprint data in each group of voiceprint data is not the same as the original grouping, it means that the preset threshold cannot accurately divide the second voiceprint data belonging to the same user into the same group of voiceprint data. Therefore, the preset threshold can be adjusted to update the multiple groups of voiceprint data and other voiceprint data.
[0078] In some embodiments, the operation of adjusting the preset threshold according to the test result in step S309 can have the following different situations:
[0079] If the verification result of any set of voiceprint data indicates that the second voiceprint data it includes does not belong to the same user, increase the preset threshold. That is, after the user verifies a preset number of sets of voiceprint data, if there is a set of voiceprint data in which the second voiceprint data does not belong to the same user, that is, the verification result shows that there is a voiceprint feature corresponding to a certain second voiceprint data in a set of voiceprint data that is different from the voiceprint features corresponding to other second voiceprint data. This is because the preset threshold is set too low, making the similarity of two voiceprint data that actually do not belong to the same user greater than the preset threshold, resulting in these two voiceprint data being determined to belong to the same user. Therefore, by increasing the preset threshold, the voiceprint data that does not belong to the same user in each set of voiceprints can be excluded.
[0080] If the verification results of different sets of voiceprint data indicate that any two second voiceprint data belonging to different sets belong to the same user, decrease the preset threshold. That is, the user's verification result shows that there are some second voiceprint data belonging to the same user among different sets, which means that the preset threshold is relatively high at this time. The preset threshold can be lowered, and after updating the preset threshold to the lowered one, determine again whether the similarity of any two voiceprint data is greater than the preset threshold, so as to update the division of multiple sets of voiceprint data and other voiceprint data.
[0081] Step S310: Update the multiple sets of voiceprint data and the other voiceprint data based on the adjusted preset threshold.
[0082] In the embodiment of the present application, if it is found in the user's verification that there is second voiceprint data in a set of voiceprint data that does not belong to the same user, after increasing the preset threshold, the multiple sets of voiceprint data and the other voiceprint data can be updated, that is, update the preset threshold to the increased one, and based on the new preset threshold, determine again the size relationship between the similarity of every two second voiceprint data and the preset threshold. If the similarity of any two second voiceprint data is greater than the preset threshold, determine these two second voiceprint data as the second voiceprint data belonging to the same user.
[0083] Step S311: Label the identity information corresponding to each set of voiceprint data in the multiple sets of voiceprint data as different identity information, and label each voiceprint data in the other voiceprint data as different identity information, where the identity information corresponding to the second voiceprint data in the same set of voiceprint data in the multiple sets of voiceprint data is the same, and the identity information corresponding to each set of voiceprint data is different from the identity information corresponding to each voiceprint data in the other voiceprint data.
[0084] In the embodiment of the present application, the content of steps S301 to S306 and step S311 can refer to the content of other embodiments and will not be elaborated here.
[0085] The annotation method for voiceprint data based on a voiceprint model provided by an embodiment of the present application, after dividing multiple second voiceprint data into multiple groups of voiceprint data and other voiceprint data, randomly selects a preset number of groups of voiceprint data by the user for verification, and updates the preset threshold according to the verification result, and then updates the multiple groups of voiceprint data and other voiceprint data, which can correct the possible errors in the identity information annotation of multiple second voiceprint data based on the initial voiceprint recognition model, and improve the accuracy of the annotation of voiceprint data.
[0086] Please refer to Figure 5 , Figure 5 which shows a schematic flowchart of an annotation method for voiceprint data based on a voiceprint model provided by another embodiment of the present application. The following will elaborate in detail on Figure 5 the process shown. The annotation method for voiceprint data based on a voiceprint model may specifically include the following steps:
[0087] Step S401: Obtain multiple voiceprint data, where the multiple voiceprint data includes multiple first voiceprint data with identity information already annotated, and multiple second voiceprint data without identity information annotated.
[0088] Step S402: Train an initial model based on the first voiceprint data to obtain an initial voiceprint recognition model.
[0089] Step S403: Perform voiceprint recognition on the multiple second voiceprint data based on the initial voiceprint recognition model to obtain the voiceprint features corresponding to each second voiceprint data.
[0090] Step S404: Based on the voiceprint features corresponding to each second voiceprint data, obtain the similarity between every two second voiceprint data among the multiple second voiceprint data.
[0091] Step S405: Determine whether the similarity between every two second voiceprint data is greater than a preset threshold.
[0092] Step S406: If the similarity between any two target voiceprint data is greater than the preset threshold, then determine the two target voiceprint data as second voiceprint data belonging to the same user, where the target voiceprint data is any voiceprint data among the multiple second voiceprint data.
[0093] Step S407: Label the identity information corresponding to each group of voiceprint data in the multiple groups of voiceprint data as different identity information, and label each voiceprint data in the other voiceprint data as different identity information, where the identity information corresponding to the second voiceprint data in the same group of voiceprint data in the multiple groups of voiceprint data is the same, and the identity information corresponding to each group of voiceprint data is different from the identity information corresponding to each voiceprint data in the other voiceprint data.
[0094] In the embodiments of the present application, steps S401 to S407 may refer to the content of other embodiments and will not be elaborated herein.
[0095] Step S408: Obtain a target similarity in the similarities, where the absolute value of the difference between the target similarity and the preset threshold is less than a target value.
[0096] In the embodiments of the present application, after labeling the identity information corresponding to multiple groups of voiceprint data and other voiceprint data as different identity information, a target similarity in the similarities can be obtained. The absolute value of the difference between the target similarity and the preset threshold is less than the target value, that is, the magnitude of the target similarity can be between a first value and a second value. The first value is the value obtained by adding the target value to the preset threshold, and the second value is the value obtained by subtracting the target value from the preset threshold. Thus, the identity information corresponding to the second voiceprint data corresponding to the target similarity is reconfirmed. Based on the target value and the preset threshold, the magnitude and quantity of the target similarity can be obtained. The target value can be set in advance and is used to limit the quantity of target similarities existing in the interval. If the target value is large, the quantity of target similarities is correspondingly large; if the target value is small, the quantity of target similarities is correspondingly small.
[0097] Step S409: Obtain the second voiceprint data corresponding to the target similarity and the second voiceprint data having the same identity information as the second voiceprint data corresponding to the target similarity as the voiceprint data to be determined, and remove the identity information corresponding to the voiceprint data to be determined.
[0098] In the embodiments of the present application, if the magnitude of the similarity between any two second voiceprint data is the same as the obtained target similarity, then these two second voiceprint data are used as the voiceprint data to be determined. That is, the similarity corresponding to the voiceprint data to be determined is close to the preset threshold, and its corresponding identity information is not clear enough. Therefore, the identity information corresponding to the voiceprint data to be determined can be removed for performing operations such as voiceprint recognition on the voiceprint data to be determined again and relabeling the identity information.
[0099] Step S410: Obtain the voiceprint data other than the voiceprint data to be determined among the multiple voiceprint data as the determined voiceprint data.
[0100] In an embodiment of the present application, after obtaining the to-be-determined voiceprint data based on the target similarity, other second voiceprint data among the multiple second voiceprint data except the to-be-determined voiceprint data can be used as the determined voiceprint data to perform transfer training on the initial voiceprint recognition model through the determined voiceprint data to obtain a new training model. The determined voiceprint data can include multiple groups of voiceprint data and other voiceprint data, and the similarity between any two second voiceprint data in each group of voiceprint data is greater than the target similarity, and the similarity between any two second voiceprint data in different groups is less than the target similarity.
[0101] Step S411: Based on the determined voiceprint data and its labeled identity information, perform transfer training on the initial voiceprint recognition model to obtain a new voiceprint recognition model, update the initial voiceprint recognition model to the new voiceprint recognition model, and update the multiple second voiceprint data to the to-be-determined voiceprint data.
[0102] In an implementation manner of the present application, after obtaining the determined voiceprint data, transfer training can be performed on the initial voiceprint recognition model based on the determined voiceprint data and its labeled identity information, and a new voiceprint recognition model can be obtained. The new voiceprint recognition model has a higher recognition accuracy for the input voiceprint data. Update the initial voiceprint recognition model to the new voiceprint recognition model, and update the multiple second voiceprint data to the to-be-determined voiceprint data. At this time, the accuracy of the initial voiceprint recognition model is higher, and it can be used to perform voiceprint recognition on the multiple second voiceprint data through the initial voiceprint recognition model.
[0103] In some implementation manners, the method of transfer training can be to transfer samples, or to transfer parameters or models, which is not limited herein.
[0104] Step S412: Repeat the step of performing voiceprint recognition on the multiple second voiceprint data based on the initial voiceprint recognition model to obtain the voiceprint features corresponding to each second voiceprint data until the identity information corresponding to each group of voiceprint data in the multiple groups of voiceprint data is labeled as different identity information, and each voiceprint data in the other voiceprint data is labeled as different identity information, where the identity information corresponding to the to-be-determined voiceprint data is different from the identity information corresponding to the determined voiceprint data.
[0105] In the embodiments of the present application, after performing transfer training on the initial voiceprint recognition model with the determined voiceprint data and updating the model and multiple second voiceprint data, the steps in the embodiments of the present application from performing voiceprint recognition on multiple second voiceprint data based on the initial voiceprint recognition model to labeling the identity information corresponding to each voiceprint data in the multiple groups of voiceprint data as different identity information can be repeated to again recognize the voiceprint features of multiple second voiceprint data based on the initial voiceprint recognition model, obtain multiple groups of voiceprint data and other voiceprint data based on the relationship between the similarity of every two second voiceprint data and the preset threshold, and label different identity information for the multiple groups of voiceprint data and other voiceprint data, so as to realize re-labeling of the identity information of the to-be-determined voiceprint data above.
[0106] Among them, the identity information labeled for the multiple groups of voiceprint data and other voiceprint data during the execution of the repeated steps is different from the identity information labeled for the multiple groups of voiceprint data and other voiceprint data during the first execution of the steps, and can be used to label the corresponding identity information for multiple second voiceprint data through the above repeated steps, further improving the labeling accuracy of the identity information corresponding to multiple second voiceprint data.
[0107] The method for labeling voiceprint data provided by the embodiments of the present application takes the second voiceprint data corresponding to the target similarity as the to-be-determined voiceprint data and removes the corresponding identity information, uses the remaining second voiceprint data among the multiple second voiceprint data as the determined voiceprint data to perform transfer training on the initial voiceprint recognition model and update it, updates the multiple second voiceprint data to the to-be-determined voiceprint data, repeats steps such as performing voiceprint recognition on multiple second voiceprint data with the initial voiceprint recognition model, and re-labels the to-be-determined voiceprint data, which can further improve the accuracy of labeling of the second voiceprint data.
[0108] Please refer to Figure 6 , Figure 6 which shows a schematic flowchart of a method for labeling voiceprint data based on a voiceprint model provided by another embodiment of the present application. The following will elaborate in detail on the Figure 6 shown process. The method for labeling voiceprint data based on a voiceprint model may specifically include the following steps:
[0109] Step S510: Obtain multiple voiceprint data, where the multiple voiceprint data includes multiple first voiceprint data with identity information labeled thereon and multiple second voiceprint data without identity information labeled thereon.
[0110] Step S520: Train an initial model based on the first voiceprint data to obtain an initial voiceprint recognition model.
[0111] Step S530: Perform voiceprint recognition on the multiple second voiceprint data based on the initial voiceprint recognition model to obtain the voiceprint features corresponding to each second voiceprint data.
[0112] Step S540: Cluster the multiple second voiceprint data based on the voiceprint features corresponding to each second voiceprint data to obtain multiple categories.
[0113] In an embodiment of the present application, after obtaining the voiceprint features corresponding to the multiple second voiceprint data through the initial voiceprint recognition model, the multiple second voiceprint data can be clustered based on the voiceprint features to obtain multiple categories. The clustering method used can be a partitioning method or a hierarchical method, etc., which is not limited herein. For example, in the partitioning clustering algorithm, a preset threshold can be used as the center of clustering, and the distance can be regarded as the absolute value of the difference between the similarity and the preset threshold. Based on the clustering algorithm, the multiple second voiceprint data can be divided into multiple categories for labeling the identity information corresponding to the second voiceprint data in different categories. Among them, each category in the multiple categories can include at least two second voiceprint data.
[0114] Step S550: Take the second voiceprint data in each category as a group of voiceprint data to obtain multiple groups of voiceprint data.
[0115] In an implementation manner of the present application, after obtaining multiple categories based on the clustering algorithm, each category can include multiple second voiceprint data, and then the second voiceprint data in each category can be taken as a group of voiceprint data to obtain multiple groups of voiceprint data. Each group of voiceprint data can include multiple second voiceprint data.
[0116] In some implementation manners, after obtaining multiple categories based on the clustering algorithm, any two second voiceprint data with a difference in the distance from the center of clustering less than the target difference and in different categories can be used as the voiceprint data to be determined, and the remaining second voiceprint data in the multiple second voiceprint data can be used as the determined voiceprint data. The initial voiceprint recognition model can be migrated and trained and updated based on the determined voiceprint data. Then, the voiceprint data model after update can be used to re-divide the voiceprint data to be determined into multiple categories again, and corresponding identity information can be labeled based on the multiple categories obtained after the re-division. Moreover, the identity information labeled in the repeated execution steps is different from the identity information labeled in the first execution step.
[0117] Step S560: Label the identity information corresponding to each set of voiceprint data in the multiple sets of voiceprint data as different identity information, and label each voiceprint data in the other voiceprint data as different identity information. Among them, the identity information corresponding to the second voiceprint data in the same set of voiceprint data in the multiple sets of voiceprint data is the same, and the identity information corresponding to each set of voiceprint data is different from the identity information corresponding to each voiceprint data in the other voiceprint data.
[0118] In the embodiments of the present application, steps S510, S520, S530, and S560 can refer to the content of other embodiments and will not be elaborated here.
[0119] The method for labeling voiceprint data based on a voiceprint model provided by the embodiments of the present application divides multiple unlabeled second voiceprint data into different categories by clustering, and uses the second voiceprint data in each category as a set of voiceprint data to obtain multiple sets of voiceprint data, and labels the identity information corresponding to each set of voiceprint data as different identity information. By grouping and labeling the identity information of multiple second voiceprint data through different implementation methods, the voiceprint data without labeled identity information can be labeled quickly and accurately.
[0120] Please refer to Figure 7 , Figure 7 which shows a schematic flowchart of a method for labeling voiceprint data based on a voiceprint model provided by still another embodiment of the present application. The following will elaborate in detail on the Figure 7 shown process. The method for labeling voiceprint data based on a voiceprint model may specifically include the following steps:
[0121] Step S610: Obtain multiple voiceprint data, where the multiple voiceprint data includes multiple first voiceprint data with labeled identity information and multiple second voiceprint data without labeled identity information.
[0122] Step S620: Train an initial model based on the first voiceprint data to obtain an initial voiceprint recognition model.
[0123] Step S630: Perform voiceprint recognition on the multiple second voiceprint data based on the initial voiceprint recognition model to obtain the voiceprint features corresponding to each second voiceprint data.
[0124] Step S640: Based on the voiceprint features corresponding to each second voiceprint data, obtain multiple sets of voiceprint data existing in the multiple second voiceprint data, and the second voiceprint data other than the multiple sets of voiceprint data in the multiple second voiceprint data as other voiceprint data. Each second voiceprint data in each set of voiceprint data belongs to the same user, and each set of voiceprint data includes at least two second voiceprint data.
[0125] Step S650: Label the identity information corresponding to each set of voiceprint data in the multiple sets of voiceprint data as different identity information, and label each voiceprint data in the other voiceprint data as different identity information, where the identity information corresponding to the second voiceprint data in the same set of voiceprint data in the multiple sets of voiceprint data is the same, and the identity information corresponding to each set of voiceprint data is different from the identity information corresponding to each voiceprint data in the other voiceprint data.
[0126] In the embodiments of the present application, steps S610 to S650 may refer to the content of other embodiments and will not be elaborated here.
[0127] Step S660: Train the target model according to the voiceprint data marked with identity information to obtain a trained voiceprint recognition model.
[0128] In the embodiments of the present application, after the identity information of the voiceprint data without identity information is marked through the above steps, the target model can be trained in a supervised learning manner according to the voiceprint data marked with identity information, and a trained voiceprint recognition model can be obtained. The accuracy rate of this voiceprint recognition model is higher than that of the target model before training.
[0129] In some embodiments, the target model may be an untrained initial model or a model trained through the above steps, which is not limited here.
[0130] The method for labeling voiceprint data based on a voiceprint model provided by the present application uses a plurality of first voiceprint data partially marked with identity information and a plurality of second voiceprint data partially not marked with identity information, and uses an initial voiceprint recognition model to label the identity information corresponding to the plurality of second voiceprint data. The target model can be trained based on these voiceprint data with identity labels to obtain a voiceprint recognition model with a relatively high accuracy rate after training.
[0131] Please refer to Figure 8 , Figure 8The figure shows a labeling device for voiceprint data based on a voiceprint model provided by an embodiment of the present application. The labeling device for voiceprint data includes: a data acquisition module, a model training module, a voiceprint recognition module, a voiceprint comparison module, and a data labeling module. Among them, the data acquisition module is used to acquire a plurality of voiceprint data, and the plurality of voiceprint data includes a plurality of first voiceprint data already labeled with identity information, and a plurality of second voiceprint data not labeled with identity information; the model training module is used to train an initial model based on the first voiceprint data to obtain an initial voiceprint recognition model; the voiceprint recognition module is used to perform voiceprint recognition on the plurality of second voiceprint data based on the initial voiceprint recognition model to obtain the voiceprint features corresponding to each second voiceprint data; the voiceprint comparison module is used to obtain multiple groups of voiceprint data existing in the plurality of second voiceprint data based on the voiceprint features corresponding to each second voiceprint data. Each second voiceprint data in each group of voiceprint data belongs to the same user, and each group of voiceprint data includes at least two second voiceprint data; the data labeling module is used to label the identity information corresponding to the second voiceprint data in each group of voiceprint data as the identity information of the same user.
[0132] As a possible implementation manner, the voiceprint recognition module may include a similarity acquisition unit and a grouping unit. Among them, the similarity acquisition unit is used to obtain the similarity between every two second voiceprint data in the plurality of second voiceprint data based on the voiceprint features corresponding to each second voiceprint data; the grouping unit is used to obtain the second voiceprint data belonging to the same user from the plurality of second voiceprint data based on the similarity between every two second voiceprint data in the plurality of second voiceprint data, and use the second voiceprint data belonging to the same user as a group of voiceprint data to obtain multiple groups of voiceprint data.
[0133] As a possible implementation manner, the grouping unit may further be used to: determine whether the similarity between every two second voiceprint data is greater than a preset threshold; if the similarity between any two target voiceprint data is greater than the preset threshold, determine the two target voiceprint data as the second voiceprint data belonging to the same user, where the target voiceprint data is any one of the plurality of second voiceprint data.
[0134] As a possible implementation manner, the grouping unit may further be used to: randomly extract a preset number of groups of voiceprint data from the multiple groups of voiceprint data; obtain the test result of the user's test on the preset number of groups of voiceprint data, and the test result is used to characterize whether the second voiceprint data in each group of voiceprint data belongs to the same user; if there is any group of voiceprint data whose verification result indicates that the second voiceprint data it includes does not belong to the same user, increase the preset threshold; update the multiple groups of voiceprint data and other voiceprint data based on the changed preset threshold.
[0135] As a possible implementation, the data annotation module may include a target acquisition unit, a to-be-determined data unit, a determined data unit, a data update unit, and a repeated loop unit. Among them, the target acquisition unit is used to acquire the target similarity in the similarity, and the absolute value of the difference between the target similarity and the preset threshold is less than the target value; the to-be-determined data unit is used to acquire the second voiceprint data corresponding to the target similarity and the second voiceprint data having the same identity information as the second voiceprint data corresponding to the target similarity as the to-be-determined voiceprint data, and remove the identity information corresponding to the to-be-determined voiceprint data; the determined data unit is used to acquire the voiceprint data other than the to-be-determined voiceprint data among the multiple voiceprint data as the determined voiceprint data; the data update unit is used to perform transfer training on the initial voiceprint recognition model based on the determined voiceprint data and its labeled identity information to obtain a new voiceprint recognition model, update the initial voiceprint recognition model to the new voiceprint recognition model, and update the multiple second voiceprint data to the to-be-determined voiceprint data; the repeated loop unit is used to repeat the step of performing voiceprint recognition on the multiple second voiceprint data based on the initial voiceprint recognition model to obtain the voiceprint features corresponding to each second voiceprint data until the identity information corresponding to each group of voiceprint data in the multiple groups of voiceprint data is labeled with different identity information, and the identity information corresponding to the to-be-determined voiceprint data is different from the identity information corresponding to the determined voiceprint data.
[0136] As a possible implementation, the voiceprint comparison module may also be used to cluster the multiple second voiceprint data based on the voiceprint features corresponding to each second voiceprint data to obtain multiple categories; and use the second voiceprint data in each category as a group of voiceprint data to obtain multiple groups of voiceprint data.
[0137] As a possible implementation, the data annotation module may also include a model training unit for training the target model according to the voiceprint data marked with identity information to obtain a trained voiceprint recognition model.
[0138] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0139] In several embodiments provided in the present application, the coupling between modules may be electrical, mechanical, or other forms of coupling.
[0140] In addition, in each embodiment of the present application, the various functional modules may be integrated in one processing module, or each module may exist physically alone, or two or more modules may be integrated in one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0141] In summary, in the method for annotating voiceprint data based on a voiceprint model provided in this application, an initial voiceprint recognition model is obtained by training an initial model with multiple first voiceprint data already annotated with identity information, and the voiceprint recognition is performed on multiple second voiceprint data not annotated with identity information through this model to obtain corresponding voiceprint features, and all second voiceprint data are grouped according to the voiceprint features to obtain multiple groups of voiceprint data and other voiceprint data, and then the multiple groups of voiceprint data and other voiceprint data are annotated with different identity information, which not only improves the efficiency of data annotation; moreover, using the manually annotated voiceprint data to train the initial voiceprint recognition model to annotate the unannotated voiceprint data can also improve the accuracy of data annotation.
[0142] Please refer to Figure 9 , which shows a structural block diagram of a computer device 200 provided in an embodiment of this application. The computer device 200 in this application may include one or more of the following components: a processor 210, a memory 220, and one or more application programs, where one or more application programs may be stored in the memory 220 and configured to be executed by one or more processors 210, and one or more programs are configured to execute the method described in the foregoing method embodiments.
[0143] The processor 210 may include one or more processing cores. The processor 210 connects various parts within the entire computer device using various interfaces and lines, and executes various functions of the computer device and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 220, and by calling data stored in the memory 220. Optionally, the processor 210 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 210 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the display content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 210 and may be implemented separately through a communication chip.
[0144] The memory 220 may include a Random Access Memory (RAM), or may also include a Read-Only Memory. The memory 220 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 220 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the computer device (such as a phone book, audio and video data, chat record data, etc.).
[0145] Please refer to Figure 10 , which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable medium 800, and the program code can be called by a processor to execute the methods described in the above method embodiments.
[0146] The computer-readable storage medium 800 can be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium 800 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 800 has a storage space for the program code 810 for executing any method step in the above methods. These program codes can be read out from or written into one or more computer program products. The program code 810 can be compressed in an appropriate form, for example.
[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A method for annotating voiceprint data based on a voiceprint model, characterized in that, The method includes: Obtaining a plurality of voiceprint data, where the plurality of voiceprint data includes a plurality of first voiceprint data with identity information labeled thereon, and a plurality of second voiceprint data without identity information labeled thereon; Training an initial model based on the first voiceprint data to obtain an initial voiceprint recognition model; Performing voiceprint recognition on the plurality of second voiceprint data based on the initial voiceprint recognition model to obtain voiceprint features corresponding to each second voiceprint data; Based on the voiceprint features corresponding to each second voiceprint data, obtaining multiple groups of voiceprint data existing in the plurality of second voiceprint data, and second voiceprint data other than the multiple groups of voiceprint data in the plurality of second voiceprint data as other voiceprint data, where each second voiceprint data in each group of the multiple groups of voiceprint data belongs to the same user, and each group of voiceprint data includes at least two second voiceprint data; Labeling the identity information corresponding to each group of voiceprint data in the multiple groups of voiceprint data as different identity information, and labeling each voiceprint data in the other voiceprint data as different identity information, where the identity information corresponding to the second voiceprint data in the same group of the multiple groups of voiceprint data is the same, and the identity information corresponding to each group of voiceprint data is different from the identity information corresponding to each voiceprint data in the other voiceprint data.
2. The method according to claim 1, characterized in that, The obtaining multiple groups of voiceprint data existing in the plurality of second voiceprint data based on the voiceprint features corresponding to each second voiceprint data includes: Obtaining the similarity between every two second voiceprint data in the plurality of second voiceprint data based on the voiceprint features corresponding to each second voiceprint data; Based on the similarity between every two second voiceprint data in the plurality of second voiceprint data, obtaining second voiceprint data belonging to the same user from the plurality of second voiceprint data, and taking the second voiceprint data belonging to the same user as a group of voiceprint data to obtain multiple groups of voiceprint data.
3. The method according to claim 2, wherein The obtaining second voiceprint data belonging to the same user from the plurality of second voiceprint data based on the similarity between every two second voiceprint data in the plurality of second voiceprint data includes: Determining whether the similarity between every two second voiceprint data is greater than a preset threshold; If the similarity between any two target voiceprint data is greater than the preset threshold, then determining the two target voiceprint data as second voiceprint data belonging to the same user, where the target voiceprint data is any voiceprint data in the plurality of second voiceprint data.
4. The method according to claim 3, wherein After obtaining second voiceprint data belonging to the same user from the plurality of second voiceprint data based on the similarity between every two second voiceprint data in the plurality of second voiceprint data, and taking the second voiceprint data belonging to the same user as a group of voiceprint data to obtain multiple groups of voiceprint data, the method further includes: Randomly extracting a preset number of groups of voiceprint data from the multiple groups of voiceprint data; Obtaining a test result of the user's test on the preset number of groups of voiceprint data, where the test result is used to characterize whether the second voiceprint data in each group of voiceprint data belongs to the same user and whether the second voiceprint data in different groups of voiceprint data does not belong to the same user; Adjusting the preset threshold according to the test result; Update the multiple sets of voiceprint data and the other voiceprint data based on the adjusted preset threshold.
5. The method according to claim 4, characterized in that, Adjusting the preset threshold according to the inspection result includes: If the verification result of any set of voiceprint data indicates that the included second voiceprint data does not belong to the same user, increase the preset threshold.
6. The method according to claim 4, characterized in that, Adjusting the preset threshold according to the inspection result includes: If the verification results of different sets of voiceprint data indicate that any two second voiceprint data belonging to different groups belong to the same user, decrease the preset threshold.
7. The method according to claim 4, characterized in that Randomly extracting a preset number of sets of voiceprint data from the multiple sets of voiceprint data includes: Obtain the number of second voiceprint data in each set of voiceprint data; Adjust the weight of each set of voiceprint data according to the number of the second voiceprint data, and the weight is positively correlated with the number of the second voiceprint data; Randomly extract a preset number of sets of voiceprint data from the multiple sets of voiceprint data based on the weight of each set of voiceprint data.
8. The method according to claim 3, wherein After labeling the identity information corresponding to each set of voiceprint data in the multiple sets of voiceprint data as different identity information, and labeling each voiceprint data in the other voiceprint data as different identity information, the method further includes: Obtain the target similarity in the similarities, and the absolute value of the difference between the target similarity and the preset threshold is less than the target value; Obtain the second voiceprint data corresponding to the target similarity and the second voiceprint data having the same identity information as the second voiceprint data corresponding to the target similarity as the voiceprint data to be determined, and remove the identity information corresponding to the voiceprint data to be determined; Obtain the voiceprint data other than the voiceprint data to be determined among the multiple voiceprint data as the determined voiceprint data; Based on the determined voiceprint data and its labeled identity information, perform transfer training on the initial voiceprint recognition model to obtain a new voiceprint recognition model, update the initial voiceprint recognition model to the new voiceprint recognition model, and update the multiple second voiceprint data to the voiceprint data to be determined; Repeat the step of performing voiceprint recognition on the multiple second voiceprint data based on the initial voiceprint recognition model to obtain the voiceprint features corresponding to each second voiceprint data until the step of labeling the identity information corresponding to each set of voiceprint data in the multiple sets of voiceprint data as different identity information, and labeling each voiceprint data in the other voiceprint data as different identity information, wherein the identity information corresponding to the voiceprint data to be determined is different from the identity information corresponding to the determined voiceprint data.
9. The method according to claim 1, wherein Obtaining multiple sets of voiceprint data existing in the multiple second voiceprint data based on the voiceprint features corresponding to each second voiceprint data includes: Cluster the multiple second voiceprint data based on the voiceprint features corresponding to each second voiceprint data to obtain multiple categories; Take the second voiceprint data in each category as a set of voiceprint data to obtain multiple sets of voiceprint data.
10. The method according to any one of claims 1-9, characterized in that, After labeling the identity information corresponding to the second voiceprint data in each set of voiceprint data as the identity information of the same user, the method further includes: Train a target model based on voiceprint data marked with identity information to obtain a trained voiceprint recognition model.
Citation Information
Patent Citations
Voiceprint model training method, voiceprint recognition method and device
CN106057206A
Training method and device of voiceprint recognition model, electronic equipment and storage medium
CN109801636A