Voiceprint determination threshold value determination method, device and equipment
By calculating the inter-class variance and precision-recall balance score in voiceprint recognition technology to determine the voiceprint judgment threshold, the problem of inappropriate voiceprint recognition threshold selection in existing technologies is solved, enabling fast and accurate verification of user identity and improving user experience and security.
Patent Information
- Application Number
- CN202411919302.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-12-24
AI Technical Summary
Existing voiceprint recognition technologies struggle to select appropriate voiceprint judgment thresholds in dual-terminal and multi-number personal platforms, leading to poor user experience or security risks. Common methods such as the empirical method, the equal error rate method, and the highest accuracy method have limitations and cannot meet the accuracy requirements.
By acquiring users' registered voiceprints and training datasets, a comparison score set is calculated. The target threshold is determined using inter-class variance and precision-recall balance scores. The voiceprint judgment threshold is adaptively set, taking into account both accuracy and security.
It enables rapid and accurate verification of user identity on dual-device and multi-account platforms, improving the accuracy and security of voiceprint recognition.
Smart Images

Figure CN119811425B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of voiceprint recognition technology, and in particular relates to a method, apparatus and equipment for determining the voiceprint judgment threshold. Background Technology
[0002] With the continuous advancement of communication technology, new communication platforms such as dual-device and multiple personal numbers are becoming increasingly popular, providing users with more flexible and convenient communication services. These platforms achieve seamless information connection and sharing by integrating multiple terminal devices or phone numbers. However, this also brings challenges to identity verification and security management. To address these challenges, voiceprint recognition technology is widely used in these platforms to achieve rapid and accurate verification of user identity.
[0003] In the application of voiceprint recognition technology, selecting the optimal voiceprint judgment threshold is a crucial issue. The choice of threshold directly affects the accuracy and security of voiceprint recognition. If the threshold is set too high, a large number of legitimate users may be mistakenly identified as illegitimate users, affecting the user experience; if the threshold is set too low, illegitimate users may easily pass verification, creating security risks.
[0004] Currently, common methods for selecting voiceprint recognition thresholds include empirical methods, the highest accuracy method, and the equal error rate method. However, these methods all have their limitations and are difficult to meet the precise requirements of dual-device and multi-number personal accounts for voiceprint recognition threshold selection. Summary of the Invention
[0005] This application provides a method, apparatus, device, computer storage medium, and computer program product for determining a voiceprint determination threshold, which can adaptively determine a suitable voiceprint determination threshold for each user, thereby facilitating rapid and accurate verification of user identity.
[0006] In a first aspect, embodiments of this application provide a method for determining a voiceprint detection threshold, including:
[0007] In response to the first user's voiceprint registration request, obtain the first user's first registered voiceprint;
[0008] Obtain a training dataset, which includes at least one positive sample and at least one negative sample. The positive sample includes the first voiceprint data of the first user, and the negative sample includes the second voiceprint data of the second user.
[0009] The comparison scores of each positive sample and each negative sample with the first registered voiceprint are calculated to obtain the comparison score set;
[0010] Based on the comparison score set, multiple undetermined thresholds are determined;
[0011] For each undetermined threshold, the comparison scores in the comparison score set are classified using the undetermined threshold as the segmentation value, and the inter-class variance and precision-recall balance score corresponding to the classification result are calculated.
[0012] Based on the inter-class variance and precision-recall balance score corresponding to each undetermined threshold, the target threshold is determined from multiple undetermined thresholds;
[0013] The target threshold is used as the voiceprint determination threshold for the first user.
[0014] In one optional implementation, the first voiceprint data is the recording data of the first user; the second user includes a registered user, and the second voiceprint data is the second registered voiceprint of the second user;
[0015] Obtain the training dataset, including:
[0016] Obtain at least one recording from the actual production recording database;
[0017] Retrieve at least one second registered voiceprint corresponding to the second user from the registered voiceprint database;
[0018] A training dataset is created based on at least one recording of a first user and at least one second registered voiceprint of a second user.
[0019] In one optional implementation, the comparison scores of each positive sample and each negative sample with the first registered voiceprint are calculated to obtain a comparison score set, including:
[0020] Calculate the similarity between each positive sample and each negative sample and the first registered voiceprint;
[0021] Based on the similarity between each positive sample and each negative sample and the first registered voiceprint, the comparison score of each positive sample and each negative sample and the first registered voiceprint is determined, and the comparison score set is obtained.
[0022] In one alternative implementation, multiple undetermined thresholds are determined based on a set of comparison scores, including:
[0023] Based on the maximum and minimum values in the comparison score set, determine the range of values for the undetermined threshold.
[0024] Multiple undetermined thresholds are determined from the value range.
[0025] In one alternative implementation, determining multiple undetermined thresholds from a value range includes:
[0026] According to the preset step value, multiple undetermined thresholds are determined within the value range.
[0027] In one alternative implementation, the classification results include a first set of scores less than or equal to a pending threshold and a second set of scores greater than the pending threshold.
[0028] Calculate the inter-class variance corresponding to the classification results, including:
[0029] Calculate the first average value of each comparative score in the first rating set, and the first proportion of the first rating set, where the first proportion is the ratio of the number of comparative scores in the first rating set to the number of comparative scores in the comparative rating set.
[0030] Calculate the second average value of each comparative score within the second rating set, and the second proportion of the second rating set, where the second proportion is the ratio of the number of comparative scores in the second rating set to the number of comparative scores in the comparative rating set.
[0031] Calculate the first difference of squares and the second difference of squares respectively, where the first difference of squares is the square of the difference between the first mean and the third mean, the second difference of squares is the square of the difference between the second mean and the third mean, and the third mean is the average of the comparison scores in the comparison score set;
[0032] Calculate the product of the first squared difference and the first proportion to obtain the first product;
[0033] Calculate the product of the second squared difference and the second proportion to obtain the second product;
[0034] Add the first product to the second product to obtain the inter-class variance corresponding to the classification result.
[0035] In one optional implementation, calculating the precision-recall balance score corresponding to the classification result includes:
[0036] Determine the first number of comparison scores corresponding to positive examples in the second set of scores;
[0037] Calculate the ratio of the number of comparison scores in the first set to the number of comparison scores in the second set to obtain the accuracy of the classification result;
[0038] Calculate the ratio of the first quantity to the second quantity to obtain the recall rate corresponding to the classification result, where the second quantity is the number of positive samples in the training dataset;
[0039] Based on precision and recall, and according to preset precision and recall weights, the precision-recall balance score corresponding to the classification result is calculated.
[0040] In one optional implementation, a target threshold is determined from multiple undetermined thresholds based on the inter-class variance and precision-recall balance score corresponding to each undetermined threshold, including:
[0041] Normalize the inter-class variance and precision-recall balance score corresponding to each undetermined threshold to obtain the normalized inter-class variance and normalized precision-recall balance score corresponding to each undetermined threshold.
[0042] The target threshold is determined from multiple undetermined thresholds based on the normalized inter-class variance and the normalized precision-recall balance score.
[0043] In one alternative implementation, a target threshold is determined from a plurality of undetermined thresholds based on the normalized inter-class variance and the normalized precision-recall equilibrium score, including:
[0044] Calculate the target sum for each undetermined threshold. The target sum is the sum of the normalized inter-class variance and the normalized precision-recall equilibrium score of the undetermined threshold.
[0045] The threshold with the largest target value is determined as the first candidate threshold.
[0046] The target threshold is determined from at least one first candidate threshold.
[0047] In one alternative implementation, determining the target threshold from at least one first candidate threshold includes:
[0048] When the number of first candidate thresholds is greater than one, the target difference for each first candidate threshold is calculated. The target difference is the absolute value of the difference between the normalized precision-recall equilibrium score and the normalized inter-class variance of the first candidate threshold.
[0049] The first candidate threshold with the smallest target difference is determined as the second candidate threshold;
[0050] The target threshold is determined from at least one second candidate threshold.
[0051] In one alternative implementation, determining the target threshold from at least one second candidate threshold includes:
[0052] If the number of second candidate thresholds is greater than one, the second candidate threshold with the largest value among at least one second candidate threshold shall be determined as the target threshold.
[0053] In one optional implementation, after obtaining the first registered voiceprint of the first user in response to the first user's voiceprint registration request, the method further includes performing the following steps for each second user:
[0054] Obtain an updated dataset, which includes at least one updated positive sample and at least two updated negative samples. The updated positive sample includes the recording data of the second user, and the at least two updated negative samples include the first registered voiceprint and the second registered voiceprints of the remaining second users.
[0055] The comparison scores between each updated positive sample and each updated negative sample and the first registered voiceprint are calculated to obtain the updated score set;
[0056] Based on the updated score set, multiple pending update thresholds are determined;
[0057] For each pending update threshold, the comparison scores in the update score set are classified using the pending update threshold as the segmentation value, and the inter-class variance and precision-recall balance score corresponding to the classification results are calculated.
[0058] Based on the inter-class variance and precision-recall balance score corresponding to each pending update threshold, the target update threshold is determined from multiple pending update thresholds.
[0059] Update the voiceprint determination threshold in the second user's voiceprint registration information to the target update threshold.
[0060] In an optional implementation, after using the target threshold as the voiceprint determination threshold for the first user, the method further includes:
[0061] In response to a user identification request, obtain the third voiceprint data of the user to be identified;
[0062] Calculate the comparison score between the third voiceprint data and the first registered voiceprint;
[0063] If the score of the comparison between the third voiceprint data and the first registered voiceprint is greater than the voiceprint determination threshold of the first user, the user to be identified is determined to be the first user.
[0064] Secondly, embodiments of this application provide a voiceprint determination threshold determination device, comprising:
[0065] The acquisition module is used to acquire the first registered voiceprint of the first user in response to the first user's voiceprint registration request;
[0066] The acquisition module is also used to acquire a training dataset, which includes at least one positive sample and at least one negative sample. The positive sample includes the first voiceprint data of the first user, and the negative sample includes the second voiceprint data of the second user.
[0067] The calculation module is used to calculate the comparison score between each positive sample and each negative sample and the first registered voiceprint, and obtain the comparison score set.
[0068] The determination module is used to determine multiple pending thresholds based on the comparison score set;
[0069] The processing module is used to classify the comparison scores in the comparison score set by comparing each undetermined threshold with the undetermined threshold as the segment value, and calculate the inter-class variance and precision-recall balance score corresponding to the classification result.
[0070] The determination module is also used to determine the target threshold from multiple undetermined thresholds based on the inter-class variance and precision-recall balance score corresponding to each undetermined threshold.
[0071] The processing module is also used to use the target threshold as the voiceprint determination threshold for the first user.
[0072] Thirdly, embodiments of this application provide an electronic device, the device including: a processor and a memory storing computer program instructions;
[0073] When the processor executes computer program instructions, it implements the voiceprint determination threshold determination method as described in any optional embodiment of the first aspect of this application.
[0074] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the voiceprint determination threshold determination method as described in any optional embodiment of the first aspect of this application.
[0075] Fifthly, embodiments of this application provide a computer program product in which instructions are executed by the processor of an electronic device, causing the electronic device to perform a voiceprint determination threshold determination method as described in any optional embodiment of the first aspect of this application.
[0076] The voiceprint determination threshold determination method, apparatus, device, computer storage medium, and computer program product of this application embodiment obtain the first registered voiceprint of the first user in response to the voiceprint registration request of the first user. Next, a training dataset is obtained, including at least one positive sample and at least one negative sample. The positive sample includes the first voiceprint data of the first user, and the negative sample includes the second voiceprint data corresponding to the second user. This improves the reliability of the training dataset by using the voiceprint data of the first user as the positive sample and the voiceprint data of other users as the negative sample. Subsequently, comparison scores between each positive sample and each negative sample and the first registered voiceprint are calculated to obtain a comparison score set. Based on the comparison score set, multiple undetermined thresholds are determined. For each undetermined threshold, the comparison scores in the comparison score set are classified using the undetermined threshold as a segmentation value, and the inter-class variance and precision-recall balance score corresponding to the classification result are calculated. Based on the inter-class variance and precision-recall balance score corresponding to each undetermined threshold, a target threshold is determined from the multiple undetermined thresholds. The target threshold is used as the voiceprint determination threshold for the first user. In this way, the inter-class variance algorithm and precision-recall balance score can be combined to balance multiple selection indicators and adaptively determine the unique voiceprint judgment threshold for each user based on different registered voiceprints. Thus, the determined voiceprint judgment threshold can adapt to actual production and business environments. When applied to scenarios such as dual-device use or multiple personal phone numbers, it can achieve fast and accurate user identity verification based on the actual input voiceprint data and the determined voiceprint judgment threshold. Attached Figure Description
[0077] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0078] Figure 1 This is a flowchart illustrating a method for determining a voiceprint threshold according to an embodiment of this application;
[0079] Figure 2 This is a schematic diagram of the structure of a voiceprint determination threshold determination device provided in another embodiment of this application;
[0080] Figure 3 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0081] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0082] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0083] Voiceprint recognition technology is widely used in new communication platforms such as dual-number and multiple personal numbers to achieve fast and accurate user identity verification. Taking dual-number as an example, in this scenario, two devices share a single phone number. This service is implemented by binding the user's mobile phone number to a wearable device. After binding, the user's mobile phone becomes the "primary device," and the wearable device becomes the "secondary device." Whether the user is making or receiving calls, both the primary and secondary devices present the same number. When a user applies for dual-number, they are required to record their voiceprint and register it in the database. Subsequent identity verification can then be performed using voiceprint recognition technology.
[0084] In the application of voiceprint recognition technology, selecting the optimal voiceprint judgment threshold is a crucial issue. The choice of threshold directly affects the accuracy and security of voiceprint recognition. If the threshold is set too high, a large number of legitimate users may be mistakenly identified as illegitimate users, affecting the user experience; if the threshold is set too low, illegitimate users may easily pass verification, creating security risks.
[0085] The relevant technology involves selecting a voiceprint recognition threshold using an empirical method. Specifically, this empirical method can include statistically analyzing data from numerous past real-world cases of the voiceprint recognition system to understand the system's recognition accuracy under different conditions. Subsequently, based on factors such as voiceprint sample quality (signal-to-noise ratio, etc.), the recognition algorithm, and database size, an appropriate recognition accuracy requirement is determined, such as 90%. Finally, based on this accuracy requirement and combined with historical system data, a threshold that achieves this accuracy is selected. This threshold is the voiceprint recognition threshold selected using the empirical method.
[0086] However, the empirical method lacks theoretical basis, is difficult to optimize quantitatively, and relies heavily on the accumulation of historical experience of the system. For new systems without sufficient data, it can serve as a preliminary reference selection method. With the development of new communication platforms such as dual-terminal and multiple personal numbers, the amount of voiceprint data in the system has increased significantly, and the drawbacks of the empirical method have become increasingly apparent.
[0087] Related technologies also involve selecting a voiceprint recognition threshold using the equal error rate method. The main goal of the equal error rate method is to find a threshold at which the model's false positive rate and false negative rate are equal. The false positive rate refers to the proportion of samples that the model predicts as positive but are actually negative out of the total samples, while the false negative rate refers to the proportion of samples that the model predicts as negative but are actually positive out of the total samples.
[0088] However, the equal error rate method only considers the balance between the false positive rate and the false negative rate, without taking into account the balance between the true positive rate and the false negative rate. In some cases, this may lead to an overemphasis on controlling the false positive rate, resulting in a lower true positive rate.
[0089] Related technologies also involve selecting the voiceprint determination threshold using the highest accuracy method. The highest accuracy method calculates the recognition accuracy at different thresholds and selects the threshold with the highest accuracy as the voiceprint determination threshold. Specifically, first, a voiceprint database and test voiceprint data are collected and labeled. Then, a threshold range is set, such as 0.3-0.9, and multiple threshold points are evenly selected within this range. For each threshold point, the test voiceprint is matched against the database, and the number of matching samples is counted. The recognition accuracy at this threshold is calculated as: Recognition Accuracy = Number of Matching Samples / Total Number of Test Samples. The recognition accuracy calculated at different threshold points is compared, and the threshold point with the highest accuracy is selected as the voiceprint determination threshold.
[0090] However, the highest precision method only considers precision and does not balance other metrics such as recall, which may overlook some important information. Furthermore, the highest precision method is heavily reliant on test data; if the distribution of test data differs significantly from the actual application scenario, the chosen threshold may be unsuitable.
[0091] In view of this, after in-depth consideration, the inventors have ingeniously proposed a method for determining the voiceprint judgment threshold, an information receiving method, a device, an equipment, a computer storage medium, and a computer program product.
[0092] The voiceprint determination threshold method provided in this application will be described below with reference to the accompanying drawings, through specific embodiments and application scenarios. The device executing the transmission in the voiceprint determination threshold method provided in this application can be a voiceprint determination threshold determination device, or a portion of the voiceprint determination threshold determination device used to execute the voiceprint determination threshold method. This application uses the execution of the voiceprint determination threshold method by a voiceprint determination threshold determination device as an example to describe in detail the voiceprint determination threshold method provided in this application.
[0093] The following is in conjunction with the appendix Figure 1 The method for determining the voiceprint determination threshold provided in the embodiments of this application will be described in detail.
[0094] Figure 1 A flowchart illustrating a method for determining a voiceprint detection threshold according to an embodiment of this application is shown. Figure 1 As shown, the method for determining the voiceprint determination threshold may specifically include the following steps S110 to 170.
[0095] S110, in response to the first user's voiceprint registration request, obtain the first user's first registered voiceprint.
[0096] In step S110, the voiceprint registration request can be used to request the first user's voiceprint to be entered into the registered voiceprint database, so as to perform voiceprint recognition on the first user based on the first registered voiceprint entered by the first user. This voiceprint registration request may include registration requests generated when a user applies for new communication platform services such as dual-terminal or multiple personal numbers.
[0097] S120, Obtain the training dataset, which includes at least one positive sample and at least one negative sample. The positive sample includes the first voiceprint data of the first user, and the negative sample includes the second voiceprint data of the second user.
[0098] In step S120, the number of positive and negative samples is not specifically limited and can be selected according to actual needs. In one example, the number of positive samples can be equal to the number of negative samples. When the ratio of positive to negative samples is 1:1, it is beneficial to optimize the decision boundary and improve the efficiency of determining the voiceprint judgment threshold.
[0099] S130, calculate the comparison score between each positive sample and each negative sample and the first registered voiceprint, and obtain the comparison score set.
[0100] In step S130, the comparison score can be calculated in various ways, and is not limited here. For example, the similarity between the sample and the first registered voiceprint can be calculated separately to obtain the comparison score. For example, the similarity between the sample and the first registered voiceprint can also be calculated based on a preset voiceprint comparison model.
[0101] S140, Based on the comparison score set, determine multiple pending thresholds.
[0102] Step S140 described above can be implemented in various ways. For example, multiple comparison scores can be selected from the comparison score set as pending thresholds. Alternatively, multiple pending thresholds can be selected based on the range of comparison scores, and so on. This application does not limit the implementation in this regard.
[0103] S150, for each undetermined threshold, classify the comparison scores in the comparison score set by using the undetermined threshold as the segmentation value, and calculate the inter-class variance and precision-recall balance score corresponding to the classification result.
[0104] In step S150, the inter-class variance can be a statistical indicator used to measure the degree of difference between different categories or groups. For example, the inter-class variance can be calculated using the inter-class variance formula in Otsu's algorithm. The precision-recall balance score can represent the combined score of precision and recall. For example, the precision-recall balance score can be calculated using the F-score method, and can be characterized by the F-score. The F-score can be a statistical indicator used to measure the precision of a binary classification model; it is the harmonic mean of precision and recall. The F-score can be seen as a trade-off between precision and recall in the classification results, used to comprehensively evaluate the classification results.
[0105] S160, determine the target threshold from multiple undetermined thresholds based on the inter-class variance and precision-recall balance score corresponding to each undetermined threshold.
[0106] S170, the target threshold is used as the voiceprint determination threshold for the first user.
[0107] In step S170, after recording the voiceprint determination threshold in the voiceprint registration information of the first user, the first user can be identified based on this voiceprint determination threshold during subsequent voiceprint recognition processes. Specifically, in one example, after using the target threshold as the voiceprint determination threshold for the first user, this voiceprint determination threshold can be recorded in the first user's voiceprint registration information. When user identification is required, the voiceprint data of the user to be identified can be obtained, and a comparison score between the voiceprint data and the first registered voiceprint can be calculated. If the comparison score is higher than the voiceprint determination threshold of the first user, then the user to be identified can be determined as the first user.
[0108] The voiceprint determination threshold method of this application embodiment obtains the first registered voiceprint of the first user in response to the voiceprint registration request of the first user. Next, a training dataset is obtained, including at least one positive sample and at least one negative sample. The positive sample includes the first voiceprint data of the first user, and the negative sample includes the second voiceprint data corresponding to the second user. This improves the reliability of the training dataset by using the voiceprint data of the first user as the positive sample and the voiceprint data of other users as the negative sample. Subsequently, the comparison scores between each positive sample and each negative sample and the first registered voiceprint are calculated to obtain a comparison score set. Based on the comparison score set, multiple undetermined thresholds are determined. For each undetermined threshold, the comparison scores in the comparison score set are classified using the undetermined threshold as a segmentation value, and the inter-class variance and precision-recall balance score corresponding to the classification result are calculated. Based on the inter-class variance and precision-recall balance score corresponding to each undetermined threshold, a target threshold is determined from the multiple undetermined thresholds. The target threshold is used as the voiceprint determination threshold for the first user. In this way, the inter-class variance algorithm and precision-recall balance score can be combined to balance multiple selection indicators and adaptively determine the unique voiceprint judgment threshold for each user based on different registered voiceprints. Thus, the determined voiceprint judgment threshold can adapt to actual production and business environments. When applied to scenarios such as dual-device use or multiple personal phone numbers, it can achieve fast and accurate user identity verification based on the actual input voiceprint data and the determined voiceprint judgment threshold.
[0109] In one embodiment, the first voiceprint data is the recording data of a first user. The second user includes a registered user, and the second voiceprint data is the second registered voiceprint of the second user.
[0110] Obtaining the training dataset can specifically include:
[0111] Obtain at least one recording from the actual production recording database.
[0112] Retrieve at least one second registered voiceprint corresponding to the second user from the registered voiceprint database.
[0113] A training dataset is created based on at least one recording of a first user and at least one second registered voiceprint of a second user.
[0114] In the above embodiments, the actual production recording database can be a database used to record publicly available recording data of users. The recording data can include recording data of users in actual production and business environments, such as publicly available call recording data generated when users talk to third-party platforms.
[0115] According to the above implementation method, the recording data of the first user in the actual generation and business environment is used as positive example samples, and the second registered voiceprint of the registered second user is used as negative example samples. In this way, the sources of both positive and negative example samples are authentic and reliable. Based on the training dataset created from the above positive and negative example samples, the voiceprint judgment threshold of the first user is determined using this training dataset. This can effectively distinguish the voiceprint of the first user from that of other users in the registered voiceprint database, thus facilitating accurate recognition of user voiceprints.
[0116] In one embodiment, the comparison scores of each positive sample and each negative sample with the first registered voiceprint are calculated to obtain a comparison score set, which may specifically include:
[0117] Calculate the similarity between each positive sample and each negative sample and the first registered voiceprint.
[0118] Based on the similarity between each positive sample and each negative sample and the first registered voiceprint, the comparison score of each positive sample and each negative sample and the first registered voiceprint is determined, and the comparison score set is obtained.
[0119] In the above embodiments, similarity can be measured by various parameters, and this application does not limit this. As an example, similarity may include cosine similarity, Euclidean distance, etc.
[0120] According to the above implementation method, the similarity between each positive sample and each negative sample and the first registered voiceprint is calculated, and the comparison score between each positive sample and each negative sample and the first registered voiceprint is determined by the similarity. This facilitates an accurate measurement of the matching degree between each sample and the first registered voiceprint, thereby providing a basis for determining an appropriate voiceprint determination threshold.
[0121] In one embodiment, multiple undetermined thresholds are determined based on a set of comparison scores, which may specifically include:
[0122] The range of values for the undetermined threshold is determined by comparing the maximum and minimum values in the scoring set.
[0123] Multiple undetermined thresholds are determined from the value range.
[0124] According to the above implementation method, the range of values for the undetermined threshold is determined based on the maximum and minimum values in the comparison score set. Then, multiple undetermined thresholds are determined from this range. In this way, a suitable undetermined threshold can be selected based on the distribution range of the comparison scores. This facilitates the determination of a suitable voiceprint determination threshold, thereby improving the accuracy of voiceprint recognition.
[0125] In one embodiment, determining multiple undetermined thresholds from a value range may specifically include:
[0126] According to the preset step value, multiple undetermined thresholds are determined within the value range.
[0127] In the above implementation, the step value can be set according to the actual required voiceprint recognition accuracy. Multiple undetermined thresholds are determined within a range according to a preset step value. This can include recursively determining multiple undetermined thresholds starting from the left endpoint of the range and following the preset step value until the right endpoint of the range is reached or exceeded. For example, if the range of undetermined thresholds is assumed to be [0.30, 0.85], and the step value is set to 0.01, then 55 undetermined thresholds can be determined: 0.30, 0.31, 0.32, 0.33…0.84, 0.85.
[0128] According to the above implementation method, multiple undetermined thresholds can be determined within a value range according to preset step values. This allows for flexible adjustment of the number of undetermined thresholds based on the required voiceprint recognition accuracy, ensuring that the multiple undetermined thresholds are evenly distributed within the value range. This improves the accuracy and flexibility of voiceprint determination thresholds.
[0129] In one embodiment, the classification results include a first set of scores less than or equal to a pending threshold and a second set of scores greater than the pending threshold.
[0130] Calculating the inter-class variance corresponding to the classification results can specifically include:
[0131] Calculate the first average value of each comparative score in the first rating set, and the first proportion of the first rating set, where the first proportion is the ratio of the number of comparative scores in the first rating set to the number of comparative scores in the comparative rating set.
[0132] Calculate the second average value of each comparative score within the second rating set, and the second proportion of the second rating set, where the second proportion is the ratio of the number of comparative scores in the second rating set to the number of comparative scores in the comparative rating set.
[0133] Calculate the first square difference and the second square difference respectively, where the first square difference is the square of the difference between the first average and the third average, the second square difference is the square of the difference between the second average and the third average, and the third average is the average of the comparative scores in the comparative score set.
[0134] Calculate the product of the first squared difference and the first proportion to obtain the first product.
[0135] Calculate the product of the second squared difference and the second proportion to obtain the second product.
[0136] Add the first product to the second product to obtain the inter-class variance corresponding to the classification result.
[0137] To facilitate understanding, an example is given below to illustrate the above implementation method. For instance, for a training dataset with a total of N samples, a set of alignment scores containing N alignment scores can be obtained. Let L represent the maximum value of the alignment score in the set of alignment scores, and n... i p represents the number of comparison scores for value i. i p represents the probability of a comparison score of i. i It can be calculated using the following formula 1.
[0138]
[0139] Using an undetermined threshold T as the segmentation value, the comparison scores in the comparison score set are classified to obtain a first score set less than or equal to T and a second score set greater than T. The average value μ of the comparison score set is then calculated. T It can be calculated using the following formula 2.
[0140]
[0141] In Equation 2, T1 can represent the minimum value of the comparison score in the comparison score set.
[0142] Let ω0 be the proportion of the number of comparison scores in the first set to the total number of comparison scores in the comparison set, and let μ0 be the average value of the comparison scores in the first set. ω0 and μ0 can be calculated using Equations 3 and 4, respectively.
[0143]
[0144] Let ω1 be the proportion of the number of comparison scores in the second rating set to the total number of comparison scores in the comparison rating set, and let μ1 be the average value of the comparison scores in the second rating set. ω1 and μ1 can be calculated using Equations 5 and 6, respectively.
[0145] ω1 = 1 - ω0 (Equation 5)
[0146]
[0147] In Equation 6, T2 can represent the smallest alignment score greater than T in the alignment score set.
[0148] Using the undetermined threshold T as the segmentation value, the inter-class variance corresponding to the classification result is... It can be done
[0149] Calculate using Equation 7.
[0150]
[0151] According to the above implementation method, the inter-class variance can be calculated using the inter-class variance calculation method in Otsu's algorithm. The inter-class variance can effectively measure the segmentation effect of the undetermined threshold, thus providing a theoretical basis for selecting a suitable voiceprint determination threshold.
[0152] In one embodiment, calculating the precision-recall balance score corresponding to the classification result may specifically include:
[0153] Determine the first number of comparison scores corresponding to positive samples in the second set of scores.
[0154] Calculate the ratio of the number of comparison scores in the first set to the number of comparison scores in the second set to obtain the accuracy of the classification result.
[0155] Calculate the ratio of the first quantity to the second quantity to obtain the recall rate corresponding to the classification result, where the second quantity is the number of positive samples in the training dataset.
[0156] Based on precision and recall, and according to preset precision and recall weights, the precision-recall balance score corresponding to the classification result is calculated.
[0157] In the above implementation, precision can represent the proportion of samples predicted as positive that are actually positive. Recall can represent the proportion of all positive samples in the training dataset that are correctly identified as positive. The second score set is the set of alignment scores greater than a predetermined threshold; in other words, the second score set can represent the set of alignment scores for samples predicted as positive.
[0158] In the above implementation, the precision recall equalization score can be calculated using the following formula 8.
[0159]
[0160] In Equation 8, F can represent the precision-recall balance score. Precision can represent precision. Recall can represent recall. α can represent the weighting parameter. For example, α can be 0.5, indicating that precision and recall are given equal importance.
[0161] For example, suppose the training dataset contains 50 first voiceprint data points of a first user and 50 second voiceprint data points of a second user. For a threshold of 0.55, the corresponding classification results show that the second score set contains 60 comparison scores, of which 45 comparison scores correspond to the first voiceprint data.
[0162] At this point, the accuracy rate is calculated:
[0163]
[0164] Calculate recall rate:
[0165]
[0166] Substituting into Equation 8, with α set to 0.5, we get F = 1.6765.
[0167] According to the above implementation method, the segmentation effect of the undetermined threshold can be evaluated by combining precision and recall, thereby providing a theoretical basis for selecting a suitable voiceprint determination threshold.
[0168] In one embodiment, a target threshold is determined from multiple undetermined thresholds based on the inter-class variance and precision-recall balance score corresponding to each undetermined threshold. Specifically, this may include:
[0169] Normalize the inter-class variance and precision-recall balance score corresponding to each undetermined threshold to obtain the normalized inter-class variance and normalized precision-recall balance score corresponding to each undetermined threshold.
[0170] The target threshold is determined from multiple undetermined thresholds based on the normalized inter-class variance and the normalized precision-recall balance score.
[0171] In one example, for each undetermined threshold, the comparison scores in the score set are classified using that undetermined threshold as the segmentation value. After calculating the inter-class variance and precision-recall balance score corresponding to the classification results, the inter-class variance and precision-recall balance score corresponding to each undetermined threshold can be obtained. An inter-class variance curve can be plotted with the undetermined threshold as the x-axis and the inter-class variance as the y-axis; a precision-recall balance score curve can be plotted with the undetermined threshold as the x-axis and the precision-recall balance score as the y-axis. According to the above implementation method, the y-axis of the two curves can be normalized so that the y-axis value falls between 0 and 1.
[0172] According to the above implementation method, the inter-class variance and precision-recall balance scores corresponding to each undetermined threshold are normalized, which is conducive to more objective analysis of each index, thereby helping to identify a suitable voiceprint judgment threshold and improve the accuracy and precision of voiceprint recognition.
[0173] In one embodiment, a target threshold is determined from a plurality of undetermined thresholds based on normalized inter-class variance and normalized precision-recall balance score, which may specifically include:
[0174] Calculate the target sum for each undetermined threshold. The target sum is the sum of the normalized inter-class variance and the normalized precision-recall equilibrium score of the undetermined threshold.
[0175] The threshold with the largest target value is determined as the first candidate threshold.
[0176] The target threshold is determined from at least one first candidate threshold.
[0177] According to the above implementation method, the undetermined threshold with the largest target sum can be considered as a suitable voiceprint determination threshold determined by integrating multiple indicators in the inter-class variance and precision-recall balance score. Determining the target threshold from at least one first candidate threshold is beneficial to improving the accuracy of voiceprint recognition.
[0178] In one embodiment, determining the target threshold from at least one first candidate threshold may specifically include:
[0179] When the number of first candidate thresholds is greater than one, the target difference for each first candidate threshold is calculated. The target difference is the absolute value of the difference between the normalized precision-recall equilibrium score and the normalized inter-class variance of the first candidate threshold.
[0180] The first candidate threshold with the smallest target difference is determined as the second candidate threshold.
[0181] The target threshold is determined from at least one second candidate threshold.
[0182] In one example, when the number of first candidate thresholds is equal to one, the target threshold is determined from at least one first candidate threshold, which may specifically include: determining the first candidate threshold as the target threshold.
[0183] According to the above implementation method, when multiple first candidate thresholds exist, the first candidate threshold with the smallest target difference can be determined as the second candidate threshold. The smaller the target difference, the better the consistency of the segmentation effect measured by the first candidate threshold through inter-class variance and precision-recall balance score. Determining the target threshold from at least one second candidate threshold can effectively take into account multiple indicators such as inter-class variance and precision-recall balance score, thereby improving the accuracy of voiceprint recognition.
[0184] In one embodiment, determining the target threshold from at least one second candidate threshold may specifically include:
[0185] If the number of second candidate thresholds is greater than one, the second candidate threshold with the largest value among at least one second candidate threshold shall be determined as the target threshold.
[0186] In one example, when the number of second candidate thresholds is equal to one, the target threshold is determined from at least one second candidate threshold, which may specifically include: determining the second candidate threshold as the target threshold.
[0187] According to the above implementation method, when multiple second candidate thresholds exist, the second candidate threshold with the largest value can be determined as the target threshold. This allows for more rigorous voiceprint recognition, reducing the risk of a non-first user being identified as the first user. Thus, when a third-party user impersonates the first user, accurate identity verification can be performed using the voiceprint determination threshold, reducing the risk of the third-party user being identified as the first user. Therefore, in scenarios such as fraud prevention, timely warning information can be issued when a third-party user impersonates the first user, reducing the risk of financial loss for legitimate users.
[0188] In one embodiment, after obtaining the first registered voiceprint of the first user in response to the first user's voiceprint registration request, the method may further include performing the following steps for each second user:
[0189] Obtain an updated dataset, which includes at least one updated positive sample and at least two updated negative samples. The updated positive sample includes the recording data of the second user, and the at least two updated negative samples include the first registered voiceprint and the second registered voiceprints of the remaining second users.
[0190] The comparison scores between each updated positive sample and each updated negative sample and the first registered voiceprint are calculated to obtain the updated score set.
[0191] Based on the updated score set, multiple pending update thresholds are determined.
[0192] For each pending update threshold, the comparison scores in the update score set are classified using the pending update threshold as the split value, and the inter-class variance and precision-recall balance score corresponding to the classification results are calculated.
[0193] The target update threshold is determined from multiple pending update thresholds based on the inter-class variance and precision-recall balance score corresponding to each pending update threshold.
[0194] Update the voiceprint determination threshold in the second user's voiceprint registration information to the target update threshold.
[0195] According to the above implementation method, when a new user registers their voiceprint, the voiceprint determination threshold for already registered users can be updated and adjusted. This allows for the adaptive acquisition of a new, optimal voiceprint determination threshold for each user, thereby improving the accuracy of voiceprint recognition.
[0196] In one example, the ratio of positive to negative samples is 1:1. Positive samples include recording data of a first user obtained from an actual production recording database. Negative samples include the second registered voiceprint corresponding to a second user obtained from a registered voiceprint database.
[0197] For the first user A, the voiceprints in the registered voiceprint database can be divided into two categories: the first registered voiceprint of the first user A, and the second registered voiceprints of N-1 second users. N-1 recordings of the first user A can be obtained from the actual production recording database, and these, along with the N-1 second registered voiceprints, can be used to create a training dataset. The number of samples in this training dataset is then 2N-2.
[0198] After determining the voiceprint determination threshold of the first user A according to the methods of the above embodiments, the second registered voiceprint of each second user in the registered voiceprint database can be traversed to redetermine the voiceprint determination threshold of each second user. It is understood that when a first user A registers a voiceprint, each second user only needs to add one recording of that second user from the actual production recording database, maintaining a 1:1 ratio of positive to negative samples, to adaptively update the voiceprint determination threshold of the second user.
[0199] In this way, for business scenarios such as one number on two devices or multiple personal numbers, a single voiceprint recognition threshold can be generated or updated for each user for identification. This facilitates the accurate and efficient operation of the voiceprint recognition system, better protecting users of multiple personal numbers from fraud risks and improving user experience.
[0200] In one embodiment, after using the target threshold as the voiceprint determination threshold for the first user, the method may further include:
[0201] In response to a user identification request, obtain the third voiceprint data of the user to be identified.
[0202] Calculate the comparison score between the third voiceprint data and the first registered voiceprint.
[0203] If the score of the comparison between the third voiceprint data and the first registered voiceprint is greater than the voiceprint determination threshold of the first user, the user to be identified is determined to be the first user.
[0204] In the above embodiments, the user identification request may include requests generated when a user binds a secondary device or handles other services requiring identity authentication, or requests automatically generated by the operating platform during risk identification. Correspondingly, the third voiceprint data may include voiceprint data entered by the user during authentication, or voiceprint data obtained by the operating platform from the user during interaction.
[0205] In one example, the user identification request may include a user identity identifier. If the user identity identifier is a first user identity identifier, a comparison score between the third voiceprint data and the first registered voiceprint can be calculated. In another example, the user identification request may not include a user identity identifier. Instead, it may iterate through the registered voiceprints in the registered voiceprint database, compare the comparison score of each registered voiceprint with the third voiceprint data, and combine the voiceprint determination threshold of each registered voiceprint to perform user identification.
[0206] According to the above implementation method, by calculating the comparison score between the third voiceprint data and the first registered voiceprint, and combining it with the voiceprint determination threshold of the first user, it is possible to identify whether the user to be identified is the first user. In this way, the accuracy and precision of user identification can be improved.
[0207] Based on the same inventive concept as the voiceprint determination threshold determination method, this application also provides a voiceprint determination threshold determination device.
[0208] like Figure 2 As shown, the voiceprint determination threshold determination device 200 may include a first acquisition module 201, a first calculation module 202, a first determination module 203, and a first processing module 204.
[0209] The first acquisition module 201 is used to acquire the first registered voiceprint of the first user in response to the voiceprint registration request of the first user.
[0210] The first acquisition module 201 is also used to acquire a training dataset, which includes at least one positive sample and at least one negative sample. The positive sample includes the first voiceprint data of the first user, and the negative sample includes the second voiceprint data of the second user.
[0211] The first calculation module 202 is used to calculate the comparison scores of each positive sample and each negative sample with the first registered voiceprint, and obtain a comparison score set.
[0212] The first determining module 203 is used to determine multiple pending thresholds based on the comparison score set.
[0213] The first processing module 204 is used to classify the comparison scores in the comparison score set with the undetermined threshold as the segmentation value for each undetermined threshold, and to calculate the inter-class variance and precision-recall balance score corresponding to the classification result.
[0214] The first determining module 203 is also used to determine the target threshold from multiple undetermined thresholds based on the inter-class variance and precision-recall balance score corresponding to each undetermined threshold.
[0215] The first processing module 204 is also used to use the target threshold as the voiceprint determination threshold of the first user.
[0216] The voiceprint determination threshold determination device of this application embodiment obtains the first registered voiceprint of the first user in response to the voiceprint registration request of the first user. Next, it obtains a training dataset, which includes at least one positive sample and at least one negative sample. The positive sample includes the first voiceprint data of the first user, and the negative sample includes the second voiceprint data corresponding to the second user. By using the voiceprint data of the first user as the positive sample and the voiceprint data of other users as the negative sample, the reliability of the training dataset can be improved. Subsequently, the comparison scores of each positive sample and each negative sample with the first registered voiceprint are calculated to obtain a comparison score set. Based on the comparison score set, multiple undetermined thresholds are determined. For each undetermined threshold, the comparison scores in the comparison score set are classified using the undetermined threshold as a segmentation value, and the inter-class variance and precision-recall balance score corresponding to the classification result are calculated. Based on the inter-class variance and precision-recall balance score corresponding to each undetermined threshold, a target threshold is determined from the multiple undetermined thresholds. The target threshold is used as the voiceprint determination threshold for the first user. In this way, the inter-class variance algorithm and precision-recall balance score can be combined to balance multiple selection indicators and adaptively determine the unique voiceprint judgment threshold for each user based on different registered voiceprints. Thus, the determined voiceprint judgment threshold can adapt to actual production and business environments. When applied to scenarios such as dual-device use or multiple personal phone numbers, it can achieve fast and accurate user identity verification based on the actual input voiceprint data and the determined voiceprint judgment threshold.
[0217] In one embodiment, the first voiceprint data can be the recording data of a first user. The second user can include a registered user, and the second voiceprint data can be the second user's second registered voiceprint.
[0218] The first acquisition module is used to acquire the training dataset, and may specifically include:
[0219] The acquisition submodule is used to retrieve at least one recording of the first user from the actual production recording database.
[0220] The acquisition submodule is also used to retrieve at least one second registered voiceprint corresponding to the second user from the registered voiceprint database.
[0221] Create a submodule to create a training dataset based on at least one recording of a first user and at least one second registered voiceprint of a second user.
[0222] In one embodiment, the first calculation module is used to calculate the comparison scores between each positive sample and each negative sample and the first registered voiceprint, respectively, to obtain a comparison score set, which may specifically include:
[0223] The first calculation submodule is used to calculate the similarity between each positive sample and each negative sample and the first registered voiceprint.
[0224] The first determining submodule is used to determine the comparison score of each positive sample and each negative sample with the first registered voiceprint based on the similarity between each positive sample and each negative sample and the first registered voiceprint, and obtain the comparison score set.
[0225] In one embodiment, the first determining module is used to determine multiple pending thresholds based on a set of comparison scores, which may specifically include:
[0226] The second determining submodule is used to determine the range of values for the undetermined threshold based on the maximum and minimum values in the comparison score set.
[0227] The second determining submodule is also used to determine multiple undetermined thresholds from the value range.
[0228] In one embodiment, the first determining module is used to determine multiple undetermined thresholds from a value range, which may specifically include:
[0229] The third determination submodule is used to determine multiple undetermined thresholds within a value range according to preset step values.
[0230] In one embodiment, the classification result may include a first set of scores less than or equal to a pending threshold and a second set of scores greater than the pending threshold.
[0231] The first processing module is used to calculate the inter-class variance corresponding to the classification results, and may specifically include:
[0232] The second calculation submodule is used to calculate the first average value of each comparison score in the first rating set and the first proportion of the first rating set, wherein the first proportion is the ratio of the number of comparison scores in the first rating set to the number of comparison scores in the comparison rating set.
[0233] The second calculation submodule is also used to calculate the second average value of each comparison score in the second rating set, and the second proportion of the second rating set, wherein the second proportion is the ratio of the number of comparison scores in the second rating set to the number of comparison scores in the comparison rating set.
[0234] The second calculation submodule is also used to calculate the first square difference and the second square difference respectively, wherein the first square difference is the square of the difference between the first average and the third average, the second square difference is the square of the difference between the second average and the third average, and the third average is the average of the comparison scores in the comparison score set.
[0235] The second calculation submodule is also used to calculate the product of the first squared difference and the first proportion to obtain the first product.
[0236] The second calculation submodule is also used to calculate the product of the second squared difference and the second proportion to obtain the second product.
[0237] The second calculation submodule is also used to add the first product and the second product to obtain the inter-class variance corresponding to the classification result.
[0238] In one embodiment, the first processing module is used to calculate the precision-recall balance score corresponding to the classification result, which may specifically include:
[0239] The fourth determination submodule is used to determine the first number of comparison scores corresponding to positive samples in the second score set.
[0240] The third calculation submodule is used to calculate the ratio of the number of comparison scores in the first set to the number of comparison scores in the second set, and to obtain the accuracy corresponding to the classification result.
[0241] The third calculation submodule is also used to calculate the ratio of the first quantity to the second quantity to obtain the recall rate corresponding to the classification result, where the second quantity is the number of positive samples in the training dataset.
[0242] The third calculation submodule is also used to calculate the precision-recall balance score corresponding to the classification result based on the precision and recall, according to the preset precision weight and recall weight.
[0243] In one embodiment, the first determining module is used to determine a target threshold from a plurality of undetermined thresholds based on the inter-class variance and precision-recall balance score corresponding to each undetermined threshold, specifically including:
[0244] The processing submodule is used to normalize the inter-class variance and precision-recall balance score corresponding to each undetermined threshold, so as to obtain the normalized inter-class variance and normalized precision-recall balance score corresponding to each undetermined threshold.
[0245] The fifth determination submodule is used to determine the target threshold from multiple undetermined thresholds based on the normalized inter-class variance and the normalized precision recall balanced score.
[0246] In one embodiment, the fifth determining submodule is used to determine a target threshold from multiple undetermined thresholds based on the normalized inter-class variance and the normalized precision-recall equilibrium score, specifically including:
[0247] The calculation unit is used to calculate the target sum for each undetermined threshold. The target sum is the sum of the normalized inter-class variance and the normalized precision-recall equilibrium score of the undetermined threshold.
[0248] The determination unit is used to determine the undetermined threshold with the largest target value as the first candidate threshold.
[0249] The determining unit is also used to determine the target threshold from at least one first candidate threshold.
[0250] In one embodiment, the determining unit is configured to determine a target threshold from at least one first candidate threshold, which may specifically include:
[0251] The calculation subunit is used to calculate the target difference of each first candidate threshold when the number of first candidate thresholds is greater than one. The target difference is the absolute value of the difference between the normalized precision-recall equilibrium score and the normalized inter-class variance of the first candidate threshold.
[0252] A sub-unit is determined to identify the first candidate threshold with the smallest target difference as the second candidate threshold.
[0253] The sub-unit is also used to determine the target threshold from at least one second candidate threshold.
[0254] In one embodiment, the determining subunit is used to determine the target threshold from at least one second candidate threshold, which may specifically include:
[0255] If the number of second candidate thresholds is greater than one, the second candidate threshold with the largest value among at least one second candidate threshold shall be determined as the target threshold.
[0256] In one embodiment, the apparatus may further include:
[0257] The second acquisition module is used to acquire the first registered voiceprint of the first user in response to the voiceprint registration request of the first user, and then acquire an updated dataset for each second user. The updated dataset includes at least one updated positive sample and at least two updated negative samples. The updated positive sample includes the recording data of the second user, and the at least two updated negative samples include the first registered voiceprint and the second registered voiceprints of the remaining second users.
[0258] The second calculation module is used to calculate the comparison scores between each updated positive sample and each updated negative sample and the first registered voiceprint, so as to obtain the updated score set.
[0259] The second determination module is used to determine multiple pending update thresholds based on the updated score set.
[0260] The second processing module is used to classify the comparison scores in the updated score set for each pending update threshold, using the pending update threshold as the segmentation value, and calculate the inter-class variance and precision-recall balance score corresponding to the classification results.
[0261] The second determining module is also used to determine the target update threshold from multiple pending update thresholds based on the inter-class variance and precision-recall balance score corresponding to each pending update threshold.
[0262] The update module is used to update the voiceprint determination threshold in the voiceprint registration information of the second user to the target update threshold.
[0263] In one embodiment, the apparatus may further include:
[0264] The third acquisition module is used to obtain the third voiceprint data of the user to be identified after using the target threshold as the voiceprint determination threshold of the first user and responding to the user identification request.
[0265] The third calculation module is used to calculate the comparison score between the third voiceprint data and the first registered voiceprint.
[0266] The third determining module is used to determine the user to be identified as the first user if the comparison score between the third voiceprint data and the first registered voiceprint is greater than the voiceprint determination threshold of the first user.
[0267] The voiceprint determination threshold determination device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0268] Figure 3 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.
[0269] An electronic device may include a processor 301 and a memory 302 storing computer program instructions.
[0270] Specifically, the processor 301 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0271] Memory 302 may include mass storage for data or instructions. For example, and not limitingly, memory 302 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 302 may include removable or non-removable (or fixed) media. Where appropriate, memory 302 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 302 is non-volatile solid-state memory.
[0272] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to one aspect of this disclosure.
[0273] The processor 301 reads and executes computer program instructions stored in the memory 302 to implement any of the voiceprint determination threshold methods in the above embodiments.
[0274] As an example, the electronic device may also include a communication interface 303 and a bus 310. Wherein, such as Figure 3 As shown, the processor 301, memory 302, and communication interface 303 are connected through bus 310 and complete communication with each other.
[0275] The communication interface 303 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0276] Bus 310 includes hardware, software, or both, that couples components of the voiceprint determination threshold determining device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 310 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.
[0277] The electronic device can execute the voiceprint determination threshold method in the embodiments of this application, thereby achieving a combination of Figure 1 and Figure 2 The method and apparatus for determining the voiceprint recognition threshold are described.
[0278] Furthermore, in conjunction with the voiceprint determination threshold method in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the voiceprint determination threshold methods in the above embodiments.
[0279] This application also provides a computer program product, including a computer program, which, when executed, implements any of the methods for determining the voiceprint determination threshold in the above embodiments.
[0280] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0281] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0282] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0283] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0284] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A method for determining a voiceprint recognition threshold, characterized in that, include: In response to the first user's voiceprint registration request, obtain the first registered voiceprint of the first user; Obtain a training dataset, which includes at least one positive sample and at least one negative sample. The positive sample includes the first voiceprint data of the first user, and the negative sample includes the second voiceprint data of the second user. The comparison scores of each positive sample and each negative sample with the first registered voiceprint are calculated to obtain the comparison score set. Based on the comparison score set, multiple undetermined thresholds are determined; For each undetermined threshold, the comparison scores in the comparison score set are classified using the undetermined threshold as the segmentation value, and the inter-class variance and precision-recall balance score corresponding to the classification result are calculated. Based on the inter-class variance and precision-recall balance score corresponding to each undetermined threshold, the target threshold is determined from the plurality of undetermined thresholds; The target threshold is used as the voiceprint determination threshold for the first user.
2. The method according to claim 1, characterized in that, The first voiceprint data is the recording data of the first user; the second user includes registered users, and the second voiceprint data is the second registered voiceprint of the second user; The acquisition of the training dataset includes: Obtain at least one recording from the actual production recording database; Retrieve at least one second registered voiceprint corresponding to the second user from the registered voiceprint database; The training dataset is created based on the recording data of at least one first user and the second registered voiceprint corresponding to at least one second user.
3. The method according to claim 1 or 2, characterized in that, The comparison scores of each positive sample and each negative sample with the first registered voiceprint are calculated respectively to obtain a comparison score set, including: Calculate the similarity between each positive sample and each negative sample and the first registered voiceprint. Based on the similarity between each positive sample and each negative sample and the first registered voiceprint, the comparison score of each positive sample and each negative sample and the first registered voiceprint is determined, and a comparison score set is obtained.
4. The method according to claim 1 or 2, characterized in that, The step of determining multiple undetermined thresholds based on the comparison score set includes: The range of values for the undetermined threshold is determined based on the maximum and minimum values in the comparison score set. The plurality of undetermined thresholds are determined from the value range.
5. The method according to claim 4, characterized in that, Determining the plurality of undetermined thresholds from the value range includes: According to the preset step value, the plurality of undetermined thresholds are determined within the value range.
6. The method according to claim 1 or 2, characterized in that, The classification results include a first set of scores less than or equal to the undetermined threshold and a second set of scores greater than the undetermined threshold; Calculate the inter-class variance corresponding to the classification results, including: Calculate the first average value of each comparative score in the first rating set, and the first proportion of the first rating set, wherein the first proportion is the ratio of the number of comparative scores in the first rating set to the number of comparative scores in the comparative rating set; Calculate the second average value of each comparative score in the second rating set, and the second proportion of the second rating set, wherein the second proportion is the ratio of the number of comparative scores in the second rating set to the number of comparative scores in the comparative rating set; Calculate the first square difference and the second square difference respectively, wherein the first square difference is the square of the difference between the first average and the third average, the second square difference is the square of the difference between the second average and the third average, and the third average is the average of the comparison scores in the comparison score set; Calculate the product of the first squared difference and the first proportion to obtain the first product; Calculate the product of the second squared difference and the second proportion to obtain the second product; Add the first product to the second product to obtain the inter-class variance corresponding to the classification result.
7. The method according to claim 6, characterized in that, The precision-recall balance score corresponding to the classification result is calculated, including: Determine the first number of comparison scores in the second set that correspond to the positive example sample; Calculate the ratio of the first number to the number of comparison scores in the second rating set to obtain the accuracy corresponding to the classification result; Calculate the ratio of the first quantity to the second quantity to obtain the recall rate corresponding to the classification result, wherein the second quantity is the number of positive samples in the training dataset; Based on the precision and recall, and according to the preset precision weight and recall weight, the precision-recall balance score corresponding to the classification result is calculated.
8. The method according to claim 1 or 2, characterized in that, The step of determining the target threshold from the plurality of undetermined thresholds based on the inter-class variance and precision-recall balance score corresponding to each undetermined threshold includes: Normalize the inter-class variance and precision-recall balance score corresponding to each undetermined threshold to obtain the normalized inter-class variance and normalized precision-recall balance score corresponding to each undetermined threshold. The target threshold is determined from the plurality of undetermined thresholds based on the normalized inter-class variance and the normalized precision-recall balance score.
9. The method according to claim 8, characterized in that, The step of determining the target threshold from the plurality of undetermined thresholds based on the normalized inter-class variance and the normalized precision-recall balance score includes: Calculate the target sum for each undetermined threshold, where the target sum is the sum of the normalized inter-class variance and the normalized precision-recall equilibrium score of the undetermined threshold. The threshold with the largest target value is determined as the first candidate threshold. The target threshold is determined from at least one of the first candidate thresholds.
10. The method according to claim 9, characterized in that, Determining the target threshold from at least one of the first candidate thresholds includes: When the number of first candidate thresholds is greater than one, the target difference of each first candidate threshold is calculated. The target difference is the absolute value of the difference between the normalized precision recall equilibrium score and the normalized inter-class variance of the first candidate threshold. The first candidate threshold with the smallest target difference is determined as the second candidate threshold; The target threshold is determined from at least one of the second candidate thresholds.
11. The method according to claim 10, characterized in that, Determining the target threshold from at least one of the second candidate thresholds includes: If the number of second candidate thresholds is greater than one, the second candidate threshold with the largest value among the at least one second candidate thresholds shall be determined as the target threshold.
12. The method according to claim 2, characterized in that, After obtaining the first registered voiceprint of the first user in response to the first user's voiceprint registration request, the method further includes performing the following steps for each second user: Obtain an updated dataset, the updated dataset including at least one updated positive sample and at least two updated negative samples, wherein the updated positive sample includes the recording data of the second user, and the at least two updated negative samples include the first registered voiceprint and the second registered voiceprints of the remaining second users; The comparison scores between each updated positive sample and each updated negative sample and the first registered voiceprint are calculated to obtain the updated score set. Based on the updated score set, multiple pending update thresholds are determined; For each pending update threshold, the comparison scores in the update score set are classified using the pending update threshold as the segmentation value, and the inter-class variance and precision-recall balance score corresponding to the classification results are calculated. Based on the inter-class variance and precision-recall balance score corresponding to each pending update threshold, the target update threshold is determined from the plurality of pending update thresholds. Update the voiceprint determination threshold in the second user's voiceprint registration information to the target update threshold.
13. The method according to claim 1 or 2, characterized in that, After using the target threshold as the voiceprint determination threshold for the first user, the method further includes: In response to a user identification request, obtain the third voiceprint data of the user to be identified; Calculate the comparison score between the third voiceprint data and the first registered voiceprint; If the score of the comparison between the third voiceprint data and the first registered voiceprint is greater than the voiceprint determination threshold of the first user, the user to be identified is determined to be the first user.
14. A device for determining a voiceprint recognition threshold, characterized in that, include: The acquisition module is used to acquire the first registered voiceprint of the first user in response to the first user's voiceprint registration request; The acquisition module is further configured to acquire a training dataset, the training dataset including at least one positive sample and at least one negative sample, the positive sample including the first voiceprint data of the first user, and the negative sample including the second voiceprint data corresponding to the second user; The calculation module is used to calculate the comparison score between each positive sample and each negative sample and the first registered voiceprint, so as to obtain a comparison score set. The determining module is used to determine multiple pending thresholds based on the comparison score set; The processing module is used to classify the comparison scores in the comparison score set for each undetermined threshold, using the undetermined threshold as the segmentation value, and calculate the inter-class variance and precision-recall balance score corresponding to the classification result. The determining module is further configured to determine the target threshold from the plurality of undetermined thresholds based on the inter-class variance and precision-recall balance score corresponding to each undetermined threshold. The processing module is further configured to use the target threshold as the voiceprint determination threshold for the first user.
15. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the voiceprint determination threshold determination method as described in any one of claims 1-13.
Citation Information
Patent Citations
Similarity threshold determination method and device, equipment and storage medium
CN114120383A
Voiceprint template updating method and related equipment
CN116168708A