Voiceprint recognition method, device, equipment and storage medium based on twin voiceprint pairs

Through the voiceprint recognition method based on twin voiceprint pairs, the speech feature set is extracted and group matching is performed, which solves the problem of inefficient individual recognition in the prior art, and achieves rapid and accurate identification of criminal gangs.

CN115376516BActive Publication Date: 2025-08-26CHINA MOBILE COMM GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110514062.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-11
Publication Date
2025-08-26
Estimated Expiration
2041-05-11

AI Technical Summary

Technical Problem

The existing voice recognition technology mainly targets individuals, which is time-consuming and labor-intensive, inefficient, and it is difficult to effectively deal with contactless fraud by criminal gangs.

Method used

The voiceprint recognition method based on twin voiceprint pairs is adopted. By extracting the voiceprint feature set of voiceprints to be identified, the voiceprint target group to be matched is determined, and the twin voiceprint pairs are used to match, so as to determine the group to which the voice is belonged when the overall coverage information meets the preset conditions, and the recognition efficiency is improved.

Benefits of technology

It realizes rapid and accurate identification of criminal gangs, improves the efficiency of voice group detection, and reduces the time and cost of identification one by one.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115376516B_ABST
    Figure CN115376516B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of speech recognition technology, and discloses a voiceprint recognition method, device, equipment and storage medium based on twin voiceprint pairs. The method comprises: extracting a voiceprint feature set to be recognized of a speech to be recognized, and then determining a target voiceprint group to be matched containing a plurality of voiceprint clusters corresponding to the speech to be recognized; then matching the twin voiceprint pairs with the voiceprint feature set to be recognized according to the voiceprint clusters; finally, determining the overall coverage information of the target voiceprint group to be matched with the voiceprint feature set to be recognized according to the matching results; and when the overall coverage information meets a preset condition, determining that the speech to be recognized belongs to the target voiceprint group to be matched. Since voiceprint recognition of a speaker is performed using the twin voiceprint pairs in the voiceprint clusters contained in the pre-constructed voiceprint group, it is possible to judge whether the speaker belongs to the voiceprint group as a whole, without having to recognize the speech to be recognized one by one according to each speaker in the group, thereby improving recognition efficiency and having obvious advantages in speech group detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition technology, and in particular to a voiceprint recognition method, device, equipment and storage medium based on twin voiceprint pairs. Background Art

[0002] Voiceprint recognition (also known as speaker identification) allows machines to automatically identify a speaker's identity from their speech. Everyone has a unique voiceprint. This is partly because everyone's acoustic organs vary in shape and size, resulting in differences in pitch and timbre. Furthermore, everyone has their own unique speaking habits, resulting in variations in word choice, rhythm, and pronunciation patterns. This uniqueness of voiceprints demonstrates the feasibility of identifying a speaker through their voice.

[0003] Voiceprint technology has been widely used in the judicial and investigative fields. In many major cases, such as ransomware, kidnapping, and terrorist threats, recordings are crucial and the only evidence. By comparing the captured voice with the suspect's voice using voiceprint recognition technology, a relatively objective identification result can be obtained, serving as one of the supporting evidence for judicial decisions.

[0004] With the development of my country's telecommunications network technology, contactless fraud schemes carried out by criminal gangs, using mobile phones, landlines, the internet, and other communication tools and modern technology, have rapidly spread in recent years, causing significant losses to the public. Currently, telecom fraud is becoming increasingly organized, with many criminals operating with clear divisions of labor and responsibility, following a pre-planned "script" to coordinate their operations. This presents significant challenges for operators and regulators.

[0005] Existing technologies are designed and developed only for single speakers. If applied to criminal gangs, they can only be identified one by one, which is time-consuming, labor-intensive, inefficient, and has low accuracy.

[0006] The above content is only used to assist in understanding the technical solution of the present invention and does not constitute an admission that the above content is prior art. Summary of the Invention

[0007] The main purpose of the present invention is to provide a voiceprint recognition method, device, equipment and storage medium based on twin voiceprint pairs, aiming to solve the technical problem that the existing voice recognition technology basically only recognizes and detects individuals, which is time-consuming, labor-intensive and inefficient.

[0008] To achieve the above object, the present invention provides a voiceprint recognition method based on twin voiceprint pairs, the method comprising the following steps:

[0009] Extracting a voiceprint feature set to be recognized of the speech to be recognized;

[0010] Determine a target voiceprint group to be matched corresponding to the speech to be recognized, wherein the target voiceprint group to be matched includes a plurality of voiceprint clusters;

[0011] Performing twin voiceprint pair matching on the voiceprint feature set to be identified according to the voiceprint cluster, and determining overall coverage information of the voiceprint target group to be matched on the voiceprint feature set to be identified according to the matching results;

[0012] When the overall coverage information meets a preset condition, it is determined that the speech to be recognized belongs to the voiceprint target group to be matched.

[0013] Preferably, the step of performing twin voiceprint pair matching on the voiceprint feature set to be identified according to the voiceprint cluster, and determining overall coverage information of the target voiceprint group to be matched on the voiceprint feature set to be identified according to the matching result, includes:

[0014] Traversing the plurality of voiceprint clusters;

[0015] Obtaining the twin voiceprint pairs contained in the currently traversed voiceprint cluster, wherein the twin voiceprint pairs contain at least two speakers and each speaker is pre-configured with a voiceprint recognition model;

[0016] Determining the coverage of the currently traversed voiceprint cluster to the voiceprint feature set to be identified according to the voiceprint recognition model corresponding to the twin voiceprint pair;

[0017] At the end of the traversal, the overall coverage information of the target voiceprint group to be matched on the voiceprint feature set to be identified is determined based on the obtained voiceprint cluster coverage of each voiceprint cluster on the voiceprint feature set to be identified.

[0018] Preferably, the voiceprint feature set to be identified includes several voiceprint features to be matched;

[0019] The step of determining, based on the voiceprint recognition model corresponding to the twin voiceprint pair, the coverage of the voiceprint cluster of the voiceprint feature set to be identified by the currently traversed voiceprint cluster, includes:

[0020] Traversing the voiceprint feature set to be identified to obtain the currently traversed voiceprint feature to be matched;

[0021] Obtaining a hit status of the currently traversed voiceprint feature to be matched according to the voiceprint recognition model corresponding to the twin voiceprint pair, wherein the hit status includes a hit or a miss;

[0022] At the end of traversal of the voiceprint feature set to be identified, counting the proportion of the voiceprint features to be matched that are hit in the voiceprint feature set to be identified in the voiceprint feature set to be identified;

[0023] The coverage of the voiceprint clusters of the to-be-identified voiceprint feature set by the currently traversed voiceprint cluster is determined according to the proportion.

[0024] Preferably, the step of obtaining a hit status of the currently traversed voiceprint feature to be matched according to the voiceprint recognition model corresponding to the twin voiceprint pair includes:

[0025] Calculate the model matching score corresponding to the currently traversed voiceprint feature to be matched according to different voiceprint recognition models corresponding to the twin voiceprint pair;

[0026] Compare the calculated model matching score with the initial threshold value;

[0027] If there is a model matching score less than the initial threshold value among the calculated model matching scores, it is determined that the currently traversed voiceprint cluster does not hit the currently traversed voiceprint feature to be matched;

[0028] If there is no model matching score less than the initial threshold value among the calculated model matching scores, selecting the maximum model matching score from the calculated model matching scores;

[0029] Comparing the maximum model matching score with a preset decision threshold;

[0030] If the maximum model matching score is greater than or equal to the preset decision threshold, it is determined that the currently traversed voiceprint cluster hits the currently traversed voiceprint feature to be matched;

[0031] If the maximum model matching score is less than the preset decision threshold, it is determined that the currently traversed voiceprint cluster does not hit the currently traversed voiceprint feature to be matched.

[0032] Preferably, after the step of determining whether the currently traversed voiceprint cluster does not match the currently traversed voiceprint feature to be matched, the method further comprises:

[0033] The remaining voiceprint clusters in the plurality of voiceprint clusters are traversed, and the step of obtaining the twin voiceprint pairs contained in the currently traversed voiceprint cluster is returned.

[0034] Preferably, before the step of extracting the voiceprint feature set to be recognized of the speech to be recognized, the method further comprises:

[0035] Obtaining a speech sample of each speaker in the speaker group;

[0036] Modeling each speaker and training the constructed model based on the voice sample to obtain a voiceprint recognition model;

[0037] In the speaker group, using the voiceprint recognition model corresponding to each speaker one by one to calculate the model matching score corresponding to the voiceprint feature of each speaker;

[0038] The twin voiceprint pairs of the speaker group are determined according to the calculated model matching scores, and a voiceprint target group of the speaker group is constructed according to the twin voiceprint pairs.

[0039] Preferably, the step of determining the twin voiceprint pairs of the speaker group based on the calculated model matching scores and constructing the voiceprint target group of the speaker group based on the twin voiceprint pairs includes:

[0040] Construct a model matching score set corresponding to each voiceprint recognition model based on the calculated model matching score;

[0041] Traversing the speaker group and obtaining the target voiceprint recognition model and target voiceprint features corresponding to the currently traversed speaker;

[0042] Searching for a target model matching score set corresponding to the target voiceprint recognition model from the model matching score set;

[0043] Reading the target model matching score corresponding to the target voiceprint feature from the target model matching score set;

[0044] Searching for a maximum model matching score other than the target model matching score in the target model matching score set;

[0045] Determine the target speaker to which the maximum model matching score belongs, and construct a twin voiceprint pair based on the currently traversed speaker and the target speaker;

[0046] When the traversal of the speaker group is completed, obtaining a twin voiceprint pair corresponding to each speaker;

[0047] Construct a voiceprint cluster to which each twin voiceprint pair belongs, and construct a voiceprint target group of the speaker group based on all the voiceprint clusters.

[0048] In addition, to achieve the above-mentioned purpose, the present invention also proposes a voiceprint recognition device based on twin voiceprint pairs, the device comprising:

[0049] A feature extraction module is used to extract a voiceprint feature set to be recognized from the speech to be recognized;

[0050] A voiceprint group acquisition module is used to determine a target voiceprint group to be matched corresponding to the speech to be recognized, wherein the target voiceprint group to be matched includes a plurality of voiceprint clusters;

[0051] a voiceprint pair matching module, configured to perform twin voiceprint pair matching on the voiceprint feature set to be identified based on the voiceprint cluster, and determine overall coverage information of the voiceprint target group to be matched on the voiceprint feature set to be identified based on the matching results;

[0052] The result determination module is configured to determine that the speech to be recognized belongs to the voiceprint target group to be matched when the overall coverage information meets a preset condition.

[0053] In addition, to achieve the above-mentioned purpose, the present invention also proposes a voiceprint recognition device based on twin voiceprint pairs, which includes: a memory, a processor, and a voiceprint recognition program based on twin voiceprint pairs stored on the memory and runnable on the processor, and the voiceprint recognition program based on twin voiceprint pairs is configured to implement the steps of the voiceprint recognition method based on twin voiceprint pairs as described above.

[0054] In addition, to achieve the above-mentioned purpose, the present invention also proposes a storage medium, on which a voiceprint recognition program based on twin voiceprint pairs is stored. When the voiceprint recognition program based on twin voiceprint pairs is executed by a processor, the steps of the voiceprint recognition method based on twin voiceprint pairs as described above are implemented.

[0055] The present invention extracts the voiceprint feature set to be recognized of the speech to be recognized, and then determines the target voiceprint group to be matched containing several voiceprint clusters corresponding to the speech to be recognized; then performs twin voiceprint pair matching on the voiceprint feature set to be recognized according to the voiceprint clusters, and finally determines the overall coverage information of the target voiceprint group to be matched with the voiceprint feature set to be recognized according to the matching results, and when the overall coverage information meets the preset conditions, it is determined that the speech to be recognized belongs to the target voiceprint group to be matched. Since the voiceprint recognition of the speaker is performed through the twin voiceprint pairs in the voiceprint cluster contained in the pre-constructed voiceprint group, it is possible to judge whether the speaker belongs to the voiceprint group as a whole, and there is no need to recognize the speech to be recognized one by one according to the speakers in the speaker group, thereby improving the recognition efficiency and having obvious advantages in speech group detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 2 is a schematic diagram of the structure of a voiceprint recognition device based on twin voiceprint pairs in the hardware operating environment involved in the embodiment of the present invention;

[0057] Figure 2 2. It is a flowchart of the first embodiment of the voiceprint recognition method based on twin voiceprint pairs of the present invention;

[0058] Figure 3 2. This is a schematic diagram of the structure of a voiceprint cluster in the first embodiment of the voiceprint recognition method based on twin voiceprint pairs of the present invention;

[0059] Figure 42. It is a flow chart of the second embodiment of the voiceprint recognition method based on twin voiceprint pairs of the present invention;

[0060] Figure 5 2. It is a flowchart of a third embodiment of a voiceprint recognition method based on twin voiceprint pairs according to the present invention;

[0061] Figure 6 This is a structural block diagram of the first embodiment of the voiceprint recognition device based on twin voiceprint pairs of the present invention.

[0062] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0063] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0064] Reference Figure 1 , Figure 1 This is a schematic diagram of the structure of a voiceprint recognition device based on twin voiceprint pairs in the hardware operating environment involved in the embodiment of the present invention.

[0065] like Figure 1 As shown, the voiceprint recognition device based on twin voiceprint pairs may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to implement communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard. Optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a wireless fidelity (WI-FI) interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk storage device. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0066] Those skilled in the art will understand that Figure 1 The structure shown in does not constitute a limitation on the voiceprint recognition device based on twin voiceprint pairs, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0067] like Figure 1As shown, the memory 1005 as a storage medium may include an operating system, a data storage module, a network communication module, a user interface module, and a voiceprint recognition program based on twin voiceprint pairs.

[0068] exist Figure 1 In the voiceprint recognition device based on twin voiceprint pairs shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the voiceprint recognition device based on twin voiceprint pairs of the present invention can be set in the voiceprint recognition device based on twin voiceprint pairs, and the voiceprint recognition device based on twin voiceprint pairs calls the voiceprint recognition program based on twin voiceprint pairs stored in the memory 1005 through the processor 1001, and executes the voiceprint recognition method based on twin voiceprint pairs provided by the embodiment of the present invention.

[0069] The embodiment of the present invention provides a voiceprint recognition method based on twin voiceprint pairs. Figure 2 , Figure 2 This is a flow chart of the first embodiment of the voiceprint recognition method based on twin voiceprint pairs of the present invention.

[0070] In this embodiment, the voiceprint recognition method based on twin voiceprint pairs includes the following steps:

[0071] Step S10: extracting a voiceprint feature set of the speech to be recognized;

[0072] It should be noted that the execution subject of the method of this embodiment can be a computing service device with voice collection, data processing, network communication and program running functions, such as a smart phone, tablet computer, personal computer, etc., or other electronic devices with the same or similar functions. This embodiment does not limit this.

[0073] It should be understood that the speech to be recognized can be a collected user voice, for example, when the user is on a phone call, the user's voice can be collected to obtain the speech to be recognized. The voiceprint feature set to be recognized can be a voiceprint feature set obtained after voiceprint feature extraction of the speech to be recognized, and the set can contain multiple voiceprint features to be matched, such as [x1, x2, ..., xn].

[0074] Step S20: determining a target voiceprint group to be matched corresponding to the speech to be recognized, wherein the target voiceprint group to be matched includes a plurality of voiceprint clusters;

[0075] It should be noted that before executing the above step S10 of this embodiment, it is necessary to first construct a voiceprint group corresponding to the speech to be recognized, that is, the above-mentioned voiceprint target group to be matched. Generally, there may be multiple voiceprint clusters in a voiceprint group, or there may be only one voiceprint cluster. The number of voiceprint clusters varies according to actual conditions.

[0076] In this embodiment, the voiceprint group can be constructed by first collecting voice samples from all speakers in a group of speakers (for example, a fraud phone gang), then modeling and identifying the speakers in the group, determining (twin) voiceprint clusters based on the identification results, and finally constructing voiceprint groups based on the (twin) voiceprint clusters. In actual applications, the voiceprint groups can be constructed for fraud gangs in different regions based on the call records of investigated or suspected fraud phone gangs, so as to subsequently perform voiceprint identification on individuals in the fraud phone gangs.

[0077] It should be emphasized that due to the certain dispersion of the distribution of fraudulent telephone gangs, and the different voiceprint groups of fraudulent telephone gangs are different, for example, the voiceprint group of group A and the voiceprint group of group B in different regions are basically impossible to be the same. If the currently collected voice to be identified originates from group A, then it is obviously inaccurate to identify it using the voiceprint group of group B. Therefore, in this step, the method of determining the target group of voiceprints to be matched corresponding to the voice to be identified can be: first determine the voice source of the voice to be identified, then query the call terminal corresponding to the voice source (such as a mobile phone, computer, smart wearable device, etc.), and then determine the fraudulent telephone gang closest to it based on the location of the call terminal (which can be a geographical location, IP address, etc.) or the daily activity area, and finally obtain the voiceprint group of the nearest fraudulent telephone gang as the target group of voiceprints to be matched.

[0078] Step S30: performing twin voiceprint pair matching on the voiceprint feature set to be identified according to the voiceprint cluster, and determining overall coverage information of the voiceprint target group to be matched on the voiceprint feature set to be identified according to the matching result;

[0079] It should be noted that a voiceprint cluster is composed of twin voiceprint pairs, which can be determined by the (pre-built and trained) voiceprint recognition model corresponding to the speaker. Specifically, in the group where the speaker is located, the voiceprint features of each speaker can be scored one by one using the voiceprint recognition model of each speaker; then for each voiceprint recognition model, the voiceprint features of other speakers m' with the highest voiceprint feature scores other than the current speaker m are searched in the group based on the scoring results. At this time, speaker m' can be called the twin voiceprint pair of speaker m. Figure 3 , Figure 3 FIG1 is a structural diagram of a voiceprint cluster in the first embodiment of the voiceprint recognition method based on twin voiceprint pairs of the present invention. Figure 3As shown, the voiceprint cluster includes three voiceprint clusters, in which the speaker m in one voiceprint cluster is a twin voiceprint pair of speaker m'.

[0080] In actual applications, by traversing the speaker group in the above manner, all twin voiceprint pairs in the speaker group can be obtained, and then the twin voiceprint pairs are analyzed. If in the twin voiceprint pair (m and m'), speaker m only has a twin voiceprint pair relationship with speaker m' in the twin voiceprint pair, and does not have a twin voiceprint pair relationship with other speakers outside the twin voiceprint pair, then m and m' are said to constitute a voiceprint cluster.

[0081] It should be understood that the voiceprint recognition model constructed for the speaker in this embodiment can select any one of the GMM mean supervector model, the identity authentication vector (i-vector) based model, and the x-vector model (a deep neural network trained to distinguish speakers, mapping variable-length speech to fixed-dimensional embedding). The selection of a specific model is not limited in this embodiment and the following embodiments.

[0082] In a specific implementation, after obtaining the voiceprint feature set to be identified, the voiceprint feature set to be identified can be matched with the twin voiceprint pairs contained in the voiceprint cluster. This can be done by first obtaining the (pre-built and trained) voiceprint recognition model corresponding to each speaker in the twin voiceprint pair, and then scoring the voiceprint features to be matched [x1, x2, ... xn] in the voiceprint feature set to be identified based on the voiceprint recognition model. The scoring results are then compared with a pre-set threshold value, and then based on the comparison results, it is determined whether the voiceprint cluster hits the voiceprint features to be matched. If the hit rate of the voiceprint features contained in the voiceprint feature set to be identified [x1, x2, ... xn] for all speakers in the twin voiceprint pairs of a voiceprint cluster exceeds a set threshold (e.g., 80%), then it is determined that the voiceprint cluster has met the coverage requirements for the voiceprint feature set to be identified.

[0083] For example, voiceprint cluster c contains a twin voiceprint pair (m and m'), and the hit rate of the voiceprint feature set to be identified [x1, x2, ... x10] is 80%, that is, 8 out of 10 voiceprint features are hit by voiceprint cluster c, then it is determined that the coverage of the voiceprint feature set to be identified [x1, x2, ... x10] meets the requirements.

[0084] It is understandable that after obtaining the coverage of each voiceprint cluster in the voiceprint group on the voiceprint feature set to be identified, the overall coverage information of the voiceprint target group to be matched on the voiceprint feature set to be identified can be determined.

[0085] Step S40: When the overall coverage information meets a preset condition, it is determined that the speech to be recognized belongs to the voiceprint target group to be matched.

[0086] It should be noted that the above-mentioned preset conditions in this embodiment can be set according to actual conditions. For example, when the number of voiceprint clusters in the entire voiceprint group that meet the coverage standard of the voiceprint feature set to be identified exceeds a set percentage (for example, 80%) of the total number of voiceprint clusters in the entire voiceprint group, it is determined that the overall coverage information meets the preset conditions, that is, the speech to be identified belongs to the target voiceprint group to be matched, thereby achieving a quick judgment on whether the speaker belongs to a specific target group.

[0087] This embodiment extracts the voiceprint feature set to be recognized of the speech to be recognized, and then determines the target voiceprint group to be matched containing several voiceprint clusters corresponding to the speech to be recognized; then performs twin voiceprint pair matching on the voiceprint feature set to be recognized according to the voiceprint clusters, and finally determines the overall coverage information of the target voiceprint group to be matched with the voiceprint feature set to be recognized based on the matching results, and when the overall coverage information meets the preset conditions, it is determined that the speech to be recognized belongs to the target voiceprint group to be matched. Since the voiceprint recognition of the speaker is performed through the twin voiceprint pairs in the voiceprint cluster contained in the pre-constructed voiceprint group, it is possible to judge whether the speaker belongs to the voiceprint group as a whole, and there is no need to recognize the speech to be recognized one by one according to the speakers in the speaker group, thereby improving the recognition efficiency and having obvious advantages in speech group detection.

[0088] refer to Figure 4 , Figure 4 This is a flow chart of the second embodiment of the voiceprint recognition method based on twin voiceprint pairs of the present invention.

[0089] Based on the above first embodiment, in this embodiment, step S30 includes:

[0090] Step S301: traversing the plurality of voiceprint clusters;

[0091] It should be understood that there may be several voiceprint clusters in the target voiceprint group to be matched. In order to accurately obtain the coverage of each voiceprint cluster on the voiceprint feature set to be identified, this embodiment adopts a traversal method to select one voiceprint cluster at a time from several voiceprint clusters to perform the following operations of this embodiment.

[0092] Step S302: obtaining the twin voiceprint pairs contained in the currently traversed voiceprint cluster, wherein the twin voiceprint pairs contain at least two speakers and each speaker is pre-configured with a voiceprint recognition model;

[0093] As described in the first embodiment, voiceprint clusters are composed of twin voiceprint pairs. Therefore, each twin voiceprint pair contains at least two speakers, and each speaker has a pre-configured and trained voiceprint recognition model during the voiceprint cluster construction phase. For example, in a twin voiceprint pair (m and m'), speaker m has a pre-configured voiceprint recognition model A1, and speaker m' has a pre-configured voiceprint recognition model A2, and so on.

[0094] Step S303: determining the coverage of the currently traversed voiceprint cluster to the voiceprint feature set to be identified based on the voiceprint recognition model corresponding to the twin voiceprint pair;

[0095] It should be noted that the coverage of a voiceprint cluster refers to the hit rate of a voiceprint cluster on the voiceprint features to be matched in the voiceprint feature set to be identified. If a voiceprint feature to be matched has no score less than the initial threshold value after being scored by the voiceprint recognition models of all voiceprint twin pairs in a voiceprint cluster, and the highest score exceeds the decision threshold value, then the voiceprint cluster can be considered to have hit the voiceprint feature to be matched. Similarly, if the hit rate of a voiceprint cluster on all voiceprint features to be matched in the voiceprint feature set to be identified exceeds a set threshold (e.g., 80%), then the voiceprint cluster is considered to have met the coverage standard for the voiceprint feature set to be identified.

[0096] Step S304: at the end of the traversal, the overall coverage information of the target voiceprint group to be matched on the voiceprint feature set to be identified is determined based on the obtained coverage of each voiceprint cluster on the voiceprint feature set to be identified.

[0097] It is understandable that when all voiceprint clusters in the target voiceprint group to be matched are traversed in the above manner, the voiceprint cluster coverage of each voiceprint cluster for the voiceprint feature set to be identified can be obtained. For example, the target voiceprint group to be matched contains 5 voiceprint clusters (M1, M2, M3, M4, M5). If the coverage of voiceprint clusters M1, M2, M3, and M5 for the voiceprint feature set [x1, x2, ... xn] meets the standard, and the coverage of M4 for the voiceprint feature set [x1, x2, ... xn] does not meet the standard, then the overall coverage information is that the overall coverage of the target voiceprint group to be matched for the voiceprint feature set to be identified is 80%. If the overall coverage exceeds the above set percentage, it can be determined that the speech to be identified belongs to the target voiceprint group to be matched.

[0098] Furthermore, in order to ensure accurate acquisition of voiceprint cluster coverage, in this embodiment, the above step S303 may specifically include:

[0099] Step S3031: traverse the voiceprint feature set to be identified to obtain the currently traversed voiceprint feature to be matched;

[0100] Step S3032: Obtaining a hit status of the currently traversed voiceprint feature to be matched according to the voiceprint recognition model corresponding to the twin voiceprint pair, wherein the hit status includes a hit or a miss;

[0101] Step S3033: After the traversal of the voiceprint feature set to be identified is completed, the proportion of the voiceprint features to be matched that are hit in the voiceprint feature set to be identified is counted;

[0102] Step S3034: determining the coverage of the voiceprint clusters of the voiceprint feature set to be identified by the currently traversed voiceprint cluster according to the proportion.

[0103] It should be noted that there may be many voiceprint features to be matched [x1, x2, ... xn] in the voiceprint feature set to be identified. To ensure the accuracy of the recognition result, this embodiment preferably adopts a traversal method to match the voiceprint features to be matched in the voiceprint feature set to be identified one by one. In this embodiment, the so-called currently traversed voiceprint feature to be matched is the voiceprint feature currently input into the voiceprint recognition model for scoring.

[0104] It should be understood that for a voiceprint cluster (assuming there is only one twin voiceprint pair), if the voiceprint features [x1, x2, x4, x5] in the voiceprint feature set to be identified [x1, x2, ... x5] are all hit by the voiceprint cluster, then the proportion of the voiceprint features to be matched [x1, x2, x4, x5] hit in the voiceprint feature set to be identified [x1, x2, ... x5] in the voiceprint feature set to be identified is (4 / 5)*100%=80% (≥80%), and it can be determined that the voiceprint cluster coverage of the voiceprint feature set to be identified [x1, x2, ... x5] by the currently traversed voiceprint cluster meets the coverage standard. On the contrary, if the proportion is less than 80%, it is determined that the coverage of the voiceprint cluster does not meet the coverage standard.

[0105] Furthermore, in this embodiment, the specific implementation of step S3032 may include the following steps:

[0106] Step S1: Calculate the model matching score corresponding to the currently traversed voiceprint feature to be matched according to different voiceprint recognition models corresponding to the twin voiceprint pair;

[0107] It should be understood that for the currently traversed voiceprint cluster, its twin voiceprint pair is determined, and accordingly, the voiceprint recognition model corresponding to the speaker in the twin voiceprint pair is also determined. For example, voiceprint cluster c contains a twin voiceprint pair (m and m'), with voiceprint recognition model A1 corresponding to speaker m and voiceprint recognition model A2 corresponding to speaker m'. The currently traversed voiceprint feature to be matched is x1. Then, by inputting the voiceprint feature to be matched x1 into voiceprint recognition models A1 and A2 respectively, the corresponding model matching scores can be calculated as A1(x1) and A2(x1).

[0108] Step S2: Compare the calculated model matching score with the initial threshold value;

[0109] In the specific implementation, after calculating the model matching scores A1(x1) and A2(x1), they can be compared with the initial threshold value (the specific data can be set according to the actual situation), and then judged based on the comparison results whether there is a model matching score A1(x1) and A2(x1) that is smaller than the initial threshold value.

[0110] Of course, in this embodiment, for the voiceprint recognition models corresponding to different twin voiceprint pairs (if there are a large number of them), a traversal method can be adopted to select one model each time to calculate the model matching score corresponding to the currently traversed voiceprint feature to be matched, and then compare the calculated model matching score with the initial threshold value. Once it is less than the initial threshold, the feature matching operation of the current voiceprint cluster is directly abandoned, and it is directly determined that the voiceprint cluster does not hit the currently traversed voiceprint feature to be matched, thereby saving feature matching time and improving the efficiency of voiceprint recognition.

[0111] Step S3: If there is a model matching score less than the initial threshold value among the calculated model matching scores, it is determined that the currently traversed voiceprint cluster does not hit the currently traversed voiceprint feature to be matched;

[0112] It is understandable that, as described above, if there is a model matching score in the calculated model matching score that is less than the initial threshold value, such as A2(x1), then it can be directly determined that the currently traversed voiceprint cluster does not hit the currently traversed voiceprint feature to be matched, and the current voiceprint cluster is abandoned, and the next voiceprint cluster is jumped to continue matching the voiceprint feature, that is, returning to the above step S301 and re-traversing other voiceprint clusters.

[0113] Step S4: if there is no model matching score less than the initial threshold value among the calculated model matching scores, selecting the maximum model matching score from the calculated model matching scores;

[0114] Correspondingly, if there is no model matching score less than the initial threshold value in the calculated model matching scores, that is, the model matching scores A1(x1) and A2(x1) are both greater than the initial threshold value, then the model matching score with the largest score can be selected from the model matching scores A1(x1) and A2(x1), that is, the above-mentioned maximum model matching score, and then the following step S5 is executed.

[0115] Step S5: comparing the maximum model matching score with a preset decision threshold;

[0116] Step S6: If the maximum model matching score is greater than or equal to the preset decision threshold, it is determined that the currently traversed voiceprint cluster hits the currently traversed voiceprint feature to be matched;

[0117] Step S7: If the maximum model matching score is less than the preset decision threshold, it is determined that the currently traversed voiceprint cluster does not hit the currently traversed voiceprint feature to be matched.

[0118] It should be noted that the value of the preset decision threshold in this embodiment is adjustable, and the absolute value of the preset decision threshold is greater than the absolute value of the initial threshold. In actual application, if there is no model matching score less than the initial threshold among the calculated model matching scores, but the largest model matching score is still lower than the preset decision threshold, it is still determined that the currently traversed voiceprint cluster does not match the currently traversed voiceprint feature to be matched.

[0119] In addition, for the voiceprint feature set to be identified [x1, x2, ... x5], if the voiceprint feature x1 to be matched is hit by a voiceprint cluster c in the target voiceprint group to be matched, a traversal method can be used to directly select a voiceprint cluster from the remaining voiceprint clusters to perform feature matching on the next voiceprint feature x2 to be matched, until all voiceprint clusters in the target voiceprint group to be matched are traversed.

[0120] This embodiment traverses all voiceprint clusters in the target voiceprint group to be matched, and then determines the hit status of each voiceprint cluster on the voiceprint features in the voiceprint feature set to be identified based on the pre-built voiceprint recognition model. It can accurately and comprehensively confirm the voiceprint cluster coverage of the voiceprint feature set to be identified. At the same time, this embodiment determines the overall coverage information of the target voiceprint group to be matched on the voiceprint feature set to be identified through the voiceprint cluster coverage, which effectively guarantees the accuracy and reliability of the final recognition result.

[0121] refer to Figure 5 , Figure 5 This is a flow chart of the third embodiment of the voiceprint recognition method based on twin voiceprint pairs of the present invention.

[0122] Based on the above embodiments, in this embodiment, before step S10, the method further includes constructing a voiceprint target group, which specifically includes the following steps:

[0123] Step S01: obtaining a speech sample of each speaker in a speaker group;

[0124] It should be noted that the speaker group in this embodiment may be determined by a specific application scenario, such as a fraud call gang, and the voice samples of each speaker in the group may be obtained during the speaker's call.

[0125] Step S02: Modeling each speaker and training the constructed model based on the voice sample to obtain a voiceprint recognition model;

[0126] It should be noted that in this step, the speaker modeling can be performed by selecting one of the GMM mean supervector model, i-vector model, and x-vector model as the initial voiceprint recognition model, and then the model is trained according to its corresponding speech sample, and the voiceprint recognition model is obtained after the model converges.

[0127] Step S03: Calculating the model matching score corresponding to the voiceprint feature of each speaker in the speaker group one by one using the voiceprint recognition model corresponding to each speaker;

[0128] It should be noted that, for a trained voiceprint recognition model, after inputting any voiceprint feature into it, a model matching score can be obtained that can represent the degree of matching between the voiceprint feature and the voiceprint recognition model.

[0129] For models with higher scores that are more likely to be identified as the target person (voiceprint recognition), the higher the model matching score, the more likely the speaker to whom the voiceprint feature belongs is to be the actual person corresponding to the model. Conversely, for models with lower scores that are more likely to be identified as the target person (voiceprint recognition), the model output can be inverted.

[0130] Step S04: determining the twin voiceprint pairs of the speaker group according to the calculated model matching scores, and constructing a target voiceprint group of the speaker group according to the twin voiceprint pairs.

[0131] It can be understood that the above-mentioned model matching score can reflect the matching degree between the current voiceprint feature and the voiceprint recognition model, that is, it reflects the matching degree between the speaker providing the current voiceprint feature and the actual corresponding person of the voiceprint recognition model. Then, if there is another voiceprint recognition model whose model matching score calculated is the highest except the model matching score calculated by the voiceprint recognition model, then the speaker providing the current voiceprint feature and the speaker corresponding to the voiceprint recognition model with the highest model matching score except the voiceprint recognition model can be called a twin voiceprint pair.

[0132] In order to accurately obtain the twin voiceprint pair, as an implementation method, the above step S04 in this embodiment may specifically include:

[0133] Step S041: constructing a model matching score set corresponding to each voiceprint recognition model according to the calculated model matching scores;

[0134] Step S042: traversing the speaker group and obtaining the target voiceprint recognition model and target voiceprint features corresponding to the currently traversed speaker;

[0135] Step S043: searching the target model matching score set corresponding to the target voiceprint recognition model from the model matching score set;

[0136] Step S044: reading the target model matching score corresponding to the target voiceprint feature from the target model matching score set;

[0137] Step S045: searching the target model matching score set for a maximum model matching score other than the target model matching score;

[0138] Step S046: Determine the target speaker to which the maximum model matching score belongs, and construct a twin voiceprint pair based on the currently traversed speaker and the target speaker;

[0139] Step S047: upon completion of the traversal of the speaker group, obtaining a twin voiceprint pair corresponding to each speaker;

[0140] Step S048: constructing the voiceprint clusters to which each twin voiceprint pair belongs, and constructing the voiceprint target group of the speaker group based on all the voiceprint clusters.

[0141] Here, the above steps S041-S048 are explained with reference to a specific example. For example, the speaker group includes three speakers (A, B, C), and the voiceprint recognition models corresponding to the speakers (A, B, C) are f A (·),f B (·),f C (·), the voiceprint features (sets) provided by speakers (A, B, C) are (X A 、X B 、X C ).

[0142] After the calculations in steps S01-S03 above, the model matching score set corresponding to each voiceprint recognition model can be obtained {[f A (X A ), f A (X B ), f A (X C )],[f B (X A ), f B (X B ), f B (X C )],[f C (X A ), f C (X B ), f C (X C )]}.

[0143] If the speaker traversed at the current moment is A, then the corresponding target voiceprint recognition model can be determined to be f A(·) and the target voiceprint feature is X A At this time, the target voiceprint recognition model f can be found from the above model matching score set A (·) The corresponding target model matching score set [f A (X A ), f A (X B ), f A (X C )], and then read the target voiceprint feature X from the target model matching score set A The corresponding target model matching score f A (X A ); then search the target model matching score set except the target model matching score f A (X A ) outside the maximum model matching score, assuming it is f A (X C ), the maximum model matching score f can be determined A (X C ) is the target speaker C, and then the twin voiceprint pair (A, C) can be constructed based on the currently traversed speaker A and target speaker C.

[0144] In a specific implementation, after constructing all the twin voiceprint pairs corresponding to the speaker group in the above manner, the twin voiceprint pairs can be analyzed, and then voiceprint clusters can be formed based on the analysis results, and then the final voiceprint target group of the speaker group can be constructed based on the voiceprint clusters.

[0145] It should be noted that the above analysis of twin voiceprint pairs can be: if in a twin voiceprint pair (m and m'), speaker m only has a twin voiceprint pair relationship with speaker m' in the twin voiceprint pair, and does not have a twin voiceprint pair relationship with other speakers outside the twin voiceprint pair, then m and m' are said to constitute a voiceprint cluster. In addition, the construction of twin voiceprint pairs in this embodiment is not limited to the specific method described above. Any other method that can determine the similarity of speaker voiceprint features and then construct twin voiceprint pairs based on the similarity can be used, and the present invention does not impose any restrictions on this.

[0146] This embodiment obtains a voiceprint recognition model by obtaining speech samples from each speaker in a speaker group, then building a model for each speaker and training the constructed model based on the speech samples. Within the speaker group, the voiceprint recognition model corresponding to each speaker is used to calculate the model matching score corresponding to each speaker's voiceprint features. Finally, twin voiceprint pairs for the speaker group are determined based on the calculated model matching scores, and a target voiceprint group for the speaker group is constructed based on the twin voiceprint pairs. This embodiment first analyzes the twin voiceprint pairs of the speaker group in this manner, and then determines the voiceprint cluster structure within the target voiceprint group based on the analysis results, ensuring that the ultimately constructed voiceprint group is consistent with the speaker group.

[0147] In addition, an embodiment of the present invention also proposes a storage medium, on which a voiceprint recognition program based on twin voiceprint pairs is stored. When the voiceprint recognition program based on twin voiceprint pairs is executed by a processor, the steps of the voiceprint recognition method based on twin voiceprint pairs as described above are implemented.

[0148] Reference Figure 6 , Figure 6 This is a structural block diagram of the first embodiment of the voiceprint recognition device based on twin voiceprint pairs of the present invention.

[0149] like Figure 6 As shown, the voiceprint recognition device based on twin voiceprint pairs proposed in the embodiment of the present invention includes:

[0150] The feature extraction module 601 is used to extract a voiceprint feature set of the speech to be recognized;

[0151] The voiceprint group acquisition module 602 is used to determine a target voiceprint group to be matched corresponding to the speech to be recognized, wherein the target voiceprint group to be matched includes a plurality of voiceprint clusters;

[0152] The voiceprint pair matching module 603 is configured to perform twin voiceprint pair matching on the voiceprint feature set to be identified based on the voiceprint cluster, and determine the overall coverage information of the voiceprint target group to be matched on the voiceprint feature set to be identified based on the matching results;

[0153] The result determination module 604 is configured to determine that the speech to be recognized belongs to the target group of voiceprints to be matched when the overall coverage information meets a preset condition.

[0154] This embodiment extracts the voiceprint feature set to be recognized of the speech to be recognized, and then determines the target voiceprint group to be matched containing several voiceprint clusters corresponding to the speech to be recognized; then performs twin voiceprint pair matching on the voiceprint feature set to be recognized according to the voiceprint clusters, and finally determines the overall coverage information of the target voiceprint group to be matched with the voiceprint feature set to be recognized based on the matching results, and when the overall coverage information meets the preset conditions, it is determined that the speech to be recognized belongs to the target voiceprint group to be matched. Since the voiceprint recognition of the speaker is performed through the twin voiceprint pairs in the voiceprint cluster contained in the pre-constructed voiceprint group, it is possible to judge whether the speaker belongs to the voiceprint group as a whole, and there is no need to recognize the speech to be recognized one by one according to the speakers in the speaker group, thereby improving the recognition efficiency and having obvious advantages in speech group detection.

[0155] Based on the first embodiment of the voiceprint recognition device based on twin voiceprint pairs of the present invention, a second embodiment of the voiceprint recognition device based on twin voiceprint pairs of the present invention is proposed.

[0156] In this embodiment, the voiceprint pair matching module 603 is further used to traverse the several voiceprint clusters; obtain the twin voiceprint pairs contained in the currently traversed voiceprint cluster, the twin voiceprint pairs contain at least two speakers and each speaker is pre-configured with a voiceprint recognition model; determine the voiceprint cluster coverage of the currently traversed voiceprint cluster on the voiceprint feature set to be identified based on the voiceprint recognition model corresponding to the twin voiceprint pairs; at the end of the traversal, determine the overall coverage information of the voiceprint target group to be matched on the voiceprint feature set to be identified based on the voiceprint cluster coverage of each voiceprint cluster obtained on the voiceprint feature set to be identified.

[0157] Furthermore, the voiceprint pair matching module 603 is also used to traverse the voiceprint feature set to be identified to obtain the currently traversed voiceprint feature to be matched; obtain the hit status of the currently traversed voiceprint feature to be matched according to the voiceprint recognition model corresponding to the twin voiceprint pair, and the hit status includes hit or miss; at the end of the traversal of the voiceprint feature set to be identified, count the proportion of the hit voiceprint features to be matched in the voiceprint feature set to be identified; determine the coverage of the voiceprint cluster of the currently traversed voiceprint cluster to the voiceprint feature set to be identified based on the proportion.

[0158] Furthermore, the voiceprint pair matching module 603 is also used to calculate the model matching score corresponding to the currently traversed voiceprint feature to be matched according to the different voiceprint recognition models corresponding to the twin voiceprint pair; compare the calculated model matching score with the initial threshold value; if there is a model matching score less than the initial threshold value in the calculated model matching scores, it is determined that the currently traversed voiceprint cluster does not hit the currently traversed voiceprint feature to be matched; if there is no model matching score less than the initial threshold value in the calculated model matching scores, the maximum model matching score is selected from the calculated model matching scores; compare the maximum model matching score with the preset decision threshold value; if the maximum model matching score is greater than or equal to the preset decision threshold value, it is determined that the currently traversed voiceprint cluster hits the currently traversed voiceprint feature to be matched; if the maximum model matching score is less than the preset decision threshold value, it is determined that the currently traversed voiceprint cluster does not hit the currently traversed voiceprint feature to be matched.

[0159] Furthermore, the voiceprint pair matching module 603 is further configured to traverse the remaining voiceprint clusters in the plurality of voiceprint clusters, and execute an operation of obtaining the twin voiceprint pairs contained in the currently traversed voiceprint cluster.

[0160] Furthermore, the voiceprint recognition device based on twin voiceprint pairs also includes: a voiceprint group construction module, which is used to obtain voice samples of each speaker in the speaker group; modeling each speaker, and training the constructed model according to the voice samples to obtain a voiceprint recognition model; in the speaker group, using the voiceprint recognition model corresponding to each speaker one by one to calculate the model matching score corresponding to the voiceprint feature of each speaker; determining the twin voiceprint pairs of the speaker group according to the calculated model matching scores, and constructing the voiceprint target group of the speaker group based on the twin voiceprint pairs.

[0161] Furthermore, the voiceprint group construction module is also used to construct a model matching score set corresponding to each voiceprint recognition model based on the calculated model matching score; traverse the speaker group and obtain the target voiceprint recognition model and target voiceprint feature corresponding to the currently traversed speaker; search for the target model matching score set corresponding to the target voiceprint recognition model from the model matching score set; read the target model matching score corresponding to the target voiceprint feature from the target model matching score set; search for the maximum model matching score other than the target model matching score in the target model matching score set; determine the target speaker to which the maximum model matching score belongs, and construct a twin voiceprint pair based on the currently traversed speaker and the target speaker; when the traversal of the speaker group is completed, obtain the twin voiceprint pair corresponding to each speaker; construct the voiceprint cluster to which each twin voiceprint pair belongs, and construct the voiceprint target group of the speaker group based on all voiceprint clusters.

[0162] Other embodiments or specific implementations of the voiceprint recognition device based on twin voiceprint pairs of the present invention can refer to the above-mentioned method embodiments and will not be repeated here.

[0163] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.

[0164] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0165] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0166] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A voiceprint recognition method based on twin voiceprint pairs, characterized in that: The method comprises: Extracting a voiceprint feature set to be recognized of the speech to be recognized; Determine a target group of voiceprints to be matched corresponding to the speech to be recognized. The target group of voiceprints to be matched includes several voiceprint clusters, each of which is composed of twin voiceprint pairs. The twin voiceprint pairs are determined by the voiceprint recognition model corresponding to the speaker. That is, in the group of speakers, the voiceprint features of each speaker are scored one by one using the voiceprint recognition model of each speaker. For each voiceprint recognition model, search the group for the voiceprint features of another speaker m' with the highest voiceprint feature score other than the current speaker m based on the scoring results. The other speaker m' is referred to as the twin voiceprint pair of the current speaker m. Performing twin voiceprint pair matching on the voiceprint feature set to be identified according to the voiceprint cluster, and determining overall coverage information of the voiceprint target group to be matched on the voiceprint feature set to be identified according to the matching results; When the overall coverage information meets a preset condition, it is determined that the speech to be recognized belongs to the voiceprint target group to be matched.

2. The voiceprint recognition method based on twin voiceprint pairs according to claim 1, characterized in that: The step of performing twin voiceprint pair matching on the voiceprint feature set to be identified according to the voiceprint cluster, and determining overall coverage information of the voiceprint target group to be matched on the voiceprint feature set to be identified according to the matching result, includes: Traversing the plurality of voiceprint clusters; Obtaining the twin voiceprint pairs contained in the currently traversed voiceprint cluster, wherein the twin voiceprint pairs contain at least two speakers and each speaker is pre-configured with a voiceprint recognition model; Determining the coverage of the currently traversed voiceprint cluster to the voiceprint feature set to be identified according to the voiceprint recognition model corresponding to the twin voiceprint pair; At the end of the traversal, the overall coverage information of the target voiceprint group to be matched on the voiceprint feature set to be identified is determined based on the obtained voiceprint cluster coverage of each voiceprint cluster on the voiceprint feature set to be identified.

3. The voiceprint recognition method based on twin voiceprint pairs according to claim 2, characterized in that: The voiceprint feature set to be identified includes a plurality of voiceprint features to be matched; The step of determining, based on the voiceprint recognition model corresponding to the twin voiceprint pair, the coverage of the voiceprint cluster of the voiceprint feature set to be identified by the currently traversed voiceprint cluster, includes: Traversing the voiceprint feature set to be identified to obtain the currently traversed voiceprint feature to be matched; Obtaining a hit status of the currently traversed voiceprint feature to be matched according to the voiceprint recognition model corresponding to the twin voiceprint pair, wherein the hit status includes a hit or a miss; At the end of traversal of the voiceprint feature set to be identified, counting the proportion of the voiceprint features to be matched that are hit in the voiceprint feature set to be identified in the voiceprint feature set to be identified; The coverage of the voiceprint clusters of the to-be-identified voiceprint feature set by the currently traversed voiceprint cluster is determined according to the proportion.

4. The voiceprint recognition method based on twin voiceprint pairs according to claim 3, characterized in that: The step of obtaining a hit status of the currently traversed voiceprint feature to be matched according to the voiceprint recognition model corresponding to the twin voiceprint pair includes: Calculate the model matching score corresponding to the currently traversed voiceprint feature to be matched according to different voiceprint recognition models corresponding to the twin voiceprint pair; Compare the calculated model matching score with the initial threshold value; If there is a model matching score less than the initial threshold value among the calculated model matching scores, it is determined that the currently traversed voiceprint cluster does not hit the currently traversed voiceprint feature to be matched; If there is no model matching score less than the initial threshold value among the calculated model matching scores, selecting the maximum model matching score from the calculated model matching scores; Comparing the maximum model matching score with a preset decision threshold; If the maximum model matching score is greater than or equal to the preset decision threshold, it is determined that the currently traversed voiceprint cluster hits the currently traversed voiceprint feature to be matched; If the maximum model matching score is less than the preset decision threshold, it is determined that the currently traversed voiceprint cluster does not hit the currently traversed voiceprint feature to be matched.

5. The voiceprint recognition method based on twin voiceprint pairs according to claim 4, characterized in that: After the step of determining that the currently traversed voiceprint cluster does not hit the currently traversed voiceprint feature to be matched, the method further includes: The remaining voiceprint clusters in the plurality of voiceprint clusters are traversed, and the step of obtaining the twin voiceprint pairs contained in the currently traversed voiceprint cluster is returned.

6. The voiceprint recognition method based on twin voiceprint pairs according to any one of claims 1 to 5, characterized in that: Before the step of extracting the voiceprint feature set to be recognized of the speech to be recognized, the method further includes: Obtaining a speech sample of each speaker in the speaker group; Modeling each speaker and training the constructed model based on the voice sample to obtain a voiceprint recognition model; In the speaker group, using the voiceprint recognition model corresponding to each speaker one by one to calculate the model matching score corresponding to the voiceprint feature of each speaker; The twin voiceprint pairs of the speaker group are determined according to the calculated model matching scores, and a voiceprint target group of the speaker group is constructed according to the twin voiceprint pairs.

7. The voiceprint recognition method based on twin voiceprint pairs according to claim 6, characterized in that: The step of determining the twin voiceprint pairs of the speaker group according to the calculated model matching scores, and constructing the voiceprint target group of the speaker group according to the twin voiceprint pairs includes: Construct a model matching score set corresponding to each voiceprint recognition model based on the calculated model matching score; Traversing the speaker group and obtaining the target voiceprint recognition model and target voiceprint features corresponding to the currently traversed speaker; Searching for a target model matching score set corresponding to the target voiceprint recognition model from the model matching score set; Reading the target model matching score corresponding to the target voiceprint feature from the target model matching score set; Searching for a maximum model matching score other than the target model matching score in the target model matching score set; Determine the target speaker to which the maximum model matching score belongs, and construct a twin voiceprint pair based on the currently traversed speaker and the target speaker; When the traversal of the speaker group is completed, obtaining a twin voiceprint pair corresponding to each speaker; Construct a voiceprint cluster to which each twin voiceprint pair belongs, and construct a voiceprint target group of the speaker group based on all the voiceprint clusters.

8. A voiceprint recognition device based on twin voiceprint pairs, characterized in that: The device comprises: A feature extraction module is used to extract a voiceprint feature set to be recognized from the speech to be recognized; A voiceprint group acquisition module is used to determine a target voiceprint group to be matched corresponding to the speech to be recognized. The target voiceprint group to be matched contains several voiceprint clusters, each of which is composed of twin voiceprint pairs. The twin voiceprint pairs are determined by the voiceprint recognition model corresponding to the speaker. That is, in the group of speakers, the voiceprint features of each speaker are scored one by one using the voiceprint recognition model of each speaker. For each voiceprint recognition model, the voiceprint features of other speakers m' with the highest voiceprint features other than the current speaker m are searched in the group based on the scoring results. The other speakers m' are called the twin voiceprint pairs of the current speaker m. a voiceprint pair matching module, configured to perform twin voiceprint pair matching on the voiceprint feature set to be identified based on the voiceprint cluster, and determine overall coverage information of the voiceprint target group to be matched on the voiceprint feature set to be identified based on the matching results; The result determination module is configured to determine that the speech to be recognized belongs to the voiceprint target group to be matched when the overall coverage information meets a preset condition.

9. A voiceprint recognition device based on twin voiceprint pairs, characterized in that: The device includes: a memory, a processor, and a voiceprint recognition program based on twin voiceprint pairs stored in the memory and runnable on the processor, wherein the voiceprint recognition program based on twin voiceprint pairs is configured to implement the steps of the voiceprint recognition method based on twin voiceprint pairs as described in any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores a voiceprint recognition program based on twin voiceprint pairs. When the voiceprint recognition program based on twin voiceprint pairs is executed by the processor, the steps of the voiceprint recognition method based on twin voiceprint pairs as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Voiceprint recognition method and device, terminal device and storage medium

    CN108900725A

  • Audio data annotation method and device, electronic equipment and storage medium

    CN109637547A