A voice authentication method, terminal device, and storage medium

By collecting pre-trained voiceprint models of multiple channels, combining Tippett diagrams and PLDA online training, selecting appropriate voiceprint models solves the problem of insufficient generalization ability of voiceprint models, achieving high-accuracy voice identification, reducing costs and reducing dependence on professionals.

CN116052688BActive Publication Date: 2025-07-25XIAMEN KUAISHANGTONG TECH CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111259396.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-28
Publication Date
2025-07-25
Estimated Expiration
2041-10-28

AI Technical Summary

Technical Problem

The existing voiceprint models lack generalization capabilities in judicial identification, resulting in insufficient recognition accuracy, and the amount of data required for retraining the model, making it difficult to achieve automated recognition.

Method used

Pre-trained vocalprint models of multiple channels are collected, model performance is evaluated through Tippett graphs and calibration sets, combined with PLDA online training, and vocalprint models that meet the requirements are selected as the identification model, and score calibration and scene adaptation training are used to achieve speech identification.

Benefits of technology

It realizes automated and low-cost voice identification, eliminates interference from personal factors, improves recognition accuracy, and reduces dependence on professionals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116052688B_ABST
    Figure CN116052688B_ABST
Patent Text Reader

Abstract

The present invention relates to a voice authentication method, a terminal device, and a storage medium. The method includes: S1: Collect multiple pre-trained voiceprint models corresponding to multiple channels; S2: Collect voice data of the same channel as the voice to be detected, and divide it into a training set, a test set, and a calibration set; S3: Determine the model performance of each voiceprint model in S1 in turn, use the voiceprint model that meets the requirements as the authentication model, and enter S5. When all the voiceprint models in S1 do not meet the requirements, enter S4; S4: Extract the voiceprint models with all cllr values less than the threshold, perform PLDA online training on the extracted voiceprint models through the training set, and determine the model performance of the models after PLDA online training until a voiceprint model that meets the requirements is found as the authentication model and then enter S5; S5: Use the authentication model to determine whether the voice to be detected and the sample belong to the same person. The present invention adopts score calibration and scenario adaptation training, ensuring the accuracy rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition, and in particular to a speech discrimination method, a terminal device, and a storage medium. Background Art

[0002] With the development of technology and the popularization of intelligent devices, the means in the field of forensic appraisal are constantly enriched. From the initial fingerprint, it has developed to DNA, and then to the recent image recognition. Automatic recognition technology is gradually replacing manual recognition. However, at present, voiceprint in China still mainly relies on manual recognition, and it is identified whether it belongs to the same person by simulating and observing the formants of fixed letter pronunciations. The main reason for this method is that the generalization ability of the current voiceprint model is insufficient, and the pre-trained model is not sufficient to support the accuracy required for forensic appraisal. If a model is re-trained for each case, the amount of data required is huge, and the feasibility is almost zero. Summary of the Invention

[0003] In order to solve the above problems, the present invention proposes a speech discrimination method, a terminal device, and a storage medium.

[0004] The specific solutions are as follows:

[0005] A speech discrimination method includes the following steps:

[0006] S1: Collect multiple pre-trained voiceprint models corresponding to multiple channels;

[0007] S2: Collect speech data of the same channel as the speech to be detected, and divide it into a training set, a test set, and a calibration set;

[0008] S3: Determine the model performance of each voiceprint model in S1 in turn, use the voiceprint model that meets the requirements as the identification model, and enter S5. When all voiceprint models in S1 do not meet the requirements, enter S4;

[0009] Step S3 specifically includes the following steps:

[0010] S301: Select a voiceprint model in S1;

[0011] S302: After putting the test set into the voiceprint model selected in S3, evaluate the performance effect of the voiceprint model on the test set, and determine whether the voiceprint model is available according to whether the performance index meets the requirements; the requirements of the performance index include that the Tippett plot distribution meets the requirements and the calibration log-likelihood ratio cllr < 0.2; when the Tippett plot distribution does not meet the requirements, return to S301 to re-select the voiceprint model. After all voiceprint models are selected, enter S4; when the Tippett plot distribution meets the requirements, enter S303;

[0012] S303: Determine whether cllr < 0.2 is satisfied. If yes, use this voiceprint model as the authentication model and proceed to S5; otherwise, perform score calibration on this voiceprint model using the calibration set.

[0013] S304: Determine whether the calibrated voiceprint model satisfies cllr < 0.2. If yes, use the calibrated voiceprint model as the authentication model and proceed to S5; otherwise, return to S3 to reselect the model until all voiceprint models have been selected, and then proceed to S4.

[0014] S4: Extract all voiceprint models with cllr values less than the threshold, perform PLDA online training on each of the extracted voiceprint models using the training set, and determine the model performance of the models after PLDA online training until a voiceprint model that meets the requirements is found as the authentication model, and then proceed to S5.

[0015] S5: Use the authentication model to determine whether the voice to be detected and the sample belong to the same person.

[0016] Further, all the multiple voiceprint models collected in step S1 are PLDA classifiers.

[0017] Further, step S2 further includes labeling the speakers in the collected voice data.

[0018] Further, the method for using the authentication model in step S5 to determine whether the voice to be detected and the sample belong to the same person is: after inputting the voice to be detected and the sample into the authentication model, obtain the log-likelihood ratio of the authentication model, and determine whether the voice to be detected and the sample belong to the same person according to the position of the log-likelihood ratio in the Tippett plot.

[0019] Further, the method for determining whether the voice to be detected and the sample belong to the same person according to the position of the log-likelihood ratio in the Tippett plot is: obtain the probabilities of the same person and different people corresponding to the log-likelihood ratio according to the Tippett plot. If the probability of the same person is greater than the probability of different people, it is determined that they belong to the same person; otherwise, it is determined that they belong to different people.

[0020] A voice authentication terminal device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method described above in the embodiments of the present invention are implemented.

[0021] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described above in the embodiments of the present invention are implemented.

[0022] The present invention adopts the above technical solutions and has the beneficial effects:

[0023] (1) Convert manual recognition to machine recognition to eliminate interference from personal factors;

[0024] (2) The usage cost is relatively low, and there is no need for professionals with many years of experience;

[0025] (3) Fractional calibration and scene adaptation training are adopted to ensure the accuracy. Brief Description of the Drawings

[0026] Figure 1 The flowchart of the first embodiment of the present invention is shown.

[0027] Figure 2 The Tippett diagram corresponding to the identification model in this embodiment is shown. Detailed Embodiments

[0028] To further illustrate each embodiment, the present invention provides drawings. These drawings are part of the disclosure of the present invention, mainly used to illustrate the embodiments, and can be combined with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention.

[0029] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0030] First Embodiment:

[0031] The first embodiment of the present invention provides a voice authentication method, as Figure 1 shown, the method includes the following steps:

[0032] S1: Collect multiple pre-trained voiceprint models corresponding to multiple channels.

[0033] In this embodiment, for the online training of the PLDA classifier, the pre-trained voiceprint models need to include the PLDA classifier. The algorithms adopted can be current mainstream algorithms such as TDNN, ECAPA-TDNN, etc. The multiple channels included are existing mainstream channels, such as WeChat voice, phone calls, etc. Each channel corresponds to a voiceprint model.

[0034] S2: Collect voice data of the same channel as the voice to be detected, and divide it into a training set, a test set, and a calibration set.

[0035] Step S2 is used to collect voice data from the same source as the voice to be detected to train the model, so that the output result of the model is more accurate. For example, if the voice to be detected is WeChat voice, the collected voice data is WeChat voice data; if the voice to be detected is a dialect, the collected voice data is also the corresponding dialect.

[0036] The collected voice data should also be labeled with the speaker to facilitate the positive and negative sample pairs of candidate components.

[0037] In this embodiment, the number of voices included in the training set is not less than 200 person - shares, the number of voices included in the test set and the calibration set is not less than 100 person - shares, and each speaker includes at least two voices.

[0038] S3: Determine the model performance of each voiceprint model in S1 in turn. Take the voiceprint model that meets the requirements as the identification model and enter S5. When all the voiceprint models in S1 do not meet the requirements, enter S4.

[0039] The implementation of step S3 in this embodiment specifically includes the following steps:

[0040] S301: Select a voiceprint model in S1;

[0041] S302: After putting the test set into the voiceprint model selected in S3, evaluate the performance effect of this voiceprint model on this test set, and determine whether this voiceprint model is available according to whether the performance index meets the requirements. The requirements of the performance index include that the Tippett plot distribution meets the requirements and Cllr (Check log - likelihood ratio) < 0.2. When the Tippett plot distribution does not meet the requirements, return to S301 to re - select the voiceprint model until all voiceprint models have been selected, and then enter S4. When the Tippett plot distribution meets the requirements, enter S303;

[0042] S303: Judge whether cllr < 0.2 is satisfied. If so, take this voiceprint model as the identification model and enter S5; otherwise, calibrate the scores of this voiceprint model through the calibration set;

[0043] S304: Judge whether the calibrated voiceprint model satisfies cllr < 0.2. If so, take the calibrated voiceprint model as the identification model and enter S5; otherwise, return to S3 to re - select the model until all voiceprint models have been selected, and then enter S4;

[0044] As Figure 2 shown is the Tippett plot corresponding to the identification model. In the figure, the two curves respectively represent the log - likelihood ratio distribution of different people and the log - likelihood ratio distribution of the same people. The abscissa represents the value of the log - likelihood ratio, and the ordinate represents the confidence level, that is, the probability.

[0045] S4: Extract all voiceprint models with cllr values less than the threshold, perform PLDA online training on each of the extracted voiceprint models through the training set, and determine the model performance of the model after PLDA online training until a voiceprint model that meets the requirements is found as the identification model and then enter S5.

[0046] Those skilled in the art can set the size of the threshold value by themselves, and it should be a value slightly greater than 0.2.

[0047] S5: Use the identification model to determine whether the voice to be detected and the sample belong to the same person.

[0048] The sample is the voice of a known speaker. In this embodiment, after inputting the voice to be detected and the sample into the identification model, the log-likelihood ratio of the identification model is obtained, and it is determined whether the voice to be detected and the sample belong to the same person according to the position of the log-likelihood ratio in the Tippett diagram. Specifically, according to the probabilities of the same person and different people corresponding to the log-likelihood ratio obtained from the Tippett diagram, if the probability of the same person is greater than the probability of different people, it is determined to belong to the same person; otherwise, it is determined to belong to different people. For example, if the value of the log-likelihood ratio is 25, the ordinate value corresponding to the curve of the log-likelihood ratio distribution of the same person in the Tippett diagram is 0.9, and the ordinate value corresponding to the curve of the log-likelihood ratio distribution of different people is 0, then the probability of belonging to the same person is 90%, the probability of belonging to different people is 0%, and since the probability of the same person is greater than the probability of different people, it is determined to belong to the same person; if the value of the log-likelihood ratio is -10, the ordinate value corresponding to the curve of the log-likelihood ratio distribution of the same person in the Tippett diagram is 0.05, and the ordinate value corresponding to the curve of the log-likelihood ratio distribution of different people is 0.41, then the probability of belonging to the same person is 5%, the probability of belonging to different people is 41%, and since the probability of the same person is less than the probability of different people, it is determined not to belong to the same person.

[0049] Embodiment 2:

[0050] The present invention also provides a voice identification terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above method embodiment of Embodiment 1 of the present invention are implemented.

[0051] Furthermore, as an executable solution, the voice identification terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The voice identification terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above composition structure of the voice identification terminal device is only an example of the voice identification terminal device, and does not constitute a limitation on the voice identification terminal device. It may include more or fewer components than the above, or combine some components, or different components. For example, the voice identification terminal device may also include input / output devices, network access devices, a bus, etc. The embodiments of the present invention do not make limitations on this.

[0052] Further, as an executable solution, the so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the voice authentication terminal device, and connects various parts of the entire voice authentication terminal device through various interfaces and lines.

[0053] The memory can be used to store the computer program and / or module. By running or executing the computer program and / or module stored in the memory, and by calling the data stored in the memory, the processor implements various functions of the voice authentication terminal device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.

[0054] The present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method in the embodiments of the present invention are implemented.

[0055] If the modules / units integrated in the voice authentication terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution medium, etc.

[0056] Although the present invention has been specifically shown and described with reference to the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in terms of form and details without departing from the spirit and scope of the present invention defined by the appended claims. All such changes are within the scope of protection of the present invention.

Claims

1. A voice authentication method, characterized in that, It includes the following steps: S1: Collect multiple pre-trained voiceprint models corresponding to multiple channels; S2: Collect voice data of the same channel as the voice to be detected, and divide it into a training set, a test set, and a calibration set; S3: Determine the model performance of each voiceprint model in S1 in turn. Use the voiceprint model that meets the requirements as the identification model and enter S5. When all voiceprint models in S1 do not meet the requirements, enter S4; Step S3 specifically includes the following steps: S301: Select a voiceprint model in S1; S302: After putting the test set into the voiceprint model selected in S3, evaluate the performance effect of this voiceprint model on this test set, and determine whether this voiceprint model is available according to whether the performance index meets the requirements. The requirements for the performance index include that the Tippett plot distribution meets the requirements and the calibration log-likelihood ratio cllr < 0.2; when the Tippett plot distribution does not meet the requirements, return to S301 to re-select the voiceprint model until all voiceprint models have been selected, and then enter S4; when the Tippett plot distribution meets the requirements, enter S303; S303: Judge whether cllr < 0.2 is satisfied. If so, use this voiceprint model as the identification model and enter S5; otherwise, perform score calibration on this voiceprint model through the calibration set; S304: Judge whether the calibrated voiceprint model meets cllr < 0.

2. If so, use the calibrated voiceprint model as the identification model and enter S5; otherwise, return to S3 to re-select the model until all voiceprint models have been selected, and then enter S4; S4: Extract all voiceprint models with cllr values less than the threshold, perform PLDA online training on each extracted voiceprint model through the training set, and determine the model performance of the model after PLDA online training until a voiceprint model that meets the requirements is found as the identification model and then enter S5; S5: Use the identification model to determine whether the voice to be detected and the sample belong to the same person.

2. The voice authentication method according to claim 1, wherein: All the multiple voiceprint models collected in step S1 are PLDA classifiers.

3. The voice authentication method according to claim 1, wherein: Step S2 further includes labeling the speakers in the collected voice data.

4. The voice authentication method according to claim 1, wherein: The method for using the identification model in step S5 to determine whether the voice to be detected and the sample belong to the same person is: After inputting the voice to be detected and the sample into the identification model, obtain the log-likelihood ratio of the identification model, and determine whether the voice to be detected and the sample belong to the same person according to the position of the log-likelihood ratio in the Tippett plot.

5. The voice authentication method according to claim 4, wherein: The method for determining whether the voice to be detected and the sample belong to the same person according to the position of the log-likelihood ratio in the Tippett plot is: Obtain the probability of the same person and the probability of different people corresponding to the log-likelihood ratio according to the Tippett plot. If the probability of the same person is greater than the probability of different people, it is determined to belong to the same person; otherwise, it is determined to belong to different people.

6. A voice authentication terminal device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the voice authentication method according to any one of claims 1 to 5.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the steps of the voice authentication method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and apparatus with registration for speaker recognition

    US20210125617A1

  • Speaker verification system

    WO1987000332A1