An evaluation method and system for the impact of speech rate on voiceprint security

Through phoneme positioning and speech time domain scaling technology, data sets of different speech speeds are constructed, the impact of speech speed on the vocal print error recognition rate is measured, and a speech speed safety evaluation algorithm is designed, which solves the problem of the reduction in accuracy of the existing voiceprint recognition system under environmental interference, and improves the security and availability of the voiceprint recognition system.

CN114038472BActive Publication Date: 2025-05-30ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111283004.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-01
Publication Date
2025-05-30
Estimated Expiration
2041-11-01

AI Technical Summary

Technical Problem

The existing voiceprint recognition system has reduced accuracy under environmental interference and lacks effective safety measurement and improvement methods, which makes it difficult to ensure the safety of voiceprints.

Method used

Through phoneme positioning and speech time domain scaling technology, a data set of different registered speech speeds and test speech speeds are constructed, the impact of speech speed on the error recognition rate of voiceprints is measured, and a speech speed safety evaluation algorithm is designed to quantify the speech speed safety of voiceprints.

Benefits of technology

Improve the security and usability of the voiceprint recognition system and provide targeted suggestions to improve the performance and security of the voiceprint recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114038472B_ABST
    Figure CN114038472B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for evaluating the impact of speech rate on voiceprint security. The method includes measuring the speech rate of speech files in an open corpus through phoneme localization and segmentation; expanding the corpus through speech time-domain stretching technology to obtain speech data of multiple speakers at different speech rates; constructing datasets with different enrollment speech rates and test speech rates to explore the impact of speech at different speech rates on the voiceprint misrecognition rate during the voiceprint recognition process; on this basis, a voiceprint speech rate security scoring algorithm is designed to quantitatively measure the speech rate security of voiceprints, which helps manufacturers and users improve the security and usability of voiceprint recognition systems, and has significant research significance and social value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of voiceprint security, and relates to a method and system for evaluating the impact of speech rate on voiceprint security. Background Art

[0002] A voiceprint is a biometric feature used to distinguish one person's voice from others, including physiological features (such as vocal tract shape) and behavioral features (such as pronunciation). Compared with other biometric recognition technologies, voiceprints have the advantages of being easy to use, having low requirements for hardware, and having a long operable distance, and can be widely used for device wake-up. Usually, the voiceprint recognition word is the device wake-up word. Nowadays, most smartphones, intelligent voice systems, and computer voice assistants are equipped with voiceprint recognition systems, and voiceprint authentication has also become one of the most commonly used biometric authentication methods in intelligent devices, public security, and financial transactions. Currently, there is a lack of measurement and research on voiceprint security.

[0003] In practical applications, the accuracy of commercial voiceprint recognition systems is relatively ordinary. Especially in the presence of environmental interference, the verification accuracy will be greatly reduced. With the increasing demand for voice-based access control, ensuring the security of voiceprint recognition is a severe challenge for current voiceprint recognition systems, and it has become particularly urgent to improve the performance and security of voiceprint recognition systems. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a method and system for evaluating the impact of speech rate on voiceprint security. This method defines the ratio of the number of phonemes to the net speech duration as the speech rate through phoneme localization, and measures the speech rate and its distribution of speech files in a public corpus; expands the corpus through speech time-domain stretching technology to obtain speech data of multiple speakers at different speech rates; constructs datasets with different registration speech rates and test speech rates, and tests the impact of speech at different speech rates on the voiceprint misrecognition rate during the voiceprint recognition process; on this basis, designs a voiceprint speech rate security evaluation algorithm to quantitatively measure the speech rate security of voiceprints, which helps manufacturers and users improve the security and usability of voiceprint recognition systems, and has significant research significance and social value.

[0005] The technical solution adopted by the present invention is as follows:

[0006] The first objective is to provide a method for evaluating the impact of speech rate on voiceprint security, including the following steps:

[0007] Step 1: Obtain speech data of different speakers at different speech rates to form a speaker corpus;

[0008] Step 2: Calculate the speech rate corresponding to each piece of speech data in the corpus, statistically analyze the speech rate distribution, and record the speech rate with the most corresponding speech data as the original speech rate;

[0009] Step 3: Construct different types of registered speech rate - test speech rate data sets;

[0010] Step 4: Conduct speaker recognition tests using the data sets obtained in Step 3: Use the speech data corresponding to the registered speech rate in the three data sets to register the speaker recognition model to obtain the speaker's voiceprint; then use the speech data corresponding to different test speech rates to test the speaker recognition model to obtain test results, and count the false acceptance rate and false rejection rate, respectively, to obtain the relationship curves between the test speech rate and the false acceptance rate and false rejection rate, and the registered speech rate - test speech rate - false acceptance rate matrix M FAR and the registered speech rate - test speech rate - false rejection rate matrix M FRR ;

[0011] Step 5: According to the registered speech rate - test speech rate false acceptance rate matrix M FAR and the registered speech rate - test speech rate false rejection rate matrix M FRR , calculate the weighted relationship matrix M = λM FAR + μM FRR , where λ and μ are the weighted coefficients of M FAR and M FRR respectively;

[0012] Step 6: Fit the diagonal data of the weighted relationship matrix M to obtain the function F 1 , and take the registered speech rate corresponding to the minimum value of the function F 1 as the reference registered speech rate; fit the different registered speech rates in the weighted relationship matrix M respectively to obtain the function F 2 ;

[0013] Step 7: Calculate the total misrecognition rate difference according to the user's registered speech rate, user's test speech rate, function F 1 and function F 2 , and perform normalization processing on the total misrecognition rate difference to obtain the speech rate security score in the range of [0, 1].

[0014] Furthermore, in Step 3, three types of registered speech rate - test speech rate data sets are constructed:

[0015] The first data set: Select the speech data with the speech rate equal to the original speech rate described in Step 2 and the corresponding speakers from the corpus, use the original speech rate as the registered speech rate, divide the speech data of the original speech rate into two parts, which are used as the registration set and the test set respectively, and use the speech time domain stretching method to obtain w different speech rates for the test set as the test speech rates;

[0016] The second data set is to select speech data and corresponding speakers whose speaking speed is equal to the original speaking speed in step 2 from the corpus, use the original speaking speed as the registration speaking speed, and continue to select w-file speaking speeds corresponding to each speaker from the corpus as the test speaking speed;

[0017] The third data set, screens out speech data and corresponding speakers whose speaking speed is equal to the original speaking speed described in step 2 from the corpus, divides the speech data with the original speaking speed into two parts, respectively as a registration set and a test set, and uses the speech time domain stretching method on the registration set and the test set to obtain w different registration speaking speeds and w different test speaking speeds.

[0018] Furthermore, the speech rate corresponding to each voice data in the corpus is calculated, specifically:

[0019] Step 2.1: Convert the English letters in each speech corresponding text into phonemes and count the number of phonemes q;

[0020] Step 2.2: Use speech endpoint detection technology to measure the net speech duration t;

[0021] Step 2.3: Use the ratio of the number of phonemes to the net speech duration Represents the speaking rate in phonemes per second.

[0022] Furthermore, the speech data corresponding to the test speech speed in the first data set is obtained by: using the speech time domain stretching algorithm to stretch the speech data at the original speech speed with a time domain stretching degree of α, the speech duration is stretched to α times the original, and the corresponding net speech duration becomes αt. Since the number of text phonemes does not change, the speech speed changes from s to When 0<α<1, the speaking speed becomes faster; when α>1, the speaking speed becomes slower; when α=1, the speaking speed remains unchanged.

[0023] Furthermore, the calculation formula of the total misrecognition rate difference is:

[0024] D s =-(|δ enroll |+|δ test |)

[0025] δ enroll =F 1 (s enroll 1)-F 1 (Th enroll )

[0026] δ test =F 2 (s test 1)-F 2 (s enroll 1)

[0027] Among them, D s is the total difference in misrecognition rate, s test 1 is the user's test speech rate, s enroll 1 is the user's registered speech rate, Th enroll is the reference registered speech rate, |.| represents the absolute value operation, F 1 (.) represents the registration speech rate - misrecognition rate fitting function when the registration speech rate in the weighted relationship matrix M is equal to the test speech rate, F 2 (.) represents the test speech rate - misrecognition rate fitting function at a fixed registration speech rate in the weighted relationship matrix M; δ test represents the difference between the misrecognition rate corresponding to the user's test speech rate and the misrecognition rate corresponding to the user's registered speech rate, δ enroll represents the difference between the misrecognition rate corresponding to the user's registered speech rate and the misrecognition rate corresponding to the reference registered speech rate.

[0028] The second objective is to provide a system for evaluating the impact of speech rate on voiceprint security, which is used to implement the above evaluation method.

[0029] The beneficial effects of the present invention are:

[0030] The present invention quantifies the speech rate by using the number of phonemes per unit time, provides more speech materials with different speech rates by changing the speech rate of the speaker through speech time-domain stretching technology, designs various types of registration speech rate - test speech rate data sets to obtain the relationship between the voiceprint misrecognition rate and the speech rate, and on this basis, designs a speech rate security scoring algorithm to quantitatively analyze the speech rate security of users using the voiceprint recognition system, which can specifically provide a reference basis for manufacturers and users to improve voiceprint security and help manufacturers improve the security and usability of the voiceprint recognition system. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a schematic flow chart of an evaluation method for the impact of speech rate on voiceprint security shown in an embodiment of the present invention;

[0032] Figure 2 is a schematic flow chart of a speech rate security scoring algorithm shown in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The technical solutions of the present invention will be further described below with specific examples. And the concept of the example embodiment will be comprehensively conveyed to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0034] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all steps. For example, some steps can be further decomposed, while some steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.

[0035] The present invention studies the correlation between voiceprint security and speech rate, including measuring the speech rate of speech files in an open corpus through phoneme localization and segmentation, obtaining speech data of multiple speakers at different speech rates by expanding the corpus through speech time-domain stretching technology, constructing data sets of different registration speech rates and test speech rates, etc. The influence of speech at different speech rates on the voiceprint misrecognition rate is obtained during the voiceprint recognition process. On this basis, a voiceprint speech rate security scoring algorithm is designed, which can quickly quantify and measure the speech rate security of voiceprints.

[0036] As Figure 1 shown, the evaluation method for the influence of speech rate on voiceprint security proposed by the present invention includes the following steps:

[0037] Specifically as follows:

[0038] Step 1: Measure the speech rate of speech files in the corpus through phoneme localization and segmentation. The corpus contains speech data of different speakers. The specific steps are as follows:

[0039] Step 1.1: Convert the English letters in the text corresponding to each piece of speech into phonemes, and count the number of phonemes q;

[0040] Step 1.2: Use voice activity detection technology to measure the net speech duration t;

[0041] Step 1.3: Use the ratio of the number of phonemes to the net speech duration to represent the speech rate, with the unit of phonemes per second.

[0042] Step 2: Obtain the speech rate distribution corresponding to the speech data in the corpus, and record the speech rate with the most corresponding speech data as the original speech rate; through speech time-domain stretching technology, expand the speech data corresponding to the original speech rate to obtain speech data of multiple speakers at different speech rates. The specific steps are as follows:

[0043] Step 2.1: Select a speech time-domain stretching algorithm. In this embodiment, the WSOLA algorithm is selected.

[0044] Step 2.2: Use the algorithm in Step 2.1 to perform time-domain stretching on the speech corresponding to the original speech rate with a time-domain stretching factor of α. The speech duration is stretched to α times the original, and the corresponding net speech duration t also becomes αt. Since the number of text phonemes remains unchanged, the speech rate changes from s′ to When 0 < α < 1, the speech rate becomes faster; when α > 1, the speech rate becomes slower; when α = 1, the speech rate remains unchanged.

[0045] Step 3: Construct datasets with different registration speech rates and test speech rates, obtain the voiceprint misrecognition rates at different speech rates, and establish a speech rate - voiceprint misrecognition rate relationship matrix. The specific steps are as follows:

[0046] Step 3.1: In this embodiment, 3 experiments are designed. In Experiment 1, the registration speech rate is the original speech rate, and the test speech rate is divided into w levels to explore the influence of the test speech rate on the voiceprint misrecognition rate. In Experiment 2, the speech rates are set the same as in Experiment 1, but the used voices are all voices that have not been processed by the voice time-domain stretching technology, to exclude the influence of the voice time-domain stretching technology on the results of Experiment 1. In Experiment 3, both the registration speech rate and the test speech rate are changed simultaneously to comprehensively explore the influence of the changes in the registration speech rate and the test speech rate on the voiceprint misrecognition rate.

[0047] Step 3.2: According to the 3 experiments designed in Step 3.1, construct 3 corresponding data sets.

[0048] The first data set: Select the voice data with the speech rate equal to the original speech rate described in Step 2 and the corresponding speakers from the corpus. Use the original speech rate as the registration speech rate, divide the voice data at the original speech rate into two parts, which are used as the registration set and the test set respectively. Use the voice time-domain stretching method on the test set to obtain w levels of different speech rates as the test speech rate. Specifically, select m pieces of voice at the original speech rate for each speaker, corresponding to n speakers. The test set speech rate is divided into w levels of different speech rates, which are obtained by time-domain stretching each piece of voice at the original speech rate. Each level of speech rate has m pieces of voice, for a total of w + 1 data sets, simply referred to as data set 1.

[0049] The second data set: Select the voice data with the speech rate equal to the original speech rate described in Step 2 and the corresponding speakers from the corpus. Use the original speech rate as the registration speech rate, and continue to select w levels of speech rates corresponding to each speaker from the corpus as the test speech rate. Specifically, select m pieces of voice at the original speech rate for each speaker, corresponding to n speakers. The test set speech rate is divided into w levels of different speech rates, and m pieces of voice are selected for each speaker at each level of speech rate, for a total of w + 1 data sets. However, all voices are original voices that have not been processed by the voice time-domain stretching technology, simply referred to as data set 2.

[0050] The third data set: Select the voice data with the speech rate equal to the original speech rate described in Step 2 and the corresponding speakers from the corpus. Divide the voice data at the original speech rate into two parts, which are used as the registration set and the test set respectively. Use the voice time-domain stretching method on the registration set and the test set respectively to obtain w levels of different registration speech rates and w levels of different test speech rates. Specifically, the registration speech rate is divided into w levels of different speech rates, and m pieces of voice are selected for each speaker at each level of speech rate. The test speech rate is divided into w levels of different speech rates, and m pieces of voice are selected for each speaker at each level of speech rate, for a total of 2w data sets, simply referred to as data set 3.

[0051] In a specific implementation of the present invention, taking Dataset 1 as an example, here n = 50, m = 10, w = 5 are selected, that is, 50 speakers. Experiments are carried out on the train-clean-360 of the public corpus LibriSpeech. The statistical results show that 12 phonemes per second is the speech rate with the largest number of sentences, and this speech rate is selected as the original speech rate. Dataset 1 selects 10 original-speed voices of each speaker as the enrollment set. The test speech rates are divided into five gears: 8, 10, 12, 14, and 16 phonemes per second. 10 voices are selected for each gear, resulting in a total of 6 datasets.

[0052] Step 3.3: For a certain speaker recognition model, use the 3 datasets constructed in Step 3.2 to conduct speaker recognition tests respectively, obtain the speaker recognition results, and statistically analyze the false recognition rate by dividing it into the false acceptance rate (FAR) and the false rejection rate (FRR). They are the false recognition rates obtained by comparing the speaker's voiceprint with the voiceprints of others and by comparing the speaker's voiceprint with their own voiceprint respectively. Here, the false recognition rate is calculated for the test set of each speech rate.

[0053] Step 3.4: Statistically analyze the results of the 3 experiments to obtain the relationship curves of FAR and FRR with the test speech rate, and the relationship matrix M of FAR and FRR with the enrollment speech rate and the test speech rate FAR and M FRR 。

[0054] Step 4: Design a voiceprint speech rate security scoring algorithm to measure the speech rate security of speaker recognition. The specific steps are as follows:

[0055] Step 4.1: In the voiceprint enrollment stage, measure the speech rate of the user's enrollment audio to obtain the user's daily speech rate s enroll 1; in the test stage, measure the speech rate s test 1 of the test audio. Here, it is assumed that s enroll 1 = 10 phonemes per second, s test 1 = 12 phonemes per second.

[0056] Step 4.2: Calculate the weighted relationship matrix M = λM FAR + μM FRR , where λ and μ are the weighted coefficients of M FAR and M FRR respectively (0 ≤ λ ≤ 1, 0 ≤ μ ≤ 1, λ + μ = 1). Fit the diagonal data M(s enroll = s test ) of the weighted relationship matrix M to obtain F 1 , and calculate the enrollment speech rate corresponding to the minimum value of F 1 as Th enroll , to determine the benchmark enrollment speech rate; then in M(s enroll = senroll 1) Calculate the fitting function F 2 ; Here, take λ = μ = 0.5 to obtain Th enroll = 12 phonemes per second.

[0057] Step 4.3: For s enroll 1 and s test 1, first calculate the difference δ 1 in the misrecognition rate between the user - registered speech rate and the benchmark - registered speech rate calculated on F enroll = F 1 (s enroll 1) - F 1 (Th enroll ), and then obtain the misrecognition rate corresponding to the user - tested speech rate and the misrecognition rate δ test = F 2 (s test 1) - F 2 (s enroll 1).

[0058] Step 4.4: Calculate the sum of the two as the total misrecognition rate difference D s = -(|δ enroll | + |δ test |), normalize D s to obtain the speech - rate security score of the user within the range of [0, 1]. The higher the score, the higher the security of the user's registered speech rate and tested speech rate. In this embodiment, for s enroll 1 = 10 phonemes per second and s test 1 = 12 phonemes per second, the speech - rate security score is 0.92, close to the full score of 1, indicating a relatively high speech - rate security.

[0059] Corresponding to the foregoing embodiment of the method for evaluating the impact of speech rate on voiceprint security, the present application also provides an embodiment of an evaluation system for the impact of speech rate on voiceprint security, which includes:

[0060] A corpus module for obtaining speech data of different speakers at different speech rates;

[0061] A speech - rate calculation module for calculating the speech rate corresponding to each piece of speech data in the corpus, statistically analyzing the speech - rate distribution, and recording the speech rate with the most corresponding speech data as the original speech rate;

[0062] A registered - speech - rate - tested - speech - rate data - set module for constructing different types of registered - speech - rate - tested - speech - rate data sets;

[0063] A voiceprint recognition test module, which is used to perform voiceprint recognition tests by using the data set in the speech rate - test speech rate data set module: registering the voiceprint model with the speech data corresponding to the registered speech rate in the data set and marking the speaker label; then authenticating the voiceprint model with the speech data corresponding to different test speech rates to obtain an authentication result, calculating the false acceptance rate and the false rejection rate, and respectively obtaining the relationship curves between the test speech rate and the false acceptance rate and the false rejection rate, and the registered speech rate - test speech rate - false acceptance rate matrix M FAR and the registered speech rate - test speech rate - false rejection rate matrix M FRR ;

[0064] A function fitting module, which is used to calculate the weighted relationship matrix M = λM FAR and the registered speech rate - test speech rate - false rejection rate matrix M FRR , where λ and μ are the weighted coefficients of M FAR and M FRR respectively; fitting the diagonal data of the weighted relationship matrix M to obtain the function F FAR , and taking the registered speech rate corresponding to the minimum value of the function F FRR as the reference registered speech rate; fitting the function F 1 respectively for different registered speech rates in the weighted relationship matrix M; 1 2 ;

[0065] A security scoring module, which is used to calculate the total misrecognition rate difference according to the user's registered speech rate, the user's test speech rate, the function F 1 and the function F 2 , and perform normalization processing on the total misrecognition rate difference to obtain the speech rate security score in the interval [0, 1].

[0066] Regarding the system in the above embodiments, the specific manners in which each unit or module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0067] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The system embodiments described above are merely illustrative, and the so - called voiceprint recognition test module may or may not be physically separated. In addition, in the present invention, each functional module may be integrated in a processing unit, or each module may exist physically alone, or two or more modules may be integrated in one unit. The above - integrated modules or units may be implemented in the form of hardware or in the form of software functional units, and some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of the present application.​

[0068] The specific embodiments of the present invention are only listed above. Obviously, the present invention is not limited to the above embodiments and there can be many variations. All variations that can be directly derived or associated by those of ordinary skill in the art from the disclosed content of the present invention shall be considered as the protection scope of the present invention.

Claims

1. An evaluation method for the impact of speech rate on voiceprint security, characterized in that, it includes the following steps: Step 1: Obtain speech data of different speakers at different speech rates to form a speaker corpus; Step 2: Calculate the speech rate corresponding to each piece of speech data in the corpus, statistically analyze the speech rate distribution, and record the speech rate corresponding to the most speech data as the original speech rate; Step 3: Construct different types of registered speech rate - test speech rate data sets; Step 4: Use the data set obtained in step 3 to perform voiceprint recognition test: Use the voice data corresponding to the registered speaking speed in the three data sets to register the voiceprint model and obtain the speaker's voiceprint; then use the voice data corresponding to different test speaking speeds to test the voiceprint model to obtain the test results, and calculate the false acceptance rate and false rejection rate to obtain the relationship curve between the test speaking speed and the false acceptance rate and the false rejection rate, and the registration speaking speed-test speaking speed-false acceptance rate matrix M FAR , registration speed-test speed-false rejection rate matrix M FRR ; Step 5: According to the false acceptance rate matrix M of the registered speech rate - test speech rate FAR and the false rejection rate matrix M of the registered speech rate - test speech rate FRR , calculate the weighted relationship matrix M = λM FAR + μM FRR , where λ and μ are the weighted coefficients of M FAR and M FRR respectively; Step 6: Fit the diagonal data of the weighted relationship matrix M to obtain the function F 1 , and use the registration speed corresponding to the minimum value of the function F 1 as the benchmark registration speed; Fit the function F for different registration speeds in the weighted relationship matrix M respectively 2 ; Step 7: Calculate the total misrecognition rate difference based on the user's registered speech rate, the user's tested speech rate, function F 1 and function F 2 Normalize the total misrecognition rate difference to obtain a speech rate security score within the range of [0, 1].

2. The evaluation method for the impact of speech rate on voiceprint security according to claim 1, characterized in that, in step 3, three types of registered speech rate - test speech rate data sets are constructed: The first data set: Select the speech data with the speech rate equal to the original speech rate described in step 2 and the corresponding speakers from the corpus. Use the original speech rate as the registered speech rate, divide the speech data with the original speech rate into two parts, which are used as the registration set and the test set respectively. Use the speech time domain stretching method on the test set to obtain w different speech rates as the test speech rates; The second data set: Select the speech data with the speech rate equal to the original speech rate described in step 2 and the corresponding speakers from the corpus. Use the original speech rate as the registered speech rate, and continue to select w speech rates corresponding to each speaker from the corpus as the test speech rates; The third data set: Select the speech data with the speech rate equal to the original speech rate described in step 2 and the corresponding speakers from the corpus. Divide the speech data with the original speech rate into two parts, which are used as the registration set and the test set respectively. Use the speech time domain stretching method on the registration set and the test set respectively to obtain w different registered speech rates and w different test speech rates.

3. The evaluation method for the impact of speech rate on voiceprint security according to claim 1, characterized in that, calculating the speech rate corresponding to each piece of speech data in the corpus is specifically: Step 2.1: Convert the English letters in the text corresponding to each piece of speech to phonemes, and count the number of phonemes q; Step 2.2: Use speech endpoint detection technology to measure the net speech duration t; Step 2.3: Use the ratio of the number of phonemes to the net speech duration to represent the speech rate, with the unit of phonemes per second.

4. The evaluation method for the impact of speech rate on voiceprint security according to claim 2, characterized in that, The method for obtaining the speech data corresponding to the test speech rate in the first data set is as follows: The speech data at the original speech rate is subjected to time-domain stretching with a time-domain stretching degree of α using the speech time-domain stretching algorithm. The speech duration is lengthened to α times the original, and the corresponding net speech duration becomes αt. Since the number of text phonemes remains unchanged, the speech rate changes from s to When 0 < α < 1, the speech rate becomes faster; when α > 1, the speech rate becomes slower; when α = 1, the speech rate remains unchanged.

5. The evaluation method for the impact of speech rate on voiceprint security according to claim 1, characterized in that, The calculation formula for the difference in total false recognition rate is: D s = -(|δ enroll | + |δ test |) δ enroll = F 1 (s enroll 1) - F 1 (Th enroll ) δ test = F 2 (s test 1) - F 2 (s enroll 1) Among them, D s is the total difference in misrecognition rate, s test 1 is the user's test speech rate, s enroll 1 is the user's registered speech rate, Th enroll is the reference registered speech rate, |.| represents the absolute value operation, F 1 (.) represents the registration speech rate - misrecognition rate fitting function when the registration speech rate in the weighted relationship matrix M is equal to the test speech rate, F 2 (.) represents the test speech rate - misrecognition rate fitting function at a fixed registration speech rate in the weighted relationship matrix M; δ test represents the difference between the misrecognition rate corresponding to the user's test speech rate and the misrecognition rate corresponding to the user's registered speech rate, δ enroll represents the difference between the misrecognition rate corresponding to the user's registered speech rate and the misrecognition rate corresponding to the reference registered speech rate.

6. An evaluation system for the impact of speech rate on voiceprint security, characterized in that, it is used to implement the evaluation method described in any one of claims 1 - 5.

Citation Information

Patent Citations

  • Method and apparatus for adjusting playing speed

    CN109147802A

  • Identity recognition method based on voiceprint recognition

    CN110610709A