Voice identity verification method
By combining environmental voiceprint recognition and contactless voice encryption technology, the voiceprint characteristics of communication are analyzed in real time to encrypt voice signals and verify identity, solving the security and identity verification problems of traditional voice communication and achieving convenient and efficient voice communication security.
Patent Information
- Application Number
- CN202511449559.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2025-12-26
AI Technical Summary
Traditional voice communication security and authentication technologies suffer from risks of password leakage, high complexity, and difficulty in remote implementation.
Combining environmental voiceprint recognition technology and contactless voice encryption technology, this system automatically encrypts and authenticates voice signals by analyzing communication voiceprint features in real time. It uses Fast Fourier Transform and Gaussian weighting for encryption and decryption in the frequency domain, and performs authentication through various feature extraction and similarity calculations.
It achieves convenient, efficient and reliable voice communication security, supports identity recognition on telephones and mobile devices, eliminates interference, and improves recognition accuracy and security.
Smart Images

Figure SMS_4 
Figure SMS_6 
Figure SMS_20
Abstract
Description
Technical Field
[0001] This invention relates to the field of information security and communication technology, and is a voice authentication method. Background Technology
[0002] With the rapid development of information technology, voice communication has become an indispensable part of people's daily lives, and its security issues are becoming increasingly prominent. Traditional voice communication encryption technologies mostly rely on complex passwords or encryption algorithms, but these methods are often at risk of password leakage or algorithm cracking. In addition, traditional authentication technologies also mainly rely on passwords, biometrics, etc., which are not only easy to imitate or steal, but may also be difficult to implement in certain situations (such as remote communication).
[0003] Traditional voice encryption technologies often require users to enter complex passwords or perform specific encryption operations, which not only increases the difficulty of use for users but also poses a risk of password leakage. Voice encryption technology, on the other hand, utilizes the voiceprint characteristics in the environment, combining voice signals with these characteristics to achieve automatic encryption without user intervention. This technology not only improves the convenience of encryption but also, due to its unique environmental dependence, increases the difficulty of cracking it.
[0004] Traditional authentication technologies, such as passwords and biometrics, while improving communication security to some extent, also have significant limitations. Passwords are easily forgotten, stolen, or guessed; while biometric identification technology is highly accurate, it may be difficult to implement or be interfered with in certain situations (such as remote communication). Therefore, finding a more convenient, reliable, and adaptable authentication technology has become an important research direction in the field of information security. Summary of the Invention
[0005] This invention proposes a voice authentication method that combines environmental voiceprint recognition technology, contactless voice encryption technology, and authentication technology. By analyzing the voiceprint characteristics of communication in real time, it achieves automatic encryption and authentication of voice signals. The purpose of this invention is to provide an efficient, convenient, and reliable voice communication security mechanism. By combining voiceprint recognition technology, contactless voice encryption technology, and authentication technology, it aims to solve the security risks and limitations of authentication in traditional voice communication.
[0006] To solve the above-mentioned technical problems, the technical content of the present invention is: a voice authentication method, the steps of which are as follows:
[0007] Step 1: First, use the device at the signal transmitting end to capture the voiceprint characteristics of the communication environment in real time.
[0008] Step 2: The receiver receives the original speech signal S(t), and then uses Fast Fourier Transform (FFT) to transform S(t) from the time domain to the frequency domain to obtain the complex sequence S'F[k]. In the frequency domain, S(t) is encrypted using Gaussian weighting to obtain the encrypted speech signal. ;
[0009] The receiver uses Fast Fourier Transform (FFT) to transform the received speech signal S(t) from the time domain to the frequency domain. Assuming S(t) is the original signal, the result of the FFT is a complex sequence S'. F [k], where k is the index in the frequency domain:
[0010]
[0011] In the frequency domain, S' is weighted using a Gaussian weighting method. F Encryption is performed on [k], which means multiplying the Gaussian function with the frequency domain signal to change its spectral characteristics; assuming the Gaussian function is G[k], the encrypted frequency domain signal C' F [k] is:
[0012]
[0013] The Gaussian function G[k] is usually defined as:
[0014]
[0015] Where A is the amplitude; k0 is the center frequency, which usually corresponds to a specific time point or sample point in the speech signal, depending on the part of the signal to be processed or the features of interest. For example, when processing a frame of a speech signal, the center point of that frame is considered the center frequency; σ is the standard deviation of the Gaussian distribution, which controls the strength of the encryption or the width of the frequency range, and determines the weighting range. The choice of σ value is based on experience and experimentation. You can try different σ values, observe the filtering effect, and choose the value that best suits your application.
[0016] Step 3: After receiving the encrypted signal, the receiving end uses a decryption algorithm to process the frequency domain signal to obtain the decrypted voice signal DF[k].
[0017] To decrypt the signal, the Gaussian weighting effect needs to be removed. Since the Gaussian function is known during encryption, the signal is decrypted by dividing by the Gaussian function.
[0018] .
[0019] Step 4: Transfer the known voiceprint S' F [k] and the decrypted D F[k] Feature extraction is performed on the same frequency band, and the feature vectors are labeled as E in sequence. transmitted E received Two different similarity functions are used to calculate the similarity between two feature vectors, and a weighted calculation is performed to obtain a weighted comprehensive similarity value.
[0020] The two audio segments were subjected to the following four feature extraction methods, and the extracted features were F1, F2, F3, and F4 in that order.
[0021] Step 4.1) Mel Frequency Cepstral Coefficients (MFCC): The signal is pre-emphasized and framed. A Hanning window is used for each frame, i.e., a Fast Fourier Transform (FFT) is performed on each frame of the windowed signal to calculate the power spectrum. The power spectrum is processed using a Mel filter bank. A Discrete Cosine Transform (DCT) is performed on each frame to obtain the MFCC.
[0022]
[0023] in, Let N be the Hanning window of the nth frame, where N is the window size.
[0024] Step 4.2) Short-time energy: The formula for calculating short-time energy is:
[0025]
[0026] in, Indicates the first Frame energy, Indicates the first The first frame Each sample value Frame length;
[0027] Step 4.3) Fundamental Frequency Extraction: There is no fixed formula for fundamental frequency extraction, as it depends on specific algorithms or methods, such as autocorrelation, cepstral method, etc. A simple example of an autocorrelation formula is:
[0028]
[0029] in, Represents the autocorrelation function. It is a time series signal. It is a lag quantity. It is the signal length.
[0030] Step 4.4) Loudness: The calculation formula is:
[0031]
[0032] Among them, It's loudness, it's Sound intensity It is a reference strength. It is a constant;
[0033] After feature extraction, weighted calculations are performed to obtain E. transmitted,i and E received,i The feature extraction formula is:
[0034]
[0035] Then, similarity is calculated, and a comprehensive similarity value is obtained by weighting the results using two different similarity functions.
[0036]
[0037]
[0038] +
[0039] Among them, w i For the weight, F i It is a characteristic; E transmitted,i and E received,i These are the i-th components of the transmitted and received environmental acoustic signature feature vectors, respectively. is a small positive number to avoid division by zero; CombinedSimilarity is the weighted combined similarity value; W1 and W2 are weights, which sum to 1. The weights are assigned based on the trust level of the two similarity functions or their importance in the application. More importantly, W1 is larger.
[0040] Step 5: Compare the overall similarity value with the threshold. If the overall similarity value is greater than the threshold, the identity verification is successful; otherwise, the verification fails.
[0041]
[0042] Threshold is a preset threshold value.
[0043] Compared with existing technologies, the beneficial effects of this invention are as follows: The voice authentication method proposed in this invention extracts unique acoustic features for authentication by collecting and analyzing user voice samples. This method not only supports voice liveness detection to ensure the communicator is a real individual, but also performs identity recognition on telephone communications and mobile devices, effectively eliminating interference and thus improving recognition accuracy. The method is simple and easy to use, while possessing high accuracy and security, capable of processing large amounts of data, and meeting the needs of different users and environments. Detailed Implementation
[0044] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0045] A voice authentication method, comprising the following steps:
[0046] Step 1: First, use the device at the signal transmitting end to capture the voiceprint characteristics of the communication environment in real time.
[0047] Step 2: The receiver receives the original speech signal S(t), and then uses Fast Fourier Transform (FFT) to transform S(t) from the time domain to the frequency domain to obtain the complex sequence S'F[k]. In the frequency domain, S(t) is encrypted using Gaussian weighting to obtain the encrypted speech signal. ;
[0048] The receiver uses Fast Fourier Transform (FFT) to transform the received speech signal S(t) from the time domain to the frequency domain. Assuming S(t) is the original signal, the result of the FFT is a complex sequence S'. F [k], where k is the index in the frequency domain:
[0049]
[0050] In the frequency domain, S' is weighted using a Gaussian weighting method. F Encryption is performed on [k], which means multiplying the Gaussian function with the frequency domain signal to change its spectral characteristics; assuming the Gaussian function is G[k], the encrypted frequency domain signal C' F [k] is:
[0051]
[0052] The Gaussian function G[k] is usually defined as:
[0053]
[0054] Where A is the amplitude; k0 is the center frequency, which usually corresponds to a specific time point or sample point in the speech signal, depending on the part of the signal to be processed or the features of interest. For example, when processing a frame of a speech signal, the center point of that frame is considered the center frequency; σ is the standard deviation of the Gaussian distribution, which controls the strength of the encryption or the width of the frequency range, and determines the weighting range. The choice of σ value is based on experience and experimentation. You can try different σ values, observe the filtering effect, and choose the value that best suits your application.
[0055] Step 3: After receiving the encrypted signal, the receiving end uses a decryption algorithm to process the frequency domain signal to obtain the decrypted voice signal DF[k].
[0056] To decrypt the signal, the Gaussian weighting effect needs to be removed. Since the Gaussian function is known during encryption, the signal is decrypted by dividing by the Gaussian function.
[0057] .
[0058] Step 4: Transfer the known voiceprint S' F [k] and the decrypted D F [k] Feature extraction is performed on the same frequency band, and the feature vectors are labeled as E in sequence. transmitted E received Two different similarity functions are used to calculate the similarity between two feature vectors, and a weighted calculation is performed to obtain a weighted comprehensive similarity value.
[0059] The two audio segments were subjected to the following four feature extraction methods, and the extracted features were F1, F2, F3, and F4 in that order.
[0060] Step 4.1) Mel Frequency Cepstral Coefficients (MFCC): The signal is pre-emphasized and framed. A Hanning window is used for each frame, i.e., a Fast Fourier Transform (FFT) is performed on each frame of the windowed signal to calculate the power spectrum. The power spectrum is processed using a Mel filter bank. A Discrete Cosine Transform (DCT) is performed on each frame to obtain the MFCC.
[0061]
[0062] in, Let N be the Hanning window of the nth frame, where N is the window size.
[0063] Step 4.2) Short-time energy: The formula for calculating short-time energy is:
[0064]
[0065] in, Indicates the first Frame energy, Indicates the first The first frame Each sample value Frame length;
[0066] Step 4.3) Fundamental Frequency Extraction: There is no fixed formula for fundamental frequency extraction, as it depends on specific algorithms or methods, such as autocorrelation, cepstral method, etc. A simple example of an autocorrelation formula is:
[0067]
[0068] in, Represents the autocorrelation function. It is a time series signal. It is a lag quantity. It is the signal length.
[0069] Step 4.4) Loudness: The calculation formula is:
[0070]
[0071] Among them, It's loudness, it's Sound intensity It is a reference strength. It is a constant;
[0072] After feature extraction, weighted calculations are performed to obtain E. transmitted,i and E received,i The feature extraction formula is:
[0073]
[0074] Then, similarity is calculated, and a comprehensive similarity value is obtained by weighting the results using two different similarity functions.
[0075]
[0076]
[0077] +
[0078] Among them, w i For the weight, F i It is a characteristic; E transmitted,i and E received,i These are the i-th components of the transmitted and received environmental acoustic signature feature vectors, respectively. is a small positive number to avoid division by zero; CombinedSimilarity is the weighted combined similarity value; W1 and W2 are weights, which sum to 1. The weights are assigned based on the trust level of the two similarity functions or their importance in the application. More importantly, W1 is larger.
[0079] Step 5: Compare the overall similarity value with the threshold. If the overall similarity value is greater than the threshold, the identity verification is successful; otherwise, the verification fails.
[0080]
[0081] Threshold is a preset threshold value.
Claims
1. A voice authentication method, characterized in that, The steps are as follows: Step 1: At the signal transmitting end, first use the device to capture the acoustic signature characteristics of the communication environment in real time; Step 2: The receiver receives the original speech signal S(t), and then uses Fast Fourier Transform (FFT) to transform S(t) from the time domain to the frequency domain to obtain the complex sequence S'F[k]. In the frequency domain, S(t) is encrypted using Gaussian weighting to obtain the encrypted speech signal. ; Step 3: After receiving the encrypted signal, the receiving end uses a decryption algorithm to process the frequency domain signal to obtain the decrypted voice signal DF[k]. Step 4: Transfer the known voiceprint S' F [k] and the decrypted D F [k] Feature extraction is performed on the same frequency band, and the feature vectors are labeled as E in sequence. transmitted E received Two different similarity functions are used to calculate the similarity between two feature vectors, and a weighted calculation is performed to obtain a weighted comprehensive similarity value. Step 5: Compare the overall similarity value with the threshold. If the overall similarity value is greater than the threshold, the identity verification is successful; otherwise, the verification fails.
2. The voice authentication method according to claim 1, characterized in that: In step 2, the specific method is as follows: The receiver uses Fast Fourier Transform (FFT) to transform the received speech signal S(t) from the time domain to the frequency domain. Assuming S(t) is the original signal, the result of the FFT is a complex sequence S'. F [k], where k is the index in the frequency domain: In the frequency domain, S' is weighted using a Gaussian weighting method. F [k] performs an encryption operation by multiplying the Gaussian function with the frequency domain signal to change its spectral characteristics; assuming the Gaussian function is G[k], the encrypted frequency domain signal C' F [k] is: The Gaussian function G[k] is usually defined as: Where A is the amplitude, k0 is the center frequency, and σ is the standard deviation of the Gaussian distribution, which determines the weighting range. The choice of the value of σ is based on experience and experiments.
3. The voice authentication method according to claim 1, characterized in that: In step 3, the specific method is as follows: To remove the influence of Gaussian weighting and decrypt the signal, since the Gaussian function is known during encryption, the signal is decrypted by dividing by the Gaussian function. 。 4. The voice authentication method according to claim 1, characterized in that: In step 4, the specific method is as follows: The two audio segments were subjected to the following four feature extraction methods, and the extracted features were F1, F2, F3, and F4 in that order. Step 4.1) Mel Frequency Cepstral Coefficients (MFCC): The signal is pre-emphasized and framed. A Hanning window is used for each frame, i.e., a Fast Fourier Transform (FFT) is performed on each frame of the windowed signal to calculate the power spectrum. The power spectrum is processed using a Mel filter bank. A Discrete Cosine Transform (DCT) is performed on each frame to obtain the MFCC. in, Let N be the Hanning window of the nth frame, where N is the window size. Step 4.2) Short-time energy: The formula for calculating short-time energy is: in, Indicates the first Frame energy, Indicates the first The first frame Each sample value Frame length; Step 4.3) Extract the fundamental frequency: The autocorrelation formula is: in, Represents the autocorrelation function. It is a time series signal. It is a lag quantity. It is the signal length; Step 4.4) Loudness: The calculation formula is: Among them, It's loudness, it's Sound intensity It is a reference strength. It is a constant; After feature extraction, weighted calculations are performed to obtain E. transmitted,i and E received,i The feature extraction formula is: Then, similarity is calculated, and a comprehensive similarity value is obtained by weighting the results using two different similarity functions. + Among them, w i For the weight, F i It is a characteristic; E transmitted,i and E received,i These are the i-th components of the transmitted and received environmental acoustic signature feature vectors, respectively. W1 is a small positive number used to avoid division by zero; Combined Similarity is the weighted combined similarity value; W1 and W2 are the weights, which sum to 1. The weights are assigned based on the trust level of the two similarity functions or their importance in the application. More importantly, W1 is larger.