A voiceprint matching method
By constructing a voiceprint noise reduction model and feature fusion technology, the accuracy of voiceprint matching in noise environments is solved, efficient voiceprint recognition in complex environments is achieved, and the accuracy and robustness of voiceprint matching is improved.
Patent Information
- Application Number
- CN202510266563.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The existing voiceprint matching methods are insufficient in complex environments, making it difficult to build a voiceprint noise reduction model in noise environments, and it is difficult to extract time and frequency domain features and fusion, resulting in a decrease in the accuracy of voiceprint matching.
By constructing a voiceprint noise reduction model, the noise pattern samples are preprocessed and feature extraction are used by generators and discriminators to extract time domain features such as short-term energy, short-term average amplitude and short-term zero-crossing rate. Combined with frequency domain features, a recurrent neural network is used to form a voiceprint feature fusion model library to calculate the similarity between the voiceprint to be verified and registered voiceprints.
Effective noise reduction in noise environments, improve the quality of voiceprint samples, realize multi-dimensional feature fusion, improve the accuracy and efficiency of voiceprint matching, and can show good performance in different environments.
Smart Images

Figure CN119763585B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voiceprint recognition, and particularly relates to a voiceprint matching method. Background Art
[0002] With the rapid development of information technology, voiceprint recognition, as a convenient and efficient identity authentication method, has been widely used in many fields, such as security monitoring, financial transactions, smart homes, etc. However, the accuracy and robustness of existing voiceprint matching methods in complex environments still need to be improved. For example, when there are background noise interferences and differences in recording devices, etc., the accuracy rate of voiceprint matching may decrease significantly, unable to meet the requirements of high-precision identity authentication in practical applications. It is difficult to construct a voiceprint noise reduction model under noisy environmental conditions to perform noise reduction processing on the collected voiceprints to generate clean voiceprint samples, and it is also difficult to extract various time-domain and frequency-domain features, and add them to the corresponding weight coefficients to obtain fused time-domain feature values and fused frequency-domain feature values. There is a lack of constructing a voiceprint feature fusion model library with the fused time-domain feature values and fused frequency-domain feature values as feature vectors, and using cosine similarity to calculate the similarity between the registered voiceprint and the voiceprint to be verified to determine whether the voiceprint matching is successful. Summary of the Invention
[0003] The present invention aims to solve at least one of the technical problems existing in the prior art; for this purpose, the present invention proposes a voiceprint matching method to solve the following technical problems:
[0004] It is difficult to construct a voiceprint noise reduction model under noisy environmental conditions to perform noise reduction processing on the collected voiceprints to generate clean voiceprint samples, and it is also difficult to extract various time-domain and frequency-domain features, and add them to the corresponding weight coefficients to obtain fused time-domain feature values and fused frequency-domain feature values. There is a lack of constructing a voiceprint feature fusion model library with the fused time-domain feature values and fused frequency-domain feature values as feature vectors, and using cosine similarity to calculate the similarity between the registered voiceprint and the voiceprint to be verified to determine whether the voiceprint matching is successful.
[0005] To achieve the above object, the present invention proposes a voiceprint matching method, which includes the following steps:
[0006] S1: Collect voiceprint samples to be matched through an audio acquisition device, including registered voiceprint samples and voiceprints to be verified. After preprocessing and annotating the collected original noisy voiceprint samples, constructing a generator and a discriminator, forming a training data pair with the original noisy voiceprint samples and the corresponding generated voiceprint samples, constructing a voiceprint noise reduction model, and performing preprocessing operations on the collected voiceprint samples;
[0007] S2: Extract the short-time energy, short-time average amplitude, and short-time zero-crossing rate from the preprocessed voiceprint samples, construct a weight vector and fuse it with the time-domain features to obtain the fused time-domain feature values, extract the frequency-domain features, and perform a fused weighting operation on the energy mean, variance, and ratio values to obtain the fused frequency-domain feature values;
[0008] S3: Input the obtained fused time-domain feature values and fused frequency-domain feature values into a recurrent neural network model to form a voiceprint feature fusion model library;
[0009] S4: Calculate the similarity between the voiceprint feature vector to be verified and each registered voiceprint feature vector in the voiceprint feature fusion model library, and determine whether the similarity is less than a preset similarity threshold.
[0010] Further, the step S1 includes the following steps:
[0011] Under the conditions of different noise environments, collect the voiceprint samples of users during registration through an audio acquisition device, convert the collected voiceprint samples into a unified audio format, and input the registered voiceprint samples into a voiceprint noise reduction model for preprocessing; use the same audio acquisition device, and under the conditions similar to the registration voiceprint sample acquisition environment, collect the voiceprint samples to be verified, convert them into a unified audio format, and at the same time, perform corresponding preprocessing operations on the voiceprint samples to be verified.
[0012] Further, the voiceprint noise reduction model includes the following steps:
[0013] Collect the voiceprint samples of different users in different noise environments, perform preprocessing operations on the voiceprint data, including sampling rate conversion, silence removal, and pre-emphasis, convert them into a form suitable for model input, label the collected original noisy voiceprint data, label the noise type, noise level, speech content, and user identity it contains, and for the voiceprint generated by the generator, label its source and the corresponding original noisy voiceprint identifier;
[0014] Design a shared input layer in the generator and discriminator, perform the same preprocessing operations on the input voiceprint, including Fourier transform and Mel-frequency cepstral coefficient extraction. After converting the voiceprint signal into a unified feature representation, use the features extracted from the noisy voiceprint signal as the input to construct the generator. The discriminator adopts a dual-input structure, one input is the feature representation of the original noisy voiceprint, and the other input is the feature representation of the generated voiceprint. A comparison module is set inside the discriminator to perform a comparative analysis on the original noisy voiceprint feature and the generated voiceprint feature;
[0015] The original noisy voiceprint samples and the corresponding generated voiceprint samples are combined to form training data pairs. At the same time, they are input into the discriminator, and the model is trained using the training set. By iteratively updating the parameters of the model multiple times, the model learns the mapping relationship from noisy voiceprints to clean voiceprints, and calculates the root mean square error between the original voiceprint signal and the denoised voiceprint signal.
[0016] Further, the calculation of the root mean square error between the original voiceprint signal and the denoised voiceprint signal includes the following steps:
[0017] The original voiceprint signal is expressed as , and the denoised voiceprint signal is expressed as . For each sampling point, calculate the square of the difference between the original voiceprint signal and the denoised voiceprint signal. After summing all the squared differences, divide the sum of squared errors by the total number of all sampling points, and take the square root of the mean square error to obtain the root mean square error;
[0018] Root mean square error calculation formula:
[0019]
[0020] Among them, represents the root mean square error, represents the total number of samples, represents the th original voiceprint signal, represents the th denoised voiceprint signal;
[0021] Based on the experience of historical voiceprint denoising projects, set the RMSE empirical threshold:
[0022] If is less than the RMSE empirical threshold, it means that the difference between the original voiceprint signal and the denoised voiceprint signal is small, and the denoising effect is good. Deploy the trained model in the system to denoise the collected voiceprint samples;
[0023] If is greater than or equal to the RMSE empirical threshold, it means that the difference between the original voiceprint signal and the denoised voiceprint signal is large, and the denoising effect is poor. Adjust the model.
[0024] Further, the extraction of short-time energy, short-time average amplitude, and short-time zero-crossing rate, and the construction of a weight vector to fuse with the time-domain features to obtain the fused time-domain feature value include the following steps:
[0025] Extract features from the preprocessed voiceprint samples. The time-domain features include short-time energy, short-time average amplitude, and short-time zero-crossing rate. Determine the frame length and frame shift, segment the voiceprint signal according to the determined frame length and frame shift, and calculate the corresponding short-time energy for the short-time energy of this frame; for each frame of signal, calculate the absolute value of each sampling point therein, add up the absolute values of all sampling points in this frame to obtain the amplitude sum of this frame of signal, and then divide by the frame length to obtain the short-time average amplitude of this frame of signal; set the zero-crossing threshold, initialize two timers, record the number of times the signal value changes from positive to negative and the difference is greater than the set threshold and the number of times the signal value changes from negative to positive and the difference is greater than the preset threshold respectively, traverse the samples within the frame, and after the traversal ends, add the values of the two counters to obtain the short-time zero-crossing rate of this frame;
[0026] Determine the weight coefficients of short-time energy, short-time average amplitude, and short-time zero-crossing rate according to application requirements and data characteristics, and construct a weight vector , for each frame, perform weighted fusion on short-time energy, short-time average amplitude, and short-time zero-crossing rate according to the weight vector to obtain the fused time-domain feature value;
[0027] Fused time-domain feature value calculation formula:
[0028]
[0029] Among them, represents the fused time-domain feature value of the th frame, represents the short-time energy of the th frame, represents the short-time average amplitude of the th frame, represents the short-time zero-crossing rate of the th frame, , and represent the corresponding weight coefficients;
[0030] After performing the weighted fusion operation on all frames, obtain a fused time-domain feature value sequence , among them, represents the number of frames after frame segmentation.
[0031] Furthermore, for the extraction of frequency-domain features, perform a fusion weighted operation on the energy mean, variance, and proportion value to obtain the fused frequency-domain feature value, including the following steps:
[0032] After windowing each frame of the signal, perform a fast Fourier transform to calculate the amplitude and phase information of the signal at different frequencies. According to the result of the fast Fourier transform, calculate the power spectrum of each frame of the signal by squaring the amplitude spectrum of the frequency spectrum and then dividing by the duration. Determine the frequency resolution based on the set sampling frequency and signal length, multiply the frequency resolution by the frequency index to obtain the corresponding actual frequency. Use the actual frequency as the abscissa and the corresponding power spectrum diagram as the ordinate to plot the power spectrum diagram, calculate the energy mean and variance of each frequency component in the power spectrum diagram, and find the maximum power spectrum value and the minimum power spectrum value in the power spectrum and subtract them from the mean respectively to obtain the corresponding differences, and then divide to get the ratio value, and take the absolute value;
[0033] Perform a fusion weighting operation on the energy mean, variance and ratio value to obtain the fusion frequency domain eigenvalue;
[0034] Formula for calculating the fusion frequency domain eigenvalue:
[0035]
[0036]
[0037] Among them, represents the power spectrum mean, represents the number of frequency components, represents the th power spectrum value of the frequency component, represents the fusion frequency domain eigenvalue, represents the maximum power spectrum value, represents the minimum power spectrum value, , and represent the corresponding weight coefficients.
[0038] Furthermore, step S3 includes the following steps:
[0039] Collect the voiceprint sample data of users, label each sample, and the labeling content includes the user's name and speech content. Input the collected sample data into a noise reduction voiceprint model for noise reduction, then perform frame segmentation and window function processing, and unify the number of sampling points and normalize the amplitude. According to the selected fused time-domain feature values and fused frequency-domain feature values, extract features from the preprocessed data, splice the extracted fused time-domain feature values and fused frequency-domain feature values into a feature vector, create a user table to store the relevant information of users, and create a voiceprint sample table to store the feature vectors of samples and the corresponding labeling information, which is associated with the user table through a foreign key. Select a recurrent neural network as the framework for building the model, divide the labeled data into training set data and input it into the recurrent neural network model for training. The model learns and recognizes the voiceprint features of different users, use the trained model to predict the data for the test set part, obtain the voiceprint feature prediction results of different users, and store the feature vectors of all users and the corresponding prediction results to form a voiceprint feature fusion model library.
[0040] Store the extracted feature vector together with the unique identifier of the registered user into the voiceprint feature fusion model library. If it is necessary to verify the voiceprint sample, collect the voiceprint sample to be verified in the same way as in the registration stage and perform preprocessing operations. Obtain the feature vector of the voiceprint to be verified through the same feature extraction and calculation steps as in the registration stage for the preprocessed voiceprint sample to be verified, and use the cosine similarity method to calculate the similarity between the feature vector of the voiceprint to be verified and each registered voiceprint feature vector in the voiceprint feature fusion model library;
[0041] Determine a preset similarity threshold according to experience, and compare the calculated similarity with the preset similarity threshold:
[0042] If the similarity is less than the preset similarity threshold, it is determined that the voiceprint matching fails;
[0043] If the similarity is greater than or equal to the preset similarity threshold, it is determined that the voiceprint matching is successful.
[0044] Advantages of the present invention:
[0045] The present invention effectively identifies and weakens these background noise components for the noisy voiceprint data by constructing a voiceprint noise reduction model, making the main sound of the voiceprint sample more prominent and significantly improving the sound quality. Moreover, the features of the voiceprint often hide in the voiceprint signal, and the existence of noise may cover or distort these features. The noise reduction model makes the features of the voiceprint more obvious and prominent by reducing the interference of noise, and can automatically adjust the noise reduction parameters according to different environmental conditions to effectively reduce the noise of voiceprint samples in various environments, enabling the model to exhibit good performance in different actual environments;
[0046] The present invention extracts corresponding features from two levels of time domain and frequency domain respectively, and obtains the fused time domain feature value and the fused frequency domain feature value through processing and calculation, realizing the full utilization of the information of the voiceprint signal in different domains, reflecting the strength, rhythm and detailed information of the sound, splicing the two feature values into a feature vector, realizing the effective fusion of multi-dimensional features, making up for the limitations of single-domain feature description, selecting a recurrent neural network as the framework for building the model, training the model, the model can learn the voiceprint feature patterns of different users, storing the feature vectors of all users and the corresponding prediction results, forming a voiceprint feature fusion model library, this matching process is more efficient and can give accurate recognition results in a short time. As time goes by and new data increases, the model can be continuously trained and optimized using the data in the model library. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In order to better understand the technical content of the present invention, specific embodiments are hereby given and described in conjunction with the accompanying drawings as follows.
[0049] Please refer to Figure 1 , Figure 1 shown in the schematic flowchart of the method of the present invention. The present invention provides a voiceprint matching method, including the following steps:
[0050] S1: Collect the voiceprint samples to be matched through an audio acquisition device, including registered voiceprint samples and voiceprint samples to be verified. After preprocessing and labeling the collected original noisy voiceprint samples, constructing a generator and a discriminator, forming a training data pair with the original noisy voiceprint samples and the corresponding generated voiceprint samples, constructing a voiceprint noise reduction model, and performing preprocessing operations on the collected voiceprint samples;
[0051] S2: Extract the short-time energy, short-time average amplitude and short-time zero-crossing rate from the preprocessed voiceprint samples, construct a weight vector to fuse with the time domain features to obtain the fused time domain feature value, extract the frequency domain features, and perform a fusion weighting operation on the energy mean, variance and proportion value to obtain the fused frequency domain feature value;
[0052] S3: Input the obtained fused time domain feature value and fused frequency domain feature value into a recurrent neural network model to form a voiceprint feature fusion model library;
[0053] S4: Calculate the similarity between the voiceprint feature vector to be verified and each registered voiceprint feature vector in the voiceprint feature fusion model library, and determine whether the similarity is less than a preset similarity threshold.
[0054] Specifically, measure the noise decibels of the registration environment, record the environmental noise level, and in the selected registration environment, convert the collected voiceprint sample audio files into a standard audio format uniformly. Use the same audio acquisition device as during registration to collect the voiceprint samples to be verified. Combine the original noisy voiceprint samples and the corresponding generated voiceprint samples into training data pairs, and input the training data pairs into the discriminator. Use the training set to train the model. For each test sample, calculate the root mean square error between the original voiceprint signal and the denoised voiceprint signal; construct a weight vector, and for each frame, perform weighted fusion on the short-time energy, short-time average amplitude, and short-time zero-crossing rate according to the weight vector to obtain the fused time-domain eigenvalue. Perform fusion and weighting operations on the energy mean, variance, and ratio value to obtain the fused frequency-domain eigenvalue; splice the fused time-domain eigenvalue and the fused frequency-domain eigenvalue of the same frame into a feature vector in a certain order, combine the feature vectors of all frames to form the feature matrix of the entire voiceprint sample for model input. Divide a part of the labeled data as the training set, input the data of the training set into the recurrent neural network model, and use the trained model to predict the data in the test set part to obtain the voiceprint feature prediction results of different users; compare each calculated similarity value with the preset similarity threshold. If all similarity values are less than the preset similarity threshold, it is determined that the voiceprint matching fails, that is, the voiceprint to be verified does not match any user in the registered voiceprint library. If at least one similarity value is greater than or equal to the preset similarity threshold, it is determined that the voiceprint matching is successful, and the registered user with the maximum similarity value is output as the matching result.
[0055] In one embodiment of the present invention, step S1 includes the following steps:
[0056] Under the conditions of different noise environments, collect the voiceprint samples of the user during registration through the audio acquisition device, convert the collected voiceprint samples into a unified audio format, and input the registered voiceprint samples into the voiceprint denoising model for preprocessing; use the same audio acquisition device and under the conditions similar to the registration voiceprint sample acquisition environment, collect the voiceprint samples to be verified and convert them into a unified audio format. At the same time, perform corresponding preprocessing operations on the voiceprint samples to be verified.
[0057] Specifically, measure the noise decibels of the registration environment, record the environmental noise level, and let the user read the pre-prepared text content through an audio acquisition device in the selected registration environment. Convert the collected voiceprint sample audio files into a standard audio format, and input the registered voiceprint samples in the same unified format into a pre-trained voiceprint noise reduction model, which can remove the interference of the registered yellow environmental noise to a certain extent; try to make the acquisition environment of the voiceprint sample to be verified similar to the noise conditions during registration, use the same audio acquisition device as during registration to collect the voiceprint sample to be verified, and require the user to speak in the same way as during registration and read the same text content in the verification environment. Collect the voiceprint samples to be verified multiple times, with the time and method of each collection being consistent with that during registration. Also convert the collected voiceprint sample audio files to be verified into the previously determined audio format, and perform preprocessing operations similar to those of the registered voiceprint samples on the voiceprint samples to be verified.
[0058] In one embodiment of the present invention, the voiceprint noise reduction model includes the following steps:
[0059] Collect voiceprint samples of different users in different noise environments, and perform preprocessing operations on the voiceprint data, including sampling rate conversion, silence removal, and pre-emphasis, to convert it into a form suitable for model input. Label the collected original noisy voiceprint data, indicating the noise type, noise level, speech content, and user identity it contains. For the voiceprint generated by the generator, label its source and the corresponding original noisy voiceprint identifier;
[0060] Design a shared input layer in the generator and discriminator, and perform the same preprocessing operations on the input voiceprint, including Fourier transform and Mel-frequency cepstral coefficient extraction. After converting the voiceprint signal into a unified feature representation, use the features extracted from the noisy voiceprint signal as the input to construct the generator. The discriminator adopts a dual-input structure, with one input being the feature representation of the original noisy voiceprint and the other input being the feature representation of the generated voiceprint. A comparison module is set inside the discriminator to compare and analyze the original noisy voiceprint feature and the generated voiceprint feature;
[0061] Form a training data pair with the original noisy voiceprint sample and the corresponding generated voiceprint sample, and at the same time, input it into the discriminator. Use the training set to train the model, and update the parameters of the model through multiple iterations, so that the model learns the mapping relationship from the noisy voiceprint to the clean voiceprint, and calculate the root mean square error between the original voiceprint signal and the noise-reduced voiceprint signal.
[0062] Specifically, recruit a large number of users and let them use audio acquisition devices to record their speech in different noise environments. For example, each user reads a fixed text content, ensuring that each user records multiple times in each noise environment to increase the diversity and quantity of data. Label the noise types according to the acquisition environment. For example, street noise includes vehicle driving sounds, pedestrian footsteps, etc., and office noise includes computer fan sounds, conversations, etc. Use professional tools to measure and record the noise level of each sample in decibels, accurate to one decimal place. Transcribe the speech content of the user in each sample and label each sample with the unique identifier of the user who collected the sample; label the voiceprint generated by the generator, indicating which original noisy voiceprint samples the generated voiceprint is generated from by the generator, such as "generated from the samples of user A in the office noise environment" and associate the unique identifier of the original noisy voiceprint sample corresponding to the generated voiceprint, so as to be able to accurately match and analyze in subsequent processing; check the sampling rates of all voiceprint samples, and if the sampling rates are inconsistent, use professional audio processing software to convert all samples to a unified sampling rate. Detect the silent parts in the audio by setting an energy threshold, using the short-time energy analysis method, calculate the energy of each frame of audio, and compare it with the threshold, and delete the detected silent parts from the audio. Perform pre-emphasis processing on the audio signal to enhance the high-frequency components; create a shared input layer in the model architecture, determine that the data format of the input layer matches the data format of the preprocessed voiceprint data, and perform Fourier transform and Mel-frequency cepstral coefficient extraction on the data; the input of the generator is the noisy voiceprint features after preprocessing and feature extraction, learn the mapping relationship from the noisy voiceprint features to the clean voiceprint features, and the output of the generator is the generated clean voiceprint representation. The discriminator adopts a dual-input structure, one input is the feature representation of the original noisy voiceprint, and the other input is the feature representation of the generated voiceprint. Set a comparison module inside the discriminator to perform comparative analysis on the original noisy voiceprint features and the generated voiceprint features. The output of the discriminator is a scalar value, indicating whether a pair of input voiceprint features come from the same category, that is, judging whether the generated voiceprint is realistic enough and indistinguishable from the original noisy voiceprint. Combine the original noisy voiceprint samples and the corresponding generated voiceprint samples into training data pairs, input the training data pairs into the discriminator, and use the training set to train the model. Repeat the training process for multiple iterations. In the test stage, select a part of the preprocessed original noisy voiceprint samples as the test set, input these samples into the trained generator to obtain the generated denoised voiceprint signal. At the same time, prepare the corresponding original noisy voiceprint signal. For each test sample, calculate the root mean square error between the original voiceprint signal and the denoised voiceprint signal.
[0063] In one embodiment of the present invention, calculating the root mean square error between the original voiceprint signal and the denoised voiceprint signal includes the following steps:
[0064] The original voiceprint signal is denoted as , and the voiceprint signal after noise reduction is denoted as . For each sampling point, calculate the square of the difference between the original voiceprint signal and the voiceprint signal after noise reduction, add up all the squared differences, divide the sum of squared errors by the total number of all sampling points, and take the square root of the mean squared error to obtain the root mean square error;
[0065] Root mean square error calculation formula:
[0066]
[0067] where represents the root mean square error, represents the total number of samples, represents the th original voiceprint signal, represents the th voiceprint signal after noise reduction;
[0068] Based on the experience of historical voiceprint noise reduction projects, set the RMSE empirical threshold:
[0069] If is less than the RMSE empirical threshold, it indicates that the difference between the original voiceprint signal and the voiceprint signal after noise reduction is small, and the noise reduction effect is good. Deploy the trained model in the system to perform noise reduction on the collected voiceprint samples;
[0070] If is greater than or equal to the RMSE empirical threshold, it indicates that the difference between the original voiceprint signal and the voiceprint signal after noise reduction is large, and the noise reduction effect is poor. Adjust the model.
[0071] Specifically, obtain the original voiceprint signal dataset and the corresponding voiceprint signal dataset after noise reduction processing. For each corresponding sampling point in a pair of original voiceprint signals and noise-reduced voiceprint signals, calculate their difference. For example, for the \(i\)-th sampling point, the difference is the value of the sampling point of the original voiceprint signal - the value of the sampling point of the noise-reduced voiceprint signal. Square each difference to obtain the squared difference of each sampling point. Add up the squared differences of all sampling points to obtain the sum of squared errors. Divide the sum of squared errors by the total number of sampling points to obtain the mean squared error. Take the square root of the mean squared error to obtain the root mean squared error (RMSE); According to the experience of historical voiceprint noise reduction projects, determine an RMSE empirical threshold. This threshold is obtained by analyzing the RMSE values between the original voiceprint signals and the noise-reduced voiceprint signals in a large number of successful noise reduction cases in the past. It represents the range of acceptable signal differences in the actual application scenario; Compare the calculated RMSE value with the set empirical threshold. If the RMSE is less than the RMSE empirical threshold, it indicates that the difference between the original voiceprint signal and the noise-reduced voiceprint signal is small and the noise reduction effect is good. Deploy the trained model in the system to perform noise reduction processing on newly collected voiceprint samples. Otherwise, it indicates that the difference between the original voiceprint signal and the noise-reduced voiceprint signal is large and the noise reduction effect is poor. Adjust the model. After adjusting the model, repeat the above steps of calculating RMSE until the RMSE value is less than the empirical threshold, and then deploy the model after ensuring that it has a good noise reduction effect.
[0072] In one embodiment of the present invention, the steps of extracting short-time energy, short-time average amplitude, and short-time zero-crossing rate, constructing a weight vector and fusing it with time-domain features to obtain a fused time-domain feature value include the following steps:
[0073] Extract features from the preprocessed voiceprint samples. The time-domain features include short-time energy, short-time average amplitude, and short-time zero-crossing rate. Determine the frame length and frame shift, and segment the voiceprint signal according to the determined frame length and frame shift. Calculate the short-time energy corresponding to the frame; For each frame of the signal, calculate the absolute value of each sampling point in it, add up the absolute values of all sampling points in the frame to obtain the amplitude sum of the frame signal, and then divide by the frame length to obtain the short-time average amplitude of the frame signal; Set the zero-crossing threshold, initialize two timers, record the number of times the signal value changes from positive to negative and the difference is greater than the set threshold, and the number of times the signal value changes from negative to positive and the difference is greater than the preset threshold respectively. Traverse the samples within the frame. After the traversal ends, add the values of the two counters to obtain the short-time zero-crossing rate of the frame;
[0074] Determine the weight coefficients of short-time energy, short-time average amplitude, and short-time zero-crossing rate according to application requirements and data characteristics, and construct a weight vector , for each frame, the short-time energy, short-time average amplitude, and short-time zero-crossing rate are weighted and fused according to the weight vector to obtain the fused time-domain eigenvalue;
[0075] Fused time-domain eigenvalue calculation formula:
[0076]
[0077] where, represents the fused time-domain eigenvalue of the th frame, represents the short-time energy of the th frame, represents the short-time average amplitude of the th frame, represents the short-time zero-crossing rate of the th frame, , and represent the corresponding weight coefficients;
[0078] After performing the weighted fusion operation on all frames, a fused time-domain eigenvalue sequence is obtained, where represents the number of frames after frame segmentation.
[0079] Specifically, the frame length is determined according to the characteristics of the voiceprint signal and application requirements. For example, if the voiceprint signal changes relatively slowly, a longer frame length can be selected, such as 30 ms - 50 ms; if the voiceprint signal changes relatively fast, a shorter frame length may be required, such as 10 ms - 20 ms. The frame length is usually expressed as L sampling points, and the calculation formula is L = frame length × sampling rate. Common frame shifts can be half or a quarter of the frame length, etc. According to the determined frame length and frame shift, the preprocessed voiceprint sample signal s(n) is segmented according to the frame length and the short-time energy is calculated; for the i-th frame signal s_i(n), the absolute value |s_i(n)| of each sampling point is calculated, and then the sum of the absolute values of all sampling points within the frame is added to obtain the amplitude sum of the frame signal. The amplitude sum is divided by the frame length to obtain the short-time average amplitude of the frame signal; according to the characteristics of the voiceprint signal and application requirements, a suitable zero-crossing threshold T is set. This threshold can be a small value close to zero, such as 0.01 - 0.1. Two counters C1 and C2 are initialized to record the number of times the signal value changes from positive to negative and the difference is greater than the threshold and the number of times the signal value changes from negative to positive and the difference is greater than the threshold respectively. The samples within the frame are traversed, and it is judged whether to increment counter C1 or C2 according to the set conditions. After traversing all sampling points of the i-th frame, the values of the two counters are added to obtain the short-time zero-crossing rate of the frame. A weight vector is constructed. For each frame, the short-time energy, short-time average amplitude, and short-time zero-crossing rate are weighted and fused according to the weight vector to obtain the fused time-domain eigenvalue;
[0080] Among them, the value of is 0.4, and the values of are both 0.3.
[0081] In one embodiment of the present invention, the extracting of frequency domain features, and the fusion and weighting operation of the energy mean, variance and ratio value to obtain the fused frequency domain feature value include the following steps:
[0082] After windowing each frame of the signal, perform a fast Fourier transform to calculate the amplitude and phase information of the signal at different frequencies. According to the result of the fast Fourier transform, square the amplitude spectrum of the frequency spectrum and then divide by the duration to obtain the power spectrum of each frame of the signal. Determine the frequency resolution according to the set sampling frequency and signal length, multiply the frequency resolution by the frequency index to obtain the corresponding actual frequency. Use the actual frequency as the abscissa and the corresponding power spectrum diagram as the ordinate to draw the power spectrum diagram, calculate the energy mean and variance of each frequency component in the power spectrum and find the maximum power spectrum value and the minimum power spectrum value in the power spectrum, subtract them from the mean respectively to obtain the corresponding difference values for division to get the ratio value, and take the absolute value;
[0083] Perform a fusion and weighting operation on the energy mean, variance and ratio value to obtain the fused frequency domain feature value;
[0084] Fused frequency domain feature value calculation formula:
[0085]
[0086]
[0087] Among them, represents the power spectrum mean, represents the number of frequency components, represents the th power spectrum value of the frequency component, represents the fused frequency domain feature value, represents the maximum power spectrum value, represents the minimum power spectrum value, , and represent the corresponding weight coefficients.
[0088] Specifically, the signal is divided into several frames. A window function is applied to each frame to reduce spectral leakage. The FFT is performed on each windowed frame signal to obtain its frequency-domain representation, including amplitude and phase information. The square of the amplitude spectrum of the signal is calculated and then divided by the duration to estimate the power spectrum of the signal. Given the sampling frequency and frame length, the frequency resolution is the sampling frequency divided by the frame length. The actual frequency of the k-th frequency component is k multiplied by the corresponding frequency resolution. The power spectrum is plotted with the actual frequency as the abscissa and the power spectrum value as the ordinate. The energy mean and variance of each frequency component in the power spectrum are calculated. The maximum and minimum power spectrum values in the power spectrum are found, and the differences obtained by subtracting them from the mean respectively are divided to get the ratio values. The absolute value of the overall ratio value is taken. The corresponding weights are set and a fusion and weighting operation is performed with the energy mean, variance, and ratio values to obtain the fusion frequency-domain feature value;
[0089] Among them, The value of is 0.5, The value of is 0.3, The value of is 0.2.
[0090] In one embodiment of the present invention, step S3 includes the following steps:
[0091] Collect the voiceprint sample data of the user, label each sample, and the labeling content includes the user's name and speech content. The collected sample data is input into a noise-reducing voiceprint model for noise reduction, then framed and window function processed, and the number of sampling points is unified and the amplitude is normalized. According to the selected fusion time-domain feature value and fusion frequency-domain feature value, feature extraction is performed on the preprocessed data. The extracted fusion time-domain feature value and fusion frequency-domain feature value are concatenated into a feature vector. A user table is created to store the relevant information of the user, and a voiceprint sample table is created to store the feature vector of the sample and the corresponding labeling information, which is associated with the user table through a foreign key. The recurrent neural network is selected as the framework for building the model. The labeled data is divided into training set data and input into the recurrent neural network model for training. The model learns and identifies the voiceprint features of different users. The trained model is used to predict the data in the test set part to obtain the voiceprint feature prediction results of different users. The feature vectors of all users and the corresponding prediction results are stored to form a voiceprint feature fusion model library.
[0092] Specifically, label the user name and speech content for each voiceprint sample. The labeling information should be accurate and error-free, and the labeling format should be unified. Input the collected voiceprint sample data into the noise reduction voiceprint model, whose purpose is to remove the noise components in the audio and improve the quality of the voiceprint data. Perform frame segmentation on the denoised voiceprint data, and divide the continuous audio signal into multiple short frames according to a certain frame length and frame shift. Perform window function processing on each frame of data to make the signal gradually decay to zero at both ends, avoiding spectral leakage. Unify the number of sampling points of all frames to a fixed value, which can be achieved by interpolation or truncation. Normalize the amplitude of the voiceprint data, and map the amplitude value of the signal to the interval of [-1, 1] or [0, 1] to avoid the influence of voiceprint data with different volumes on model training. According to the requirements of voiceprint recognition and the characteristics of the data, select appropriate time-domain feature values, including short-time energy, short-time average amplitude difference, and zero-crossing rate. These features can reflect the energy change, frequency change, and time-span information of the voiceprint signal in the time domain. At the same time, select frequency-domain feature values including mean energy, variance, and maximum-minimum ratio value, which can better capture the frequency characteristics of the voiceprint signal. For each preprocessed voiceprint frame, calculate the selected time-domain feature values and frequency-domain feature values to obtain the feature vector of this frame. Concatenate the fused time-domain feature value and fused frequency-domain feature value of the same frame in a certain order to form a feature vector, which is used as the final feature representation of this frame. Combine the feature vectors of all frames to form the feature matrix of the entire voiceprint sample for subsequent model input. Store the extracted features and labeling information according to the designed storage format. Divide a part of the labeled data as the training set, input the data of the training set into the recurrent neural network model, set appropriate training parameters, and use the trained model to predict the data in the test set part to obtain the voiceprint feature prediction results of different users. During the prediction process, the same preprocessing and feature extraction operations need to be performed on the input data, and then the feature vector is input into the model to obtain the prediction result output by the model. Store the feature vectors of all users and the corresponding prediction results to form the voiceprint feature fusion model library.
[0093] In one embodiment of the present invention, step S4 includes the following steps:
[0094] Store the extracted feature vector together with the unique identifier of the registered user into the voiceprint feature fusion model library. When it is necessary to verify the voiceprint sample, collect the voiceprint sample to be verified in the same way as in the registration stage and perform preprocessing operations. Obtain the feature vector of the voiceprint to be verified through the same feature extraction and calculation steps as in the registration stage for the preprocessed voiceprint sample to be verified. Use the cosine similarity method to calculate the similarity between the feature vector of the voiceprint to be verified and each registered voiceprint feature vector in the voiceprint feature fusion model library;
[0095] Determine the preset similarity threshold according to experience, and compare the calculated similarity with the preset similarity threshold:
[0096] If the similarity is less than the preset similarity threshold, it is determined that the voiceprint matching fails;
[0097] If the similarity is greater than or equal to the preset similarity threshold, it is determined that the voiceprint matching is successful.
[0098] Specifically, according to the feature extraction steps described above, process the voiceprint samples of each registered user, and finally obtain the feature vectors of each registered user. Generate a unique identifier for each registered user, such as a user ID. Store the feature vectors together with the unique identifier in the voiceprint feature fusion model library. At the same time, record other relevant information of the registered users in the model library, such as name, contact information, etc. When it is necessary to verify the voiceprint samples, use the same recording equipment and environmental conditions as in the registration stage to collect the voiceprint samples to be verified. Preprocess the collected voiceprint samples to be verified, including steps such as noise reduction of the voiceprint model, framing, window function processing, unifying the number of sampling points, and amplitude normalization. Extract features from the preprocessed voiceprint samples to be verified, and calculate according to the selected fusion time-domain feature values and fusion frequency-domain feature values in the registration stage to obtain the feature vectors of the voiceprint samples to be verified. Their dimensions and types should be consistent with the feature vectors in the registration stage for similarity calculation. Use the cosine similarity method to calculate the similarity between the feature vectors of the voiceprint samples to be verified and each registered voiceprint feature vector in the voiceprint feature fusion model library. Determine the preset similarity threshold according to experience. The selection of the threshold needs to consider factors such as the security requirements of the system, the false recognition rate, and the rejection rate. Usually, a suitable value is determined through experiments and data analysis. Compare each calculated similarity value with the preset similarity threshold. If all similarity values are less than the preset similarity threshold, it is determined that the voiceprint matching fails, that is, the voiceprint to be verified does not match any user in the registered voiceprint library. If at least one similarity value is greater than or equal to the preset similarity threshold, it is determined that the voiceprint matching is successful, and the registered user with the maximum similarity value is output as the matching result.
[0099] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A voiceprint matching method, characterized in that, The method includes the following steps: S1: Collect the voiceprint samples to be matched through an audio acquisition device, including registered voiceprint samples and voiceprint samples to be verified. Preprocess and label the collected original noisy voiceprint samples. After constructing a generator and a discriminator, form a training data pair with the original noisy voiceprint samples and the corresponding generated voiceprint samples, construct a voiceprint noise reduction model, and perform preprocessing operations on the collected voiceprint samples; S2: Extract the short-time energy, short-time average amplitude, and short-time zero-crossing rate from the preprocessed voiceprint samples. Construct a weight vector and fuse it with the time-domain features to obtain a fused time-domain feature value. Extract the frequency-domain features, and perform a fused weighting operation on the energy mean, variance, and ratio value to obtain a fused frequency-domain feature value; S3: Input the obtained fused time-domain feature value and fused frequency-domain feature value into a recurrent neural network model to form a voiceprint feature fusion model library; S4: Calculate the similarity between the voiceprint feature vector to be verified and each registered voiceprint feature vector in the voiceprint feature fusion model library, and determine whether the similarity is less than a preset similarity threshold; The voiceprint noise reduction model includes the following steps: Collect voiceprint samples of different users in different noise environments, and perform preprocessing operations on the voiceprint data, including sampling rate conversion, silence removal, and pre-emphasis, and convert it into a form suitable for model input. Label the collected original noisy voiceprint data, and label the noise type, noise level, speech content, and user identity it contains. For the voiceprint generated by the generator, label its source and the corresponding original noisy voiceprint identifier; Design a shared input layer in the generator and the discriminator, and perform the same preprocessing operations on the input voiceprint, including Fourier transform and Mel-frequency cepstral coefficient extraction. After converting the voiceprint signal into a unified feature representation, use the features extracted from the noisy voiceprint signal as the input to construct the generator. The discriminator adopts a dual-input structure, one input is the feature representation of the original noisy voiceprint, and the other input is the feature representation of the generated voiceprint, and a comparison module is set inside the discriminator to perform comparative analysis on the original noisy voiceprint feature and the generated voiceprint feature; Form a training data pair with the original noisy voiceprint samples and the corresponding generated voiceprint samples, and at the same time, input them into the discriminator. Use the training set to train the model, and update the parameters of the model through multiple iterations to enable the model to learn the mapping relationship from the noisy voiceprint to the clean voiceprint, and calculate the root mean square error between the original voiceprint signal and the denoised voiceprint signal; The step S3 includes the following steps: Collect the voiceprint sample data of users, label each sample, and the labeling content includes the user's name and speech content. Input the collected sample data into the noise reduction voiceprint model for noise reduction, and then perform frame segmentation and window function processing. Unify the number of sampling points and normalize the amplitude. According to the selected fusion time-domain feature values and fusion frequency-domain feature values, extract features from the preprocessed data. Concatenate the extracted fusion time-domain feature values and fusion frequency-domain feature values into a feature vector. Create a user table to store the relevant information of users, and create a voiceprint sample table to store the feature vectors of samples and the corresponding labeling information, which is associated with the user table through a foreign key. Select a recurrent neural network as the framework for building the model, divide the labeled data into training set data and input it into the recurrent neural network model for training. The model learns and recognizes the voiceprint features of different users. Use the trained model to predict the data in the test set part to obtain the voiceprint feature prediction results of different users. Store the feature vectors of all users and the corresponding prediction results to form a voiceprint feature fusion model library; The extraction of frequency-domain features, and the fusion weighted operation of the energy mean, variance and ratio value to obtain the fusion frequency-domain feature value, includes the following steps: After windowing each frame of the signal, perform a fast Fourier transform to calculate the amplitude and phase information of the signal at different frequencies. According to the result of the fast Fourier transform, obtain the power spectrum of each frame of the signal by squaring the amplitude spectrum of the frequency spectrum and then dividing by the duration. Determine the frequency resolution based on the set sampling frequency and signal length, multiply the frequency resolution by the frequency index to obtain the corresponding actual frequency. Use the actual frequency as the abscissa and the corresponding power spectrum diagram as the ordinate to plot the power spectrum diagram, calculate the energy mean and variance of each frequency component in the power spectrum diagram, and in the power spectrum find the maximum power spectrum value and the minimum power spectrum value, subtract them from the mean respectively to obtain the corresponding differences, divide them to get the ratio value, and take the absolute value; Perform a fusion weighted operation on the energy mean, variance and ratio value to obtain the fusion frequency-domain feature value; Fusion frequency-domain feature value calculation formula: ; ; Among them, represents the mean power spectrum, represents the number of frequency components, represents the th power spectrum value of the frequency component, represents the fused frequency domain eigenvalue, represents the maximum power spectrum value, represents the minimum power spectrum value, , and represent the corresponding weight coefficients.
2. The voiceprint matching method according to claim 1, wherein The step S1 includes the following steps: Under the conditions of different noise environments, collect the voiceprint samples of users during registration through an audio acquisition device, convert the collected voiceprint samples into a unified audio format, and input the registered voiceprint samples into the voiceprint noise reduction model for preprocessing; use the same audio acquisition device, and under the conditions similar to the registration voiceprint sample acquisition environment, collect the voiceprint samples to be verified, convert them into a unified audio format, and at the same time, perform corresponding preprocessing operations on the voiceprint samples to be verified.
3. A voiceprint matching method according to claim 1, characterized in that The calculation of the root mean square error between the original voiceprint signal and the noise-reduced voiceprint signal includes the following steps: The original voiceprint signal is denoted as , and the voiceprint signal after noise reduction is denoted as , for each sampling point, calculate the square of the difference between the original voiceprint signal and the voiceprint signal after noise reduction, add up all the squared differences, divide the sum of squared errors by the total number of all sampling points, and take the square root of the mean squared error to obtain the root mean square error; Root mean square error calculation formula: ; Among them, represents the root mean square error, represents the total number of samples, represents the th original voiceprint signal, represents the th noise-reduced voiceprint signal; Based on the experience of historical voiceprint noise reduction projects, set the RMSE empirical threshold: If is less than the RMSE empirical threshold, it indicates that the difference between the original voiceprint signal and the noise-reduced voiceprint signal is small, and the noise reduction effect is good. The trained model is deployed in the system to perform noise reduction on the collected voiceprint samples; If is greater than or equal to the RMSE empirical threshold, it indicates that the difference between the original voiceprint signal and the denoised voiceprint signal is large, the denoising effect is poor, and the model needs to be adjusted.
4. A voiceprint matching method according to claim 1, characterized in that, The extraction of short-time energy, short-time average amplitude and short-time zero-crossing rate, and the construction of a weight vector to fuse with the time-domain features to obtain the fusion time-domain feature value, includes the following steps: Extract features from the preprocessed voiceprint samples. The time-domain features include short-time energy, short-time average amplitude, and short-time zero-crossing rate. Determine the frame length and frame shift, segment the voiceprint signal according to the determined frame length and frame shift, and calculate the corresponding short-time energy for the short-time energy of this frame; for each frame of signal, calculate the absolute value of each sampling point therein, add up the absolute values of all sampling points in this frame to obtain the amplitude sum of this frame of signal, and then divide by the frame length to obtain the short-time average amplitude of this frame of signal; set the zero-crossing threshold, initialize two timers, record the number of times the signal value changes from positive to negative and the difference is greater than the set threshold and the number of times the signal value changes from negative to positive and the difference is greater than the preset threshold respectively, traverse the samples within the frame, and after the traversal ends, add the values of the two counters to obtain the short-time zero-crossing rate of this frame; Determine the weight coefficients of short-time energy, short-time average amplitude, and short-time zero-crossing rate according to application requirements and data characteristics, and construct a weight vector , for each frame, weight and fuse the short-time energy, short-time average amplitude, and short-time zero-crossing rate according to the weight vector to obtain the fused time-domain feature value; Fusion time-domain feature value calculation formula: ; Among them, represents the fused time-domain eigenvalue of the th frame, represents the short-time energy of the th frame, represents the short-time average amplitude of the th frame, represents the short-time zero-crossing rate of the th frame, , and represent the corresponding weight coefficients; After performing a weighted fusion operation on all frames, a sequence of fused time-domain eigenvalue is obtained , where represents the number of frames after frame segmentation.
5. A voiceprint matching method according to claim 1, characterized in that, Step S4 described above includes the following steps: Store the extracted feature vector together with the unique identifier of the registered user in the voiceprint feature fusion model library. When it is necessary to verify the voiceprint sample, collect the voiceprint sample to be verified in the same way as in the registration stage and perform preprocessing operations. Obtain the feature vector of the voiceprint to be verified through the same feature extraction and calculation steps as in the registration stage for the preprocessed voiceprint sample to be verified, and use the cosine similarity method to calculate the similarity between the feature vector of the voiceprint to be verified and each registered voiceprint feature vector in the voiceprint feature fusion model library; Determine the preset similarity threshold according to experience, and compare the calculated similarity with the preset similarity threshold: If the similarity is less than the preset similarity threshold, it is determined that the voiceprint matching fails; If the similarity is greater than or equal to the preset similarity threshold, it is determined that the voiceprint matching is successful.
Citation Information
Patent Citations
Human voice recognition system
CN113077794A
Feature-based intelligent voiceprint recognition method and system
CN119296568A
Human voice enhancement and noise reduction method based on generative adversarial network
CN119360872A