Self-adaptive VAD parameter adjusting method and system based on voiceprint recognition
By introducing voiceprint recognition technology into the VAD algorithm and adjusting VAD parameters in real time, the problem of insufficient adaptability of VAD algorithms in the existing technology under complex environments and individual differences is solved, and the accuracy and robustness of speech detection are significantly improved.
Patent Information
- Application Number
- CN202510511744.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing VAD algorithm lacks adaptability under complex environments and individual differences, resulting in a decrease in the accuracy of speech detection and the problems of false detection and missed detection.
Adaptive VAD parameter adjustment method based on voiceprint recognition is adopted to determine whether the user is a registered user through pre-processing of audio signals and voiceprint feature extraction, and the VAD parameters are adjusted in real time according to the judgment results, including zero-crossing rate detection threshold, signal-to-noise ratio threshold, suspend time and duration, etc.
It realizes personalized VAD parameters and real-time dynamic adjustment, improves the sensitivity and accuracy of voice detection in complex environments, reduces the risk of misdetection and missed detection caused by noise interference, and ensures the integrity and coherence of voice detection.
Smart Images

Figure CN120048268A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to an adaptive VAD parameter adjustment method and system based on voiceprint recognition. Background Art
[0002] VAD (Voice Activity Detection) is a technology used to distinguish speech signals from non-speech signals (such as silence and noise). Traditional VAD algorithms usually use fixed energy thresholds or spectral features for speech detection. In complex scenarios, fixed parameters and lack of adaptability can easily lead to a decrease in speech detection accuracy, and there are problems with false detection and missed detection. Specifically, in multi-user scenarios, there are significant differences in the speech characteristics of different users, such as volume, speaking speed, frequency distribution, etc. If the VAD parameters are fixed, they cannot be adaptively adjusted according to individual characteristics, resulting in difficulty in accurately detecting the speech of some users, missed detection or unstable detection.
[0003] Existing VAD parameter adjustment methods are often based on preset models or static rules, lacking dynamic adaptive capabilities, and cannot automatically adjust parameters when the environment or user changes, affecting detection accuracy. In addition, in voice calls or speech recognition applications, if the VAD parameters are not suitable for the current environment or user characteristics, it may cause voice fragments to be lost or introduce unnecessary background noise, thereby affecting system performance.
[0004] In summary, the existing VAD algorithm lacks adaptability when facing complex environments and individual differences, resulting in a decrease in the accuracy of speech detection and the existence of problems of false detection and missed detection. Summary of the invention
[0005] In order to solve the above problems, the present invention proposes an adaptive VAD parameter adjustment method and system based on voiceprint recognition.
[0006] On the one hand, the present invention proposes an adaptive VAD parameter adjustment method based on voiceprint recognition, comprising: Acquire audio signals through audio acquisition equipment, pre-process the audio signals and extract voiceprint features, map the voiceprint features to the user voiceprint model, compare the current user voiceprint model with the stored model in the voiceprint library, and determine whether the user is a registered user; Determine the VAD parameter according to whether the user is a registered user; if the user is determined to be an unregistered user, adjust the VAD parameter in real time by detecting characteristic fluctuations of the silent segment and the speech segment; The adjusted VAD parameters are stored in the voiceprint library.
[0007] Furthermore, when the audio signal is acquired through the audio acquisition device, it includes: The audio signal is sampled at the same sampling rate, converted to mono, and the format is PCM format.
[0008] Furthermore, when the audio signal is preprocessed, it includes: S1: Calculate the average value of the audio signal and subtract the average value of the audio signal from each sampling point; S2: Pre-emphasis is performed according to the following relationship: ; S3: Divide the audio signal into frames of 25 ms and set the frame shift to 10 ms; S4: Windowing the audio signal based on the Hamming window, multiplying each sampling point by the Hamming window function; in, and are the audio signals of the nth sampling point before and after pre-emphasis processing, respectively. is the pre-emphasis coefficient, 0.9< ≤1, It is the audio signal of the n-1th sampling point before pre-emphasis processing.
[0009] Furthermore, when extracting voiceprint features from an audio signal, the voiceprint features include MFCC values, LPCC values, and fundamental frequencies.
[0010] Furthermore, the current user's voiceprint model is compared with the stored model in the voiceprint library to determine whether the user is a registered user, including: Obtain the MFCC value, LPCC value and fundamental frequency of the current user voiceprint model and each stored model in the voiceprint library and compare them, calculate the similarity, select the maximum similarity, and compare the maximum similarity with the similarity threshold; when the maximum similarity is greater than or equal to the similarity threshold, determine that the current user model is a registered user; when the maximum similarity is less than the similarity threshold, determine that the current user model is not a registered user; the similarity satisfies the following relationship: ; in, is the similarity, is the jth voiceprint feature of the current user’s voiceprint model, is the jth voiceprint feature of the storage model; is the number of voiceprint features, =3, when j=1, it corresponds to MFCC value; when j=2, it corresponds to LPCC value; when j=3, it corresponds to baseband.
[0011] Furthermore, when it is determined that the user is a registered user, the VAD parameters of the storage model corresponding to the maximum similarity in the voiceprint library are retrieved; When it is determined that the user is not a registered user, the preprocessed audio signal is divided into a silent segment and a speech segment according to the energy value, and the VAD parameter is adjusted in real time according to the characteristic fluctuation detection of the silent segment and the speech segment.
[0012] Furthermore, when it is determined that the user is not a registered user, when the user is divided into a silent segment and a voice segment according to the energy value, it includes: The energy value of each frame of audio signal is calculated, and the energy value is compared with the energy threshold to determine whether it is a silent segment or a speech segment. The energy value is obtained through the following relationship: ; ; in, is the energy value of the mth frame, is the kth sampling point of the mth frame, N is the number of sampling points per frame; N ≥ 2, is the energy threshold; when Greater than or equal to It is judged as a speech segment, otherwise it is a silent segment.
[0013] Furthermore, the VAD parameters are adjusted in real time according to the characteristic fluctuation detection of the silent segment, including: The VAD parameters include a zero-crossing rate detection threshold, a signal-to-noise ratio threshold, a hang-up time, and a duration; the zero-crossing rate detection threshold, the signal-to-noise ratio threshold, the hang-up time, and the duration are adjusted by the following relationship: ; ; ; ; ; ; ; in, is the adjusted zero-crossing rate detection threshold, is the initial detection threshold, is the spectrum volatility, is the h-th dimension MFCC value of the m-th frame, is the h-th dimension MFCC value of the m-1-th frame, K represents the dimension of the MFCC value, K is 13 or 20, is a constant, =10 -6 , is the signal-to-noise ratio, is the average energy value of all speech segments, is the average energy value of all silent segments, is the adjusted signal-to-noise ratio threshold, is the initial signal-to-noise ratio threshold, is the adjusted suspension time, is the initial suspension time, is the adjusted duration, is the initial duration, is the short-term energy fluctuation rate of the m-th frame, is the short-term energy fluctuation rate of the m-1th frame, is the spectrum volatility adjustment coefficient, is the spectrum stability adjustment coefficient, is the first signal-to-noise ratio adjustment coefficient, is the second signal-to-noise ratio adjustment coefficient, is the energy fluctuation influence coefficient, is the spectrum fluctuation influence coefficient, , , and The value range of is [0.05, 0.2].
[0014] Furthermore, it also includes: recording the time length of time the storage model is stored in the voiceprint library, and when the time length is greater than or equal to a preset maximum time length, deleting the storage model.
[0015] On the other hand, the present invention also proposes an adaptive VAD parameter adjustment system based on voiceprint recognition, comprising: The collection and judgment module is configured to obtain an audio signal through an audio collection device, pre-process the audio signal and extract voiceprint features, map the voiceprint features to a user voiceprint model, compare the current user voiceprint model with a stored model in a voiceprint library, and determine whether the user is a registered user; The adjustment processing module is configured to determine the VAD parameter according to whether the user is a registered user; when the user is determined to be an unregistered user, the VAD parameter is adjusted in real time by detecting characteristic fluctuations of the silent segment and the speech segment; The storage management module is provided with a voiceprint library, and the storage management module is configured to store the adjusted VAD parameters in the voiceprint library.
[0016] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides an adaptive VAD parameter adjustment method and system based on voiceprint recognition. By introducing voiceprint recognition technology, the personalized and real-time dynamic adjustment of VAD parameters is realized, and the problem of fixed parameters and lack of adaptability of VAD algorithm in the prior art is effectively overcome. When detecting audio signals, the present invention first extracts and compares the user's voiceprint to determine whether it is a registered user, so as to load the historically optimized VAD parameters in the registered user scenario to ensure the stability and accuracy of the detection. In the unregistered user scenario, the VAD parameters are adjusted in real time through the characteristic fluctuation detection of the silent segment and the voice segment, including the zero-crossing rate threshold, signal-to-noise ratio threshold, suspension time and duration, etc., which significantly improves the sensitivity and accuracy of VAD detection in complex environments.
[0017] The present invention introduces a signal-to-noise ratio adaptive adjustment mechanism, which enables the system to automatically optimize VAD parameters under different signal-to-noise ratio environments, effectively reducing the risk of false detection and missed detection caused by noise interference. At the same time, a comprehensive detection method of energy short-term volatility and spectrum volatility is adopted to enable VAD parameters to maintain a dynamic and smooth transition between silence and speech segments, reducing the loss or truncation of speech segments and improving the integrity and continuity of speech detection. In addition, the present invention also has a model maintenance mechanism, which can prevent the redundant expansion of the voiceprint library by recording the length of the storage model and regularly cleaning up expired models, thereby maintaining the efficient operation of the system and the validity of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings: Figure 1 A flowchart of a method for adaptive VAD parameter adjustment based on voiceprint recognition provided by an embodiment of the present invention.
[0019] Figure 2 A functional framework diagram of an adaptive VAD parameter adjustment system based on voiceprint recognition provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The exemplary embodiments disclosed in the present application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features described in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0021] See also Figure 1 As shown, an embodiment of the present invention provides an adaptive VAD parameter adjustment method based on voiceprint recognition, comprising: Acquire audio signals through audio acquisition equipment, pre-process the audio signals and extract voiceprint features, map the voiceprint features to the user voiceprint model, compare the current user voiceprint model with the stored model in the voiceprint library, and determine whether the user is a registered user; Determine the VAD parameters based on whether the user is a registered user; if the user is determined to be an unregistered user, adjust the VAD parameters in real time by detecting the characteristic fluctuations of the silent segment and the speech segment; The adjusted VAD parameters are stored in the voiceprint library.
[0022] It should be noted that the above-mentioned adaptive VAD parameter adjustment method based on voiceprint recognition realizes adaptive parameter optimization for individual characteristics of different users by combining voiceprint recognition with VAD parameter adjustment, and effectively overcomes the problem of fixed parameters and decreased detection accuracy of existing VAD algorithms in complex environments. First, this method compares the user model with the stored model in the voiceprint library through voiceprint recognition technology, can automatically identify registered users, and load historically optimized VAD parameters, so as to maintain the stability and accuracy of detection parameters in the registered user scenario and reduce false detection or missed detection due to individual differences.
[0023] Secondly, in the scenario of non-registered users, this method adopts the characteristic fluctuation detection mechanism of silent segments and speech segments, which can dynamically adjust the VAD parameters in real time. By detecting the characteristic changes of short-term energy fluctuation rate and spectrum fluctuation rate of silent and speech segments, the zero-crossing rate detection threshold, signal-to-noise ratio threshold, suspension time and duration parameters are automatically optimized to effectively adapt to user characteristics and environmental changes. This mechanism improves the sensitivity and adaptability of speech detection in complex scenarios, significantly reduces the interference of background noise on VAD detection, reduces the occurrence of false detection and missed detection, and ensures the accuracy and integrity of speech detection.
[0024] In addition, the adjusted VAD parameters are stored in the voiceprint library, which realizes the dynamic update of parameters and self-optimization of the model, so that the system can continuously optimize the VAD parameters during multiple uses, and improve the stability and adaptability of long-term operation. Overall, this method can effectively improve the accuracy, robustness and real-time performance of voice detection in complex environments and multi-user scenarios, and is widely used in speech recognition, voice calls, intelligent voice interaction and other fields.
[0025] In some embodiments of the present application, when an audio signal is acquired through an audio acquisition device, the process includes: The audio signal is sampled at the same sampling rate, converted to mono, and the format is PCM format.
[0026] In this embodiment, the sampling rates are uniformly 16 kHz or 48 kHz.
[0027] It should be noted that by unifying the sampling rate and audio format processing, the consistency and high-quality acquisition of audio signals are ensured, thereby improving the accuracy and robustness of subsequent voiceprint recognition and VAD parameter adjustment. In this embodiment, the sampling rate is unified to 16kHz or 48kHz, which can effectively balance the fidelity and processing efficiency of the audio signal. The 16kHz sampling rate is commonly used in speech recognition and voice calls in the field of speech processing, and can meet the feature extraction requirements of most speech signals, while reducing the amount of data and improving real-time processing performance; the 48kHz sampling rate is suitable for high-fidelity speech or audio processing, can capture richer spectral information, and is suitable for high-quality speech or audio analysis.
[0028] Converting audio signals to monophonic processing can simplify the audio data structure, avoid redundant calculations or inconsistent features when adjusting VAD parameters for dual-channel or multi-channel audio, and improve processing efficiency and model adaptability. At the same time, using PCM (pulse code modulation) format to store audio signals can ensure lossless compression of audio signals, maintain the accuracy and integrity of the original signal, avoid loss or distortion of speech features caused by lossy compression, and thus improve the accuracy of subsequent voiceprint feature extraction and model comparison.
[0029] In some embodiments of the present application, when preprocessing the audio signal, it includes: S1: Calculate the average value of the audio signal and subtract the average value of the audio signal from each sampling point; S2: Pre-emphasis is performed according to the following relationship: ; S3: Divide the audio signal into frames of 25 ms and set the frame shift to 10 ms; S4: Windowing the audio signal based on the Hamming window, multiplying each sampling point by the Hamming window function; in, and are the audio signals of the nth sampling point before and after pre-emphasis processing, respectively. is the pre-emphasis coefficient, 0.9< ≤1, It is the audio signal of the n-1th sampling point before pre-emphasis processing.
[0030] It should be noted that by calculating the average value of the audio signal in the preprocessing stage and subtracting the average value from each sampling point, the DC component and signal offset can be effectively removed, ensuring that the audio signal is in a zero-mean state for subsequent feature extraction and analysis, thereby improving the stability and consistency of the features. The pre-emphasis processing step uses a first-order high-pass filter to enhance high-frequency components and suppress low-frequency components, thereby improving the clarity and resolution of the speech signal. By setting the range of the pre-emphasis coefficient to 0.9≤α≤1, the noise caused by excessive filtering can be suppressed while ensuring the high-frequency enhancement effect, thereby maintaining the signal quality.
[0031] Dividing the audio signal into frames of 25ms and setting the frame shift to 10ms helps to balance the time resolution and frequency resolution, ensuring that the speech signal can be fully captured in the time and frequency domains, while reducing the signal overlap between adjacent frames and improving the accuracy of voice activity detection.
[0032] Using the Hamming window to perform windowing on the audio signal can effectively suppress spectrum leakage and reduce spectrum distortion caused by truncation. The Hamming window function has a smaller amplitude at the edge of the spectrum, thereby smoothing the spectrum in the frequency domain and improving the frequency resolution and accuracy of spectrum feature extraction. This preprocessing method can significantly improve the accuracy and robustness of subsequent voiceprint feature extraction and VAD parameter adjustment.
[0033] In some embodiments of the present application, when extracting voiceprint features from an audio signal, the voiceprint features include MFCC values, LPCC values, and fundamental frequencies.
[0034] In some embodiments of the present application, when comparing the current user's voiceprint model with the stored model in the voiceprint library to determine whether the user is a registered user, the following steps are included: Obtain the MFCC value, LPCC value and fundamental frequency of the current user's voiceprint model and each stored model in the voiceprint library and compare them, calculate the similarity, select the maximum similarity, and compare the maximum similarity with the similarity threshold; when the maximum similarity is greater than or equal to the similarity threshold, the current user model is judged to be a registered user, and when the maximum similarity is less than the similarity threshold, the current user model is judged not to be a registered user; the similarity satisfies the following relationship: ; in, is the similarity, is the jth voiceprint feature of the current user’s voiceprint model, is the jth voiceprint feature of the storage model; is the number of voiceprint features, =3, when j=1, it corresponds to MFCC value; when j=2, it corresponds to LPCC value; when j=3, it corresponds to baseband.
[0035] It should be noted that by comparing the current user's voiceprint model with the stored model in the voiceprint library, the identity of the registered user can be effectively recognized. Using MFCC, LPCC and fundamental frequency voiceprint features for comparison helps to comprehensively characterize the user's voice feature information. Among them, MFCC can capture the short-time spectrum characteristics of the voice signal, LPCC can reflect the linear prediction characteristics of the voice signal, and fundamental frequency can describe the pitch characteristics of the voice. The joint comparison of the three features can enhance the system's ability to distinguish the individual characteristics of the voice, thereby improving the accuracy and robustness of the comparison.
[0036] The similarity calculation uses the Euclidean distance formula or cosine similarity calculation method, which can quantify the difference between the current user's voiceprint model and the stored model in the feature space. By selecting the maximum similarity and comparing it with the preset similarity threshold, it can ensure that the system identifies the user as a registered user when the similarity between the user model and the stored model is high, and identifies the user as an unregistered user when the similarity is low. This threshold-based discrimination method effectively reduces the false recognition rate and missed recognition rate, ensuring the recognition accuracy and stability of the system.
[0037] In addition, multi-dimensional feature comparison can enhance anti-interference capabilities in complex environments. Even if there is a certain amount of noise or voice changes, model comparison can still make robust judgments based on comprehensive features, thereby improving the adaptability and reliability of the system in multiple scenarios.
[0038] The method for constructing the voiceprint model of the current user is as follows: Use the extracted voiceprint features to construct the voiceprint model of the current user. This model can be a multidimensional feature vector or a trained machine learning model (such as Gaussian mixture model GMM, support vector machine SVM or deep neural network DNN, etc.). Obtain the voiceprint model of the registered user from the voiceprint library. The voiceprint library stores the voiceprint models of all registered users, and each user's model contains representative information of its voice features. The construction of the model belongs to the existing technology, so it will not be repeated here.
[0039] In some embodiments of the present application, when it is determined that the user is a registered user, the VAD parameters of the storage model corresponding to the maximum similarity in the voiceprint library are retrieved; When it is determined that the user is not a registered user, the preprocessed audio signal is divided into a silent segment and a speech segment according to the energy value, and the VAD parameters are adjusted in real time according to the characteristic fluctuation detection of the silent segment and the speech segment.
[0040] In some embodiments of the present application, when a user is determined to be a non-registered user, when the user is divided into a silent segment and a voice segment according to the energy value, the following is included: Calculate the energy value of each frame of audio signal and determine whether it is a silent segment or a speech segment by comparing the energy value with the energy threshold. The energy value is obtained through the following relationship: ; ; in, is the energy value of the mth frame, is the kth sampling point of the mth frame, N is the number of sampling points per frame; N ≥ 2, is the energy threshold; when Greater than or equal to It is judged as a speech segment, otherwise it is a silent segment.
[0041] It should be noted that by separating the silent segment from the speech segment, only the speech part can be focused on in the speech recognition model, and useless silent or noise data can be ignored. This helps to reduce the interference of background noise and improve the accuracy of speech recognition. By removing the silent segment in advance, the consumption of computing resources for the non-speech signal is reduced, thereby improving the processing efficiency of the system. Especially in the real-time speech recognition system, reducing the silent segment can significantly reduce the delay. Dynamic judgment based on the real-time energy value and energy threshold can adapt to the changes in speech intensity in different environments, such as the speaker's speaking speed, the volume of the speech, etc., to avoid the speech signal being mistakenly divided into a silent segment due to being too weak. The energy threshold can be adjusted according to different application scenarios (such as noisy environment, speaker group, etc.), so that the method has good flexibility and adaptability, and can adapt to a variety of sound environments. If the calculation method and threshold of the energy value are properly set, the situation of misjudging the silent segment as a speech segment can be avoided, and the model's recognition ability for non-registered users or abnormal users can be further improved. This method can better handle the changes of various speech signals, enhance the system's ability to distinguish between silent and speech signals, and maintain a relatively stable recognition performance in different noise environments.
[0042] In some embodiments of the present application, VAD parameters are adjusted in real time according to characteristic fluctuation detection of the silent segment, including: VAD parameters include zero-crossing rate detection threshold, signal-to-noise ratio threshold, suspension time and duration; zero-crossing rate detection threshold, signal-to-noise ratio threshold, suspension time and duration are adjusted through the following relationship: ; ; ; ; ; ; ; in, is the adjusted zero-crossing rate detection threshold, is the initial detection threshold, is the spectrum volatility, is the h-th dimension MFCC value of the m-th frame, is the h-th dimension MFCC value of the m-1-th frame, K represents the dimension of the MFCC value, K is 13 or 20, is a constant, =10 -6 , is the signal-to-noise ratio, is the average energy value of all speech segments, is the average energy value of all silent segments, is the adjusted signal-to-noise ratio threshold, is the initial signal-to-noise ratio threshold, is the adjusted suspension time, is the initial suspension time, is the adjusted duration, is the initial duration, is the short-term energy fluctuation rate of the m-th frame, is the short-term energy fluctuation rate of the m-1th frame, is the spectrum volatility adjustment coefficient, is the spectrum stability adjustment coefficient, is the first signal-to-noise ratio adjustment coefficient, is the second signal-to-noise ratio adjustment coefficient, is the energy fluctuation influence coefficient, is the spectrum fluctuation influence coefficient, , , and The value range of is [0.05, 0.2].
[0043] It should be noted that the zero-crossing rate detection threshold is defined as follows: the zero-crossing rate is the frequency at which the signal waveform crosses the zero axis, and is usually used to measure the "variability" of the signal. In speech signals, the high or low zero-crossing rate can reflect the activity of the speech. Adjustment method: By combining the MFCC features of the current frame with the MFCC features of the previous frame, the changes in the MFCC features are calculated, which helps to adjust the zero-crossing rate detection threshold in real time. In this way, changes in speech and different noise environments can be dynamically responded to.
[0044] Definition of signal-to-noise ratio threshold: The signal-to-noise ratio (SNR) reflects the ratio of useful components to noise components in a signal. A higher SNR usually means better quality of speech signals and lower noise. Adjustment method: By calculating the average energy of speech segments and silent segments, the signal-to-noise ratio threshold is adjusted based on this information to improve the accuracy of signal recognition in noisy environments. In different noise environments, the SNR threshold needs to be adjusted dynamically.
[0045] Hang time definition: Hang time refers to the time the system needs to maintain the "voice active state" after detecting a voice segment to avoid misjudgment due to short silence. Adjustment method: Dynamically adjust the hang time according to the energy fluctuation rate and spectrum fluctuation. If the voice signal fluctuates greatly, the hang time can be appropriately extended to ensure the continuity of voice detection.
[0046] Duration refers to the length of time that speech activity is considered valid, and is usually used to prevent short, non-speech disturbances from being misjudged as speech. Adjustment method: Adjustment is made through characteristics such as energy short-term volatility, spectrum volatility, and energy and spectrum stability. In particular, the energy fluctuation influence coefficient and spectrum fluctuation influence coefficient can help determine the duration adjustment under different fluctuation conditions.
[0047] In some embodiments of the present application, the method further includes: recording the time length of time the stored model is stored in the voiceprint database, and deleting the stored model when the time length is greater than or equal to a preset maximum time length. In addition, if the current user's voiceprint model is compared with the stored model in the voiceprint database and it is determined that the user is not a registered user, the current user's voiceprint model is stored in the voiceprint database.
[0048] See also Figure 2 As shown, an embodiment of the present invention provides an adaptive VAD parameter adjustment system based on voiceprint recognition, including: The collection and judgment module is configured to obtain audio signals through an audio collection device, pre-process the audio signals and extract voiceprint features, map the voiceprint features to a user voiceprint model, compare the current user voiceprint model with a stored model in a voiceprint library, and determine whether the user is a registered user; The adjustment processing module is configured to determine the VAD parameters according to whether the user is a registered user; when the user is determined to be an unregistered user, the VAD parameters are adjusted in real time by detecting characteristic fluctuations of the silent segment and the speech segment; The storage management module is provided with a voiceprint library, and the storage management module is configured to store the adjusted VAD parameters in the voiceprint library.
[0049] In summary, the present invention provides an adaptive VAD parameter adjustment method and system based on voiceprint recognition. By introducing voiceprint recognition technology, the personalization and real-time dynamic adjustment of VAD parameters are realized, effectively overcoming the problem of fixed parameters and lack of adaptability of VAD algorithm in the prior art. When detecting audio signals, the present invention first extracts and compares the user's voiceprint to determine whether it is a registered user, thereby loading the historically optimized VAD parameters in the registered user scenario to ensure the stability and accuracy of the detection. In the unregistered user scenario, the VAD parameters are adjusted in real time through the characteristic fluctuation detection of the silent segment and the voice segment, including the zero-crossing rate threshold, signal-to-noise ratio threshold, suspension time and duration, etc., which significantly improves the sensitivity and accuracy of VAD detection in complex environments.
[0050] The present invention introduces a signal-to-noise ratio adaptive adjustment mechanism, which enables the system to automatically optimize VAD parameters under different signal-to-noise ratio environments, effectively reducing the risk of false detection and missed detection caused by noise interference. At the same time, a comprehensive detection method of energy short-term volatility and spectrum volatility is adopted to enable VAD parameters to maintain a dynamic and smooth transition between silence and speech segments, reducing the loss or truncation of speech segments and improving the integrity and continuity of speech detection. In addition, the present invention also has a model maintenance mechanism, which can prevent the redundant expansion of the voiceprint library by recording the length of the storage model and regularly cleaning up expired models, thereby maintaining the efficient operation of the system and the validity of the data.
[0051] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. An adaptive VAD parameter adjustment method based on voiceprint recognition, characterized in that: include: Acquire audio signals through audio acquisition equipment, pre-process the audio signals and extract voiceprint features, map the voiceprint features to the user voiceprint model, compare the current user voiceprint model with the stored model in the voiceprint library, and determine whether the user is a registered user; Determine the VAD parameter according to whether the user is a registered user, and if the user is determined to be an unregistered user, adjust the VAD parameter in real time by detecting characteristic fluctuations of the silent segment and the speech segment; The adjusted VAD parameters are stored in the voiceprint library.
2. The method for adaptive VAD parameter adjustment based on voiceprint recognition according to claim 1, characterized in that: When obtaining audio signals through an audio acquisition device, it includes: The audio signal is sampled at the same sampling rate, converted to mono, and the format is PCM format.
3. The method for adaptive VAD parameter adjustment based on voiceprint recognition according to claim 1, characterized in that: When preprocessing the audio signal, including: S1: Calculate the average value of the audio signal and subtract the average value of the audio signal from each sampling point; S2: Pre-emphasis is performed according to the following relationship: ; S3: Divide the audio signal into frames of 25 ms and set the frame shift to 10 ms; S4: Windowing the audio signal based on the Hamming window, multiplying each sampling point by the Hamming window function; in, and are the audio signals of the nth sampling point before and after pre-emphasis processing, respectively. is the pre-emphasis coefficient, 0.9< ≤1, It is the audio signal of the n-1th sampling point before pre-emphasis processing.
4. The method for adaptive VAD parameter adjustment based on voiceprint recognition according to claim 1, characterized in that: When extracting voiceprint features from an audio signal, the voiceprint features include MFCC values, LPCC values, and fundamental frequencies.
5. The method for adaptive VAD parameter adjustment based on voiceprint recognition according to claim 4, characterized in that: When comparing the current user's voiceprint model with the stored model in the voiceprint library to determine whether the user is a registered user, it includes: Obtain the MFCC value, LPCC value and fundamental frequency of the current user voiceprint model and each stored model in the voiceprint library and compare them, calculate the similarity, select the maximum similarity, and compare the maximum similarity with the similarity threshold; when the maximum similarity is greater than or equal to the similarity threshold, determine that the current user model is a registered user; when the maximum similarity is less than the similarity threshold, determine that the current user model is not a registered user; the similarity satisfies the following relationship: ; in, is the similarity, is the jth voiceprint feature of the current user’s voiceprint model, is the jth voiceprint feature of the storage model; is the number of voiceprint features, =3, when j=1, it corresponds to MFCC value; when j=2, it corresponds to LPCC value; when j=3, it corresponds to baseband.
6. The method for adaptive VAD parameter adjustment based on voiceprint recognition according to claim 5, characterized in that: When it is determined that the user is a registered user, the VAD parameters of the storage model corresponding to the maximum similarity in the voiceprint library are retrieved; When it is determined that the user is not a registered user, the preprocessed audio signal is divided into a silent segment and a speech segment according to the energy value, and the VAD parameter is adjusted in real time according to the characteristic fluctuation detection of the silent segment and the speech segment.
7. The method for adaptive VAD parameter adjustment based on voiceprint recognition according to claim 6, characterized in that: When a user is determined to be a non-registered user, and when the user is divided into silent and voice segments according to the energy value, the following are included: The energy value of each frame of audio signal is calculated, and the energy value is compared with the energy threshold to determine whether it is a silent segment or a speech segment. The energy value is obtained through the following relationship: ; ; in, is the energy value of the mth frame, is the kth sampling point of the mth frame, N is the number of sampling points per frame; N ≥ 2, is the energy threshold; when Greater than or equal to It is judged as a speech segment, otherwise it is a silent segment.
8. The method for adaptive VAD parameter adjustment based on voiceprint recognition according to claim 7, characterized in that: The VAD parameters are adjusted in real time according to the characteristic fluctuation detection of the silent segment, including: The VAD parameters include a zero-crossing rate detection threshold, a signal-to-noise ratio threshold, a hang-up time, and a duration; the zero-crossing rate detection threshold, the signal-to-noise ratio threshold, the hang-up time, and the duration are adjusted by the following relationship: ; ; ; ; ; ; ; in, is the adjusted zero-crossing rate detection threshold, is the initial detection threshold, is the spectrum volatility, is the h-th dimension MFCC value of the m-th frame, is the h-th dimension MFCC value of the m-1-th frame, K represents the dimension of the MFCC value, K is 13 or 20, is a constant, =10 -6 , is the signal-to-noise ratio, is the average energy value of all speech segments, is the average energy value of all silent segments, is the adjusted signal-to-noise ratio threshold, is the initial signal-to-noise ratio threshold, is the adjusted suspension time, is the initial suspension time, is the adjusted duration, is the initial duration, is the short-term energy fluctuation rate of the m-th frame, is the short-term energy fluctuation rate of the m-1th frame, is the spectrum volatility adjustment coefficient, is the spectrum stability adjustment coefficient, is the first signal-to-noise ratio adjustment coefficient, is the second signal-to-noise ratio adjustment coefficient, is the energy fluctuation influence coefficient, is the spectrum fluctuation influence coefficient, , , and The value range of is [0.05, 0.2].
9. The method for adaptive VAD parameter adjustment based on voiceprint recognition according to claim 8, characterized in that: Also includes: The time length of time that the storage model is stored in the voiceprint database is recorded, and when the time length is greater than or equal to a preset maximum time length, the storage model is deleted.
10. An adaptive VAD parameter adjustment system based on voiceprint recognition, applied to the adaptive VAD parameter adjustment method based on voiceprint recognition according to any one of claims 1 to 9, characterized in that: include: The collection and judgment module is configured to obtain an audio signal through an audio collection device, pre-process the audio signal and extract voiceprint features, map the voiceprint features to a user voiceprint model, compare the current user voiceprint model with a stored model in a voiceprint library, and determine whether the user is a registered user; The adjustment processing module is configured to determine the VAD parameter according to whether the user is a registered user; when the user is determined to be an unregistered user, the VAD parameter is adjusted in real time by detecting characteristic fluctuations of the silent segment and the speech segment; The storage management module is provided with a voiceprint library, and the storage management module is configured to store the adjusted VAD parameters in the voiceprint library.
Citation Information
Patent Citations
Sound activation detection apparatus and method
CN101320559A
Voice endpoint detection method based on voice features of speakers
CN108986844A
Speaker recognition device and method using voice signal analysis
US20100082341A1
Methods and systems for processing recorded audio content to enhance speech
US20220165289A1
Voice activity detection method, apparatus and device, and storage medium
WO2021139425A1
Cited By
Adaptive control method and device for audio equipment
CN120972573A
Multi-intelligent control switching method and system for intelligent doors and windows
CN121075329A
A method and system for switching between multiple intelligent controls for smart doors and windows
CN121075329B