Anti-voice-camouflage voiceprint verification method and device
By constructing a camouflage voiceprint verification model and using voiceprint authentication features for camouflage authentication and feature extraction, the problem of insufficient robustness in the face of voice-camouflage attacks is solved, and the accuracy of camouflage voiceprint recognition and user authentication accuracy is achieved.
Patent Information
- Application Number
- CN202510057767.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-23
AI Technical Summary
Traditional voiceprint recognition systems are not robust enough when facing voice camouflage attacks, making it difficult to effectively recognize and defend against camouflage voice.
By collecting user's registered voice samples, extracting registered voiceprint fusion features, building a camouflage voiceprint verification model, and using the signal energy and voiceprint information entropy of the voiceprint frame to determine the voiceprint identification characteristics, perform camouflage identification and feature extraction, and finally input the camouflage voiceprint fusion features into the verification model for verification.
It improves the robustness of the voiceprint verification device, enhances the recognition ability of camouflage voice, provides more accurate user identity verification, and effectively responds to voice camouflage attacks.
Smart Images

Figure CN120032648A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of voiceprint verification, and more specifically, to a voiceprint verification method and device that is resistant to voice masquerading. Background Art
[0002] Anti-spoofing voiceprint verification is a technology that aims to improve the defense capability of voiceprint recognition systems against spoofing attacks. With the development of speech synthesis, voice spoofing and deep fake technologies, traditional voiceprint recognition systems are facing more and more security threats, especially in sensitive applications such as identity authentication and financial payment, where spoofing attacks have become an increasingly serious problem.
[0003] However, traditional features have certain limitations when facing disguised voices. Disguised voices may change the spectral structure and rhythmic features of voices, or even simulate different voice samples, which makes verification methods that rely solely on traditional features easy to be deceived. In order to effectively identify disguised voiceprints, more advanced technologies and algorithms must be used to perform more comprehensive feature analysis and fusion. Therefore, how to identify whether the voice to be detected is a disguised voice and extract the features of the disguised voice for voiceprint verification to improve the robustness of the voiceprint verification device is a difficult problem faced by the industry. Summary of the invention
[0004] The present application provides a voiceprint verification method and device that are resistant to voice masquerading, which can identify whether the voice to be detected is a disguised voice, and extract the characteristics of the disguised voice for voiceprint verification, so as to improve the robustness of the voiceprint verification device.
[0005] In a first aspect, the present application provides a voiceprint verification method for resisting voice masquerading, the verification method comprising the following steps:
[0006] Collecting a registered voice sample of the user, performing voiceprint feature extraction on the registered voice sample to obtain a registered voiceprint fusion feature of the user, and using the registered voiceprint fusion feature to construct a disguised voiceprint verification model for the user;
[0007] Acquire a voice signal to be detected, determine the signal energy of each voice signal frame in the voice signal to be detected, determine the voiceprint information entropy corresponding to each voice signal frame in the voice signal to be detected through the corresponding signal energy, and determine the voiceprint identification feature of the voice signal to be detected according to the voiceprint information entropy corresponding to each voice signal frame;
[0008] Based on the voiceprint identification feature, the voice signal to be detected is subjected to disguise identification, and when the voice signal to be detected has a disguise feature, the disguise voiceprint feature is extracted from the voice signal to be detected to obtain a disguise voiceprint fusion feature;
[0009] The disguised voiceprint fusion feature is input into the disguised voiceprint verification model to perform voiceprint disguise verification, and the voiceprint verification result of the voice signal to be detected is obtained.
[0010] In this embodiment, after collecting the user's registration voice sample, the method further includes: performing audio noise reduction on the registration voice sample.
[0011] In this embodiment, voiceprint feature extraction is performed on the registered voice sample to obtain the registered voiceprint fusion feature of the user, specifically including:
[0012] Using a Mel filter bank to filter the registered speech sample, thereby obtaining a Mel frequency cepstral coefficient sequence of the registered speech sample;
[0013] Performing formant extraction on the registered speech sample to obtain formant features of the registered speech sample;
[0014] Extracting the fundamental frequency of the registered voice sample to obtain the fundamental frequency feature of the registered voice sample;
[0015] The registered voiceprint fusion feature of the user is determined according to the Mel-frequency cepstral coefficient sequence, the formant feature and the fundamental frequency feature.
[0016] In this embodiment, using the registered voiceprint fusion feature to construct a fake voiceprint verification model for the user specifically includes:
[0017] The registered voiceprint fusion features are input into a convolutional neural network for feature learning, and combined with a long short-term memory network to capture temporal features, thereby constructing a user's disguised voiceprint verification model.
[0018] In this embodiment, determining the signal energy of each speech signal frame in the speech signal to be detected specifically includes:
[0019] For each speech signal frame in the speech signal to be detected, determining the energy spectrum of the speech signal frame according to the amplitude spectrum of the frequency points in the speech signal frame;
[0020] The signal energy of the speech signal frame is determined by the energy spectrum, and then the signal energy of each speech signal frame in the speech signal to be detected is obtained.
[0021] In this embodiment, determining the voiceprint information entropy corresponding to each voice signal frame in the voice signal to be detected by using the corresponding signal energy specifically includes:
[0022] For each speech signal frame in the speech signal to be detected, determining the energy contribution of the speech signal frame based on the signal energy of the speech signal frame;
[0023] The energy contribution is mapped to voiceprint information to obtain the voiceprint information entropy corresponding to the voice signal frame, and then the voiceprint information entropy corresponding to each voice signal frame in the voice signal to be detected is obtained.
[0024] In this embodiment, performing disguise identification on the voice signal to be detected based on the voiceprint identification feature specifically includes:
[0025] The voiceprint identification feature of the voice signal to be detected is compared with a preset feature threshold. When the voiceprint identification feature is greater than the feature threshold, it is determined that the voice signal to be detected has a disguise feature. When the voiceprint identification feature is not greater than the feature threshold, it is determined that the voice signal to be detected does not have a disguise feature.
[0026] In this embodiment, extracting disguised voiceprint features from the voice signal to be detected to obtain disguised voiceprint fusion features specifically includes:
[0027] Using a Mel filter bank to filter the speech signal to be detected, thereby obtaining a Mel frequency cepstral coefficient sequence of the speech signal to be detected;
[0028] Performing formant extraction on the speech signal to be detected to obtain formant features of the speech signal to be detected;
[0029] Extracting the fundamental frequency of the speech signal to be detected to obtain the fundamental frequency characteristics of the speech signal to be detected;
[0030] The disguised voiceprint fusion feature is determined according to the Mel-frequency cepstral coefficient sequence, the formant feature and the fundamental frequency feature.
[0031] In this embodiment, the voiceprint verification result is a verification result used to indicate whether the voice signal to be detected is a disguised voice.
[0032] In a second aspect, the present application provides a user-specific password voiceprint recognition voiceprint verification device, which is used to perform a voiceprint verification method against voice masquerading. The voiceprint verification device includes a voiceprint verification unit, and the voiceprint verification unit includes:
[0033] A model building module is used to collect a registered voice sample of the user, extract voiceprint features from the registered voice sample, obtain a registered voiceprint fusion feature of the user, and use the registered voiceprint fusion feature to build a disguised voiceprint verification model for the user;
[0034] A voiceprint extraction module is used to obtain a voice signal to be detected, determine the signal energy of each voice signal frame in the voice signal to be detected, determine the voiceprint information entropy corresponding to each voice signal frame in the voice signal to be detected through the corresponding signal energy, and determine the voiceprint identification feature of the voice signal to be detected according to the voiceprint information entropy corresponding to each voice signal frame;
[0035] a disguise identification module, configured to perform disguise identification on the voice signal to be detected based on the voiceprint identification feature, and when the voice signal to be detected has a disguise feature, extract the disguise voiceprint feature of the voice signal to be detected to obtain a disguise voiceprint fusion feature;
[0036] The voiceprint verification module is used to input the disguised voiceprint fusion feature into the disguised voiceprint verification model to perform voiceprint disguise verification and obtain the voiceprint verification result of the voice signal to be detected.
[0037] The technical solution provided by the embodiments disclosed in this application has the following beneficial effects:
[0038] The method comprises the following steps: collecting a registered voice sample of a user, extracting voiceprint features from the registered voice sample, obtaining a registered voiceprint fusion feature of the user, and using the registered voiceprint fusion feature to construct a disguised voiceprint verification model for the user; obtaining a voice signal to be detected, determining the signal energy of each voice signal frame in the voice signal to be detected, determining the voiceprint information entropy corresponding to each voice signal frame in the voice signal to be detected by the corresponding signal energy, and determining the voiceprint identification feature of the voice signal to be detected according to the voiceprint information entropy corresponding to each voice signal frame; performing disguise identification on the voice signal to be detected based on the voiceprint identification feature, and when the voice signal to be detected has a disguise feature, extracting a disguise voiceprint feature of the voice signal to be detected, and obtaining a disguise voiceprint fusion feature; inputting the disguise voiceprint fusion feature into the disguise voiceprint verification model to perform voiceprint disguise verification, and obtaining a voiceprint verification result of the voice signal to be detected.
[0039] It can be seen that in the present application, firstly, a disguised voiceprint verification model is constructed based on the extracted registered voiceprint fusion features, which can effectively distinguish between real voiceprints and disguised voiceprints; then, by analyzing the energy of the voice signal frame, the local features of the voice signal to be detected can be obtained, and the voiceprint information entropy of each voice signal frame in the voice signal to be detected is calculated based on the signal energy, which can capture the multi-dimensional features of different voice signal frames, and the voiceprint identification features are determined by using the voiceprint information entropy of the voice signal frame, which can effectively quantify the disguise traces of the voice signal to be detected and improve the device's recognition ability of disguised voice; further, combining the disguise identification based on the voiceprint identification feature with the disguise voiceprint feature extraction can help the voiceprint verification device to quickly and accurately identify and extract the disguise features when facing complex disguised voices; finally, by inputting the disguised voiceprint fusion features into the disguised voiceprint verification model for verification, it can effectively identify whether the voice to be detected is a disguised voice, which can enhance the security, robustness and adaptability of the voiceprint verification device, and can effectively deal with voice disguise attacks and provide more accurate user identity authentication.
[0040] In summary, the technical solution adopted in the present application can identify whether the voice to be detected is a disguised voice, and extract the characteristics of the disguised voice for voiceprint verification, so as to improve the robustness of the voiceprint verification device. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:
[0042] Figure 1 It is a flow chart of the voiceprint verification method against voice masquerading provided by the present application;
[0043] Figure 2 is an exemplary flow chart for determining the voiceprint information entropy corresponding to each voice signal frame in the voice signal to be detected provided by the present application;
[0044] Figure 3 is an exemplary flow chart for determining camouflage voiceprint fusion features provided by the present application;
[0045] Figure 4 This is a module structure diagram of the voiceprint verification unit provided in this application. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical scheme and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and drawings. The schematic implementation modes and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention. It should be noted that the present invention is already in the actual development and use stage.
[0047] Embodiment 1: In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods. Figure 1 As shown in FIG. 1 , this figure is an exemplary flow chart of a voiceprint verification method for resisting voice masquerading according to this embodiment of the present application, and the verification method includes the following steps:
[0048] In step S1, a registered voice sample of the user is collected, voiceprint features are extracted from the registered voice sample to obtain a registered voiceprint fusion feature of the user, and a disguised voiceprint verification model of the user is constructed using the registered voiceprint fusion feature.
[0049] In the specific implementation, first, the user's registered voice sample can be collected through a voice collection device, such as a microphone; then, after collecting the user's registered voice sample, it also includes: performing audio noise reduction on the registered voice sample, and an audio noise reduction algorithm can be used to perform audio noise reduction on the registered voice sample. The audio noise reduction algorithm used in this application is the Wiener filter algorithm. Other audio noise reduction algorithms can also be used in actual implementation, which is not limited here.
[0050] In this embodiment, voiceprint feature extraction is performed on the registered voice sample to obtain the registered voiceprint fusion feature of the user, which can be specifically performed in the following manner, namely:
[0051] Using a Mel filter bank to filter the registered speech sample, thereby obtaining a Mel frequency cepstral coefficient sequence of the registered speech sample;
[0052] Performing formant extraction on the registered speech sample to obtain formant features of the registered speech sample;
[0053] Extracting the fundamental frequency of the registered voice sample to obtain the fundamental frequency feature of the registered voice sample;
[0054] The registered voiceprint fusion feature of the user is determined according to the Mel-frequency cepstral coefficient sequence, the formant feature and the fundamental frequency feature.
[0055] In specific implementation, first, the registered speech samples can be filtered using a Mel filter bank, that is, the registered speech samples are framed by a window function, and the signal of each frame is subjected to a short-time Fourier transform (STFT) to obtain a spectrum, thereby converting the spectrum into an energy spectrum on a Mel frequency scale through a Mel filter bank, and after logarithmic transformation of the Mel frequency energy spectrum, a discrete cosine transform (DCT) is applied to obtain Mel frequency cepstral coefficients (MFCC) of multiple registered speech samples. The Mel frequency cepstral coefficients contain the spectrum characteristics of the registered speech samples and can effectively capture information such as the timbre and sound quality of the speech. The sequence composed of all Mel frequency cepstral coefficients can be used as the Mel frequency cepstral coefficient sequence of the registered speech samples; then, the registered speech samples can be subjected to formant extraction to obtain the formant characteristics of the registered speech samples, wherein the formant characteristics are characteristics used to reflect the resonance characteristics of the vocal tract in the registered speech samples, and the spectral characteristics of the prediction error signal can be obtained by performing linear predictive coding (LPC) on the speech signal, thereby identifying the position and amplitude of the formant by analyzing the extreme values of the linear predictive coding spectrum, and the formant characteristics of the registered speech samples can be extracted in the above manner.
[0056] In addition, in a specific implementation, the fundamental frequency of the registered voice sample can be extracted to obtain the fundamental frequency feature of the registered voice sample, wherein the fundamental frequency feature is the lowest frequency of the registered voice sample, which is usually controlled by the user's voice frequency, and the fundamental frequency extraction algorithm (such as the autocorrelation method, the time domain method or the Fourier transform method) can be used to extract the fundamental frequency feature of the registered voice sample; finally, the registered voiceprint fusion feature of the user can be determined according to the Mel frequency cepstral coefficient sequence, the resonance peak feature and the fundamental frequency feature, wherein the registered voiceprint fusion feature represents the comprehensive voiceprint feature of the user at the time of registration, and the Mel frequency cepstral coefficient sequence, the resonance peak feature and the fundamental frequency feature can be feature spliced to obtain the registered voiceprint fusion feature of the user, and the registered voiceprint fusion feature obtained after fusion usually has a higher degree of discrimination, can provide a richer and more personalized voiceprint description, and improve the accuracy and anti-counterfeiting ability of voiceprint verification.
[0057] In this embodiment, the registered voiceprint fusion feature is used to construct a fake voiceprint verification model of the user, which can be specifically constructed in the following manner, namely:
[0058] The registered voiceprint fusion features are input into a convolutional neural network for feature learning, and combined with a long short-term memory network to capture temporal features, thereby constructing a user's disguised voiceprint verification model.
[0059] In specific implementation, convolutional neural network (CNN) is mainly used to extract local features. For voiceprint verification tasks, convolutional neural network can extract discriminative local patterns from the input registered voiceprint fusion features through convolutional layers, that is, the convolutional layer scans the input registered voiceprint fusion features through the convolution kernel, and learns the local laws of the features. The filters of each convolutional layer will learn different abstract representations of the input registered voiceprint fusion features; Long short-term memory network (LSTM) is a special recurrent neural network (RNN) that is good at learning and capturing long-term dependencies. The long short-term memory network can effectively process input data with time series characteristics and convolution. The features extracted by the neural network are input into the long short-term memory network. The role of the long short-term memory network is to learn the temporal dependencies and long-term dependencies in the input registered voiceprint fusion features through its gating mechanism (such as forget gate, input gate, and output gate). In voiceprint verification, the long short-term memory network can capture the changes in the time dimension of the voice signal by dynamically modeling the registered voiceprint fusion features. Finally, the disguised voiceprint verification model performs disguised voiceprint verification through an output layer. The output layer is usually a binary classification layer, which is used to determine whether the voice to be detected is a disguised voiceprint. The output of this output layer is usually a probability value, indicating the possibility that the input voice is a disguised voiceprint.
[0060] It should be noted that the disguised voiceprint verification model is constructed based on the extracted registered voiceprint fusion features, which can effectively distinguish between real voiceprints and disguised voiceprints. In addition, the disguised voiceprint verification model not only has high recognition accuracy, but can also effectively deal with disguised voiceprint attacks in different environments and has strong robustness.
[0061] In step S2, a speech signal to be detected is obtained, and the signal energy of each speech signal frame in the speech signal to be detected is determined. The voiceprint information entropy corresponding to each speech signal frame in the speech signal to be detected is determined through the corresponding signal energy. The voiceprint identification feature of the speech signal to be detected is determined based on the voiceprint information entropy corresponding to each speech signal frame.
[0062] In specific implementation, the voice signal to be detected can be recorded in real time by a voice acquisition device, and noise and echo are removed after preprocessing the voice signal to be detected.
[0063] In this embodiment, the signal energy of each speech signal frame in the speech signal to be detected may be determined in the following manner, namely:
[0064] For each speech signal frame in the speech signal to be detected, determining the energy spectrum of the speech signal frame according to the amplitude spectrum of the frequency points in the speech signal frame;
[0065] The signal energy of the speech signal frame is determined by the energy spectrum, and then the signal energy of each speech signal frame in the speech signal to be detected is obtained.
[0066] In specific implementation, first, when processing the voice signal to be detected, the voice signal to be detected can be divided into multiple small voice signal frames, each voice signal frame represents a section of time series data in the voice signal to be detected. The frame length of the voice signal frame in this application is 25ms, so that the voice signal frame is subjected to fast Fourier transform, and converted from the time domain to the frequency domain, and the amplitude spectrum of the voice signal frame can be calculated; then, the energy spectrum of the voice signal frame can be determined according to the amplitude spectrum of the frequency point in the voice signal frame. In actual implementation, the energy spectrum of the voice signal frame can be determined by the following formula:
[0067]
[0068] Among them, f i represents the energy spectrum of the i-th speech signal frame, y j represents the amplitude spectrum of the jth frequency point in the ith speech signal frame, and n represents the total number of frequency points in the ith speech signal frame; finally, the signal energy of the speech signal frame can be determined by the energy spectrum, wherein the signal energy is used to represent the energy activation degree of the speech signal frame, and the larger the signal energy is, the higher the energy activation degree is. In actual implementation, the square of the absolute value of the energy spectrum of the speech signal frame can be used as the signal energy of the speech signal frame. The signal energy of each speech signal frame in the speech signal to be detected can be obtained in the above manner.
[0069] Preferably, in this embodiment, reference Figure 2 As shown in FIG. 1 , this figure is an exemplary flow chart of determining the voiceprint information entropy corresponding to each voice signal frame in the voice signal to be detected in an embodiment of the present application. In this embodiment, the voiceprint information entropy corresponding to each voice signal frame in the voice signal to be detected can be determined by the corresponding signal energy, which can be implemented by the following steps:
[0070] First, in step S21, for each speech signal frame in the speech signal to be detected, the energy contribution of the speech signal frame is determined based on the signal energy of the speech signal frame;
[0071] Then, in step S22, voiceprint information mapping is performed on the energy contribution to obtain the voiceprint information entropy corresponding to the speech signal frame, and then the voiceprint information entropy corresponding to each speech signal frame in the speech signal to be detected is obtained.
[0072] In specific implementation, first, for each speech signal frame in the speech signal to be detected, the energy contribution of the speech signal frame can be determined based on the signal energy of the speech signal frame, wherein the energy contribution represents the energy contribution degree of the speech signal frame in the speech signal to be detected. In actual implementation, the energy contribution of the speech signal frame can be determined by the following formula:
[0073]
[0074] Among them, p i represents the energy contribution of the i-th speech signal frame, k i represents the signal energy of the i-th speech signal frame, k v Represents the signal energy of the vth speech signal frame, m represents the total number of speech signal frames in the speech signal to be detected; then, the energy contribution can be mapped to the voiceprint information, that is, the energy contribution of the speech signal frame is mapped to the logarithmic domain, so as to obtain the voiceprint information entropy corresponding to the speech signal frame, that is, the voiceprint information entropy corresponding to the speech signal frame = -log the energy contribution of the speech signal frame, wherein the voiceprint information entropy represents the content of voiceprint information in the speech signal frame. The voiceprint information entropy corresponding to each speech signal frame in the speech signal to be detected can be obtained in the above manner.
[0075] In this embodiment, the voiceprint identification feature of the voice signal to be detected is determined based on the voiceprint information entropy corresponding to each voice signal frame; it should be noted that, in the present application, the voiceprint identification feature is a feature used to determine whether the signal to be detected has traces of disguise; in specific implementation, the standard deviation of the voiceprint information entropy corresponding to all voice signal frames can be used as the voiceprint identification feature of the voice signal to be detected.
[0076] It should be noted that, through energy analysis of the voice signal frame, the local characteristics of the voice signal to be detected can be obtained, and the voiceprint information entropy of each voice signal frame in the voice signal to be detected is calculated based on the signal energy, so that the multi-dimensional characteristics of different voice signal frames can be captured. The voiceprint identification characteristics are determined by using the voiceprint information entropy of the voice signal frame, which can effectively quantify the disguised traces of the voice signal to be detected and improve the device's recognition ability for disguised voice.
[0077] In step S3, disguise identification is performed on the voice signal to be detected based on the voiceprint identification feature. When the voice signal to be detected has a disguise feature, disguise voiceprint feature extraction is performed on the voice signal to be detected to obtain a disguise voiceprint fusion feature.
[0078] In this embodiment, the disguise identification of the voice signal to be detected based on the voiceprint identification feature can be specifically performed in the following manner, namely:
[0079] The voiceprint identification feature of the voice signal to be detected is compared with a preset feature threshold. When the voiceprint identification feature is greater than the feature threshold, it is determined that the voice signal to be detected has a disguise feature. When the voiceprint identification feature is not greater than the feature threshold, it is determined that the voice signal to be detected does not have a disguise feature.
[0080] In the specific implementation, first, the feature threshold can be preset through analysis on the data training set and historical data statistics, which will not be repeated here; then, the voiceprint identification feature extracted from the voice signal to be detected is compared with the preset feature threshold. When the value of the voiceprint identification feature is greater than the set feature threshold, it means that the feature of the voice signal to be detected is significantly different from that of normal speech, and there may be disguised features. When the value of the voiceprint identification feature is not greater than the set feature threshold, it means that the feature of the voice signal to be detected is closer to that of normal speech, and it is unlikely to be disguised speech.
[0081] Preferably, in this embodiment, reference Figure 3 As shown in FIG. 1 , this figure is an exemplary flow chart of determining the disguised voiceprint fusion feature in an embodiment of the present application. In this embodiment, the disguised voiceprint feature extraction is performed on the voice signal to be detected, and the disguised voiceprint fusion feature is obtained, which can be specifically implemented by the following steps:
[0082] First, in step S31, the speech signal to be detected is filtered using a Mel filter bank to obtain a Mel frequency cepstral coefficient sequence of the speech signal to be detected;
[0083] Then, in step S32, formant extraction is performed on the speech signal to be detected to obtain formant features of the speech signal to be detected;
[0084] Secondly, in step S33, the fundamental frequency of the speech signal to be detected is extracted to obtain the fundamental frequency feature of the speech signal to be detected;
[0085] Finally, in step S34, the disguised voiceprint fusion feature is determined according to the Mel-frequency cepstral coefficient sequence, the formant feature and the fundamental frequency feature.
[0086] In specific implementation, the method in step S1 can be used to extract disguised voiceprint features from the voice signal to be detected, thereby obtaining disguised voiceprint fusion features, which are comprehensive features used to indicate that the voice signal to be detected contains disguised features.
[0087] It should be noted that combining camouflage identification based on voiceprint identification features with camouflage voiceprint feature extraction can help the voiceprint verification device to quickly and accurately identify and extract camouflage features when faced with complex camouflage voices. This not only helps to improve the recognition accuracy of camouflage voices, but also enhances the robustness of the device, enabling it to cope with the challenges of various camouflage technologies, thereby improving the security and practicality of the voiceprint verification device.
[0088] In step S4, the disguised voiceprint fusion feature is input into the disguised voiceprint verification model to perform voiceprint disguise verification, and obtain the voiceprint verification result of the voice signal to be detected.
[0089] In the specific implementation, first, the extracted disguised voiceprint fusion features are used as input to the disguised voiceprint verification model obtained in step S1; then, the disguised voiceprint verification model will analyze the disguised voiceprint fusion features through the existing training data and compare them with the known disguised voice patterns to determine whether the voice signal to be detected contains disguised features. During the verification process, the disguised voiceprint verification model compares the input disguised voiceprint fusion features with the standard voiceprint features based on the knowledge learned during the training process. The disguised voiceprint verification model will output a classification result based on the analysis results, that is, the voiceprint verification result of the voice signal to be detected. The voiceprint verification result is a verification result used to indicate whether the voice signal to be detected is a disguised voice. For example, if the disguised voiceprint verification model determines that the voice signal to be detected is disguised, the voiceprint verification result is "disguised voice"; if it is determined to be real voice, the voiceprint verification result is "real voice".
[0090] It should be noted that by inputting the disguised voiceprint fusion features into the disguised voiceprint verification model for verification, it is possible to effectively identify whether the voice to be detected is a disguised voice, which can enhance the security, robustness and adaptability of the voiceprint verification device, and can effectively respond to voice disguise attacks and provide more accurate user identity authentication.
[0091] It can be seen that in the present application, firstly, a disguised voiceprint verification model is constructed based on the extracted registered voiceprint fusion features, which can effectively distinguish between real voiceprints and disguised voiceprints; then, by analyzing the energy of the voice signal frame, the local features of the voice signal to be detected can be obtained, and the voiceprint information entropy of each voice signal frame in the voice signal to be detected is calculated based on the signal energy, which can capture the multi-dimensional features of different voice signal frames, and the voiceprint identification features are determined by using the voiceprint information entropy of the voice signal frame, which can effectively quantify the disguise traces of the voice signal to be detected and improve the device's recognition ability of disguised voice; further, combining the disguise identification based on the voiceprint identification feature with the disguise voiceprint feature extraction can help the voiceprint verification device to quickly and accurately identify and extract the disguise features when facing complex disguised voices; finally, by inputting the disguised voiceprint fusion features into the disguised voiceprint verification model for verification, it can effectively identify whether the voice to be detected is a disguised voice, which can enhance the security, robustness and adaptability of the voiceprint verification device, and can effectively deal with voice disguise attacks and provide more accurate user identity authentication.
[0092] In summary, the technical solution adopted in the present application can identify whether the voice to be detected is a disguised voice, and extract the characteristics of the disguised voice for voiceprint verification, so as to improve the robustness of the voiceprint verification device.
[0093] Embodiment 2: The present application provides a user-specific password voiceprint recognition and voiceprint verification device, the voiceprint verification device includes a voiceprint verification unit, referring to Figure 4 As shown, this figure is a schematic diagram of a voiceprint verification unit according to this embodiment of the present application, and the voiceprint verification unit includes:
[0094] The model building module 100 is used to collect the user's registered voice sample, extract the voiceprint feature of the registered voice sample, obtain the user's registered voiceprint fusion feature, and use the registered voiceprint fusion feature to build the user's disguised voiceprint verification model;
[0095] The voiceprint extraction module 200 is used to obtain a voice signal to be detected, determine the signal energy of each voice signal frame in the voice signal to be detected, determine the voiceprint information entropy corresponding to each voice signal frame in the voice signal to be detected through the corresponding signal energy, and determine the voiceprint identification feature of the voice signal to be detected according to the voiceprint information entropy corresponding to each voice signal frame;
[0096] The disguise identification module 300 is used to perform disguise identification on the voice signal to be detected based on the voiceprint identification feature, and when the voice signal to be detected has a disguise feature, extract the disguise voiceprint feature of the voice signal to be detected to obtain a disguise voiceprint fusion feature;
[0097] The voiceprint verification module 400 is used to input the disguised voiceprint fusion feature into the disguised voiceprint verification model to perform voiceprint disguise verification and obtain the voiceprint verification result of the voice signal to be detected.
[0098] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A voiceprint verification method for resisting voice masquerading, characterized in that: The verification method comprises the following steps: Collecting a registered voice sample of the user, performing voiceprint feature extraction on the registered voice sample to obtain a registered voiceprint fusion feature of the user, and using the registered voiceprint fusion feature to construct a disguised voiceprint verification model for the user; Acquire a voice signal to be detected, determine the signal energy of each voice signal frame in the voice signal to be detected, determine the voiceprint information entropy corresponding to each voice signal frame in the voice signal to be detected through the corresponding signal energy, and determine the voiceprint identification feature of the voice signal to be detected according to the voiceprint information entropy corresponding to each voice signal frame; Based on the voiceprint identification feature, the voice signal to be detected is subjected to disguise identification, and when the voice signal to be detected has a disguise feature, the disguise voiceprint feature is extracted from the voice signal to be detected to obtain a disguise voiceprint fusion feature; The disguised voiceprint fusion feature is input into the disguised voiceprint verification model to perform voiceprint disguise verification, and the voiceprint verification result of the voice signal to be detected is obtained.
2. A voiceprint verification method for resisting voice masquerading as claimed in claim 1, characterized in that: After collecting the user's registration voice sample, the method further includes: performing audio noise reduction on the registration voice sample.
3. A voiceprint verification method for resisting voice masquerading as claimed in claim 1, characterized in that: Extracting voiceprint features from the registered voice sample to obtain the user's registered voiceprint fusion features specifically includes: Using a Mel filter bank to filter the registered speech sample, thereby obtaining a Mel frequency cepstral coefficient sequence of the registered speech sample; Performing formant extraction on the registered speech sample to obtain formant features of the registered speech sample; Extracting the fundamental frequency of the registered voice sample to obtain the fundamental frequency feature of the registered voice sample; The registered voiceprint fusion feature of the user is determined according to the Mel-frequency cepstral coefficient sequence, the formant feature and the fundamental frequency feature.
4. A voiceprint verification method for resisting voice masquerading as claimed in claim 1, characterized in that: Using the registered voiceprint fusion feature to construct a user's disguised voiceprint verification model specifically includes: The registered voiceprint fusion features are input into a convolutional neural network for feature learning, and combined with a long short-term memory network to capture temporal features, thereby constructing a user's disguised voiceprint verification model.
5. The voiceprint verification method against voice masquerading as claimed in claim 1, characterized in that: Determining the signal energy of each speech signal frame in the speech signal to be detected specifically includes: For each speech signal frame in the speech signal to be detected, determining the energy spectrum of the speech signal frame according to the amplitude spectrum of the frequency points in the speech signal frame; The signal energy of the speech signal frame is determined by the energy spectrum, and then the signal energy of each speech signal frame in the speech signal to be detected is obtained.
6. A voiceprint verification method for resisting voice masquerading as claimed in claim 1, characterized in that: Determining the voiceprint information entropy corresponding to each voice signal frame in the voice signal to be detected by using the corresponding signal energy specifically includes: For each speech signal frame in the speech signal to be detected, determining the energy contribution of the speech signal frame based on the signal energy of the speech signal frame; The energy contribution is mapped to voiceprint information to obtain the voiceprint information entropy corresponding to the voice signal frame, and then the voiceprint information entropy corresponding to each voice signal frame in the voice signal to be detected is obtained.
7. A voiceprint verification method for resisting voice masquerading as claimed in claim 1, characterized in that: Performing disguise identification on the voice signal to be detected based on the voiceprint identification feature specifically includes: The voiceprint identification feature of the voice signal to be detected is compared with a preset feature threshold. When the voiceprint identification feature is greater than the feature threshold, it is determined that the voice signal to be detected has a disguise feature. When the voiceprint identification feature is not greater than the feature threshold, it is determined that the voice signal to be detected does not have a disguise feature.
8. The voiceprint verification method against voice masquerading as claimed in claim 1, characterized in that: Extracting the disguised voiceprint feature of the voice signal to be detected to obtain the disguised voiceprint fusion feature specifically includes: Using a Mel filter bank to filter the speech signal to be detected, thereby obtaining a Mel frequency cepstral coefficient sequence of the speech signal to be detected; Performing formant extraction on the speech signal to be detected to obtain formant features of the speech signal to be detected; Extracting the fundamental frequency of the speech signal to be detected to obtain the fundamental frequency characteristics of the speech signal to be detected; The disguised voiceprint fusion feature is determined according to the Mel-frequency cepstral coefficient sequence, the formant feature and the fundamental frequency feature.
9. The voiceprint verification method against voice masquerading as claimed in claim 1, characterized in that: The voiceprint verification result is a verification result used to indicate whether the voice signal to be detected is a disguised voice.
10. A voiceprint verification device for resisting voice masquerading, used to execute a voiceprint verification method for resisting voice masquerading as claimed in any one of claims 1 to 9, characterized in that: The voiceprint verification device includes a voiceprint verification unit, and the voiceprint verification unit includes: A model building module is used to collect a registered voice sample of the user, extract voiceprint features from the registered voice sample, obtain a registered voiceprint fusion feature of the user, and use the registered voiceprint fusion feature to build a disguised voiceprint verification model for the user; A voiceprint extraction module is used to obtain a voice signal to be detected, determine the signal energy of each voice signal frame in the voice signal to be detected, determine the voiceprint information entropy corresponding to each voice signal frame in the voice signal to be detected through the corresponding signal energy, and determine the voiceprint identification feature of the voice signal to be detected according to the voiceprint information entropy corresponding to each voice signal frame; a disguise identification module, configured to perform disguise identification on the voice signal to be detected based on the voiceprint identification feature, and when the voice signal to be detected has a disguise feature, extract the disguise voiceprint feature of the voice signal to be detected to obtain a disguise voiceprint fusion feature; The voiceprint verification module is used to input the disguised voiceprint fusion feature into the disguised voiceprint verification model to perform voiceprint disguise verification and obtain the voiceprint verification result of the voice signal to be detected.