Identity recognition method and device based on voiceprint features

Through audio separation and noise reduction processing technology, combined with feature classification models, the problem of voiceprint recognition accuracy in noisy environments is solved, achieving more efficient and accurate user identity authentication.

CN120708624APending Publication Date: 2025-09-26AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511048562.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing voiceprint recognition technology lacks an effective cleaning and processing mechanism when dealing with noise and irregular audio, resulting in reduced recognition accuracy and an inability to meet the security and accuracy requirements of bank user identity authentication.

Method used

The audio separation model and residual denoising diffusion model are used to preprocess and reduce the noise of the voice data, extract acoustic features and perform weighted processing, and then compare them with the feature classification model to determine the user's identity.

Benefits of technology

It improves the clarity of voice signals, enhances the robustness of the system, effectively reduces noise interference, and significantly improves the voiceprint recognition rate and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708624A_ABST
    Figure CN120708624A_ABST
Patent Text Reader

Abstract

The invention provides an identity recognition method and device based on voiceprint features, and the method comprises the steps: carrying out the preprocessing of obtained voice data, and obtaining a preprocessed audio; performing noise reduction processing on the preprocessed audio by using an audio separation model and a residual denoising diffusion model to obtain a noise-reduced audio; extracting acoustic features from the noise-reduced audio, and analyzing each dimension of the acoustic features under different noises to obtain a weight coefficient of each dimension; performing weighting processing on each dimension of the acoustic features based on the weight coefficient to obtain a feature matrix; and performing feature comparison according to the feature matrix and a feature classification model, and determining a user identity corresponding to the voice data. And the audio separation model and the residual denoising diffusion model are used for voice denoising, so that environmental noise is removed, the definition of voice signals is improved, and the robustness of the system is enhanced. A plurality of user audios are separated by adopting the feature classification model, so that interference is effectively reduced, and the voiceprint recognition rate is improved. Meanwhile, the feature matrix also significantly improves the recognition efficiency and accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning and artificial intelligence technology, and in particular to a method and device for identity recognition based on voiceprint features. Background Art

[0002] With the rapid development of informatization and electronicization in the financial industry, banks face increasingly severe security challenges in user identity verification. Traditional identity authentication methods, such as passwords and facial recognition, no longer meet the high security and accuracy requirements of modern banking. These methods are vulnerable to threats such as password leaks and identity theft, and cannot effectively prevent identity forgery and security risks. Therefore, improving the security and accuracy of bank user identity verification has become a key issue in the development of FinTech.

[0003] In audio recognition technology, voiceprint recognition, as an emerging biometric technology, is gradually being applied to identity authentication in the financial industry. Voiceprint recognition technology authenticates the user's identity based on the user's sound wave characteristics. It is contactless and convenient, and can improve security to a certain extent. However, existing voiceprint recognition technology still has some shortcomings. First, many existing solutions rely on machine learning models to extract audio features and lack in-depth processing of relevant acoustic features. This may lead to the loss or underutilization of voiceprint feature information. Secondly, current voiceprint recognition systems usually assume that the input audio comes from a pure environment. However, in actual production environments, noisy background noise often affects the clarity of the voiceprint, resulting in a decrease in recognition accuracy. Existing technologies lack effective cleaning and processing mechanisms when dealing with these noisy and non-standard audio, further affecting the accuracy of the recognition results.

[0004] Therefore, improving the feature processing capabilities and adaptability to noise of voiceprint recognition technology has become the key to improving the security and reliability of bank user authentication systems. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method and apparatus for identity recognition based on voiceprint features to solve the current problem of poor precision and low accuracy in user identity recognition using voiceprints.

[0006] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:

[0007] The first aspect of the present invention discloses an identity recognition method based on voiceprint features, the method comprising:

[0008] Preprocessing the acquired voice data to obtain preprocessed audio;

[0009] Performing noise reduction processing on the preprocessed audio using an audio separation model and a residual denoising diffusion model to obtain a noise-reduced audio;

[0010] Extracting acoustic features from the noise reduction frequency, analyzing dimensions of the acoustic features under different noise conditions, and obtaining weight coefficients for the dimensions;

[0011] Performing weighted processing on each dimension of the acoustic feature based on the weight coefficient to obtain a feature matrix;

[0012] A feature comparison is performed based on the feature matrix and the feature classification model to determine the user identity corresponding to the voice data.

[0013] Preferably, the preprocessing of the acquired voice data to obtain preprocessed audio includes:

[0014] Setting a sampling frequency, and acquiring voice data based on the sampling frequency;

[0015] Performing a framing operation on the voice data, and performing endpoint detection on the voice data after the framing operation;

[0016] Dividing the voice data after the framing operation into information segment voice data and silence segment voice data according to the detected breakpoints;

[0017] The information segment voice data and the silence segment voice data are pre-enhanced using a high-pass filter to obtain pre-processed audio.

[0018] Preferably, the performing noise reduction processing on the pre-processed audio using the audio separation model and the residual denoising diffusion model to obtain the noise-reduced audio includes:

[0019] Calculating short-time Fourier transform features and phase difference features of the preprocessed audio;

[0020] Normalizing the short-time Fourier transform features and the phase difference features, and inputting the results into an audio separation model to predict an amplitude spectrum mask of each sound source signal;

[0021] Determining a short-time Fourier transform feature of each sound source based on the amplitude spectrum mask and the short-time Fourier transform feature;

[0022] Perform inverse Fourier transform on the short-time Fourier transform characteristics of each sound source to obtain the time domain signal of each sound source;

[0023] plotting a spectrogram based on the time domain signal of each sound source;

[0024] Performing denoising processing on the spectrum graph using a residual denoising diffusion model to obtain a processed spectrum graph;

[0025] Convert the processed spectrogram back to the time domain signal to obtain the noise-reduced frequency spectrum.

[0026] Preferably, extracting acoustic features from the noise reduction frequency, analyzing each dimension of the acoustic features under different noises, and obtaining a weight coefficient of each dimension includes:

[0027] Extracting Mel-frequency cepstral coefficient features from the noise reduction frequency using a triangular filter;

[0028] extracting gamma-tone frequency cepstral coefficient features from the noise reduction frequency using a gamma-tone filter;

[0029] Fusing the Mel-frequency cepstral coefficient feature and the gamma-tone frequency cepstral coefficient feature to obtain an acoustic feature;

[0030] Performing extreme value normalization processing on the acoustic features, calculating the intra-class sum of square differences of each dimensional feature of the acoustic features under different noise conditions, and calculating the inter-class sum of square differences of each dimensional feature of the acoustic features under different noise conditions;

[0031] Divide the sum of square differences within each class of the features of each dimension by the sum of square differences between each class of the features of each dimension to obtain an initial weight coefficient;

[0032] The initial weight coefficients are normalized to obtain weight coefficients of each dimension.

[0033] Preferably, performing feature comparison based on the feature matrix and the feature classification model to determine the user identity corresponding to the voice data includes:

[0034] Inputting the characteristic matrix into the residual diffusion denoising model for denoising;

[0035] The feature matrix after noise reduction is input into the feature classification model to extract the voiceprint features;

[0036] The voiceprint feature is calculated and compared with the reserved user voice feature, and the user identity corresponding to the voice data is determined based on the comparison result.

[0037] A second aspect of the present invention discloses an identity recognition device based on voiceprint features, the device comprising:

[0038] A preprocessing unit, configured to preprocess the acquired voice data to obtain preprocessed audio;

[0039] a noise reduction unit, configured to perform noise reduction processing on the preprocessed audio using an audio separation model and a residual denoising diffusion model to obtain a noise-reduced audio;

[0040] an analysis unit, configured to extract acoustic features from the noise reduction frequency, analyze each dimension of the acoustic features under different noise conditions, and obtain a weight coefficient for each dimension;

[0041] a weighting processing unit, configured to perform weighted processing on each dimension of the acoustic feature based on the weight coefficient to obtain a feature matrix;

[0042] The comparison unit is used to perform feature comparison based on the feature matrix and the feature classification model to determine the user identity corresponding to the voice data.

[0043] Preferably, the pre-processing unit includes:

[0044] A setting module, configured to set a sampling frequency and obtain voice data based on the sampling frequency;

[0045] A detection module, configured to perform a framing operation on the voice data and perform endpoint detection on the voice data after the framing operation;

[0046] A division module, configured to divide the voice data after the framing operation into information segment voice data and silence segment voice data according to the detected breakpoints;

[0047] The preprocessing module is used to perform pre-enhancement processing on the information segment voice data and the silence segment voice data using a high-pass filter to obtain preprocessed audio.

[0048] Preferably, the noise reduction unit includes:

[0049] A first calculation module, configured to calculate a short-time Fourier transform feature and a phase difference feature of the preprocessed audio;

[0050] A prediction module, configured to normalize the short-time Fourier transform features and the phase difference features, input the normalized features into an audio separation model, and predict an amplitude spectrum mask for each sound source signal;

[0051] a determination module, configured to determine a short-time Fourier transform feature of each sound source based on the amplitude spectrum mask and the short-time Fourier transform feature;

[0052] A transformation module is used to perform inverse Fourier transform on the short-time Fourier transform characteristics of each sound source to obtain the time domain signal of each sound source;

[0053] A drawing module, configured to draw a spectrum diagram based on the time domain signal of each sound source;

[0054] A noise reduction processing module, configured to perform noise reduction processing on the spectrum graph using a residual denoising diffusion model to obtain a processed spectrum graph;

[0055] The conversion module is used to convert the processed spectrum graph back into a time domain signal to obtain a noise-reduced frequency signal.

[0056] Preferably, the analysis unit comprises:

[0057] A first extraction module is used to extract Mel-frequency cepstral coefficient features from the noise reduction frequency using a triangular filter;

[0058] A second extraction module is configured to extract gamma-tone frequency cepstral coefficient features from the noise reduction frequency using a gamma-tone filter;

[0059] A fusion module, configured to fuse the Mel-frequency cepstral coefficient feature and the gamma-tone frequency cepstral coefficient feature to obtain an acoustic feature;

[0060] a second calculation module, configured to perform extreme value normalization processing on the acoustic features, calculate the within-class sum of square differences of each dimensional feature of the acoustic features under different noise conditions, and calculate the between-class sum of square differences of each dimensional feature of the acoustic features under different noise conditions;

[0061] A third calculation module is used to divide the intra-class sum of square differences of the features of each dimension by the inter-class sum of square differences of the features of each dimension to obtain an initial weight coefficient;

[0062] The normalization processing module is used to perform normalization processing on the initial weight coefficient to obtain the weight coefficient of each dimension.

[0063] Preferably, the comparison unit is specifically used to: input the feature matrix into the residual diffusion denoising model for noise reduction processing; input the feature matrix after noise reduction processing into the feature classification model to extract voiceprint features; calculate the voiceprint features and compare them with the reserved user voice features, and determine the user identity corresponding to the voice data based on the comparison results.

[0064] Based on the above-mentioned embodiment of the present invention, a method and device for identity recognition based on voiceprint features are provided. The acquired voice data is preprocessed to obtain preprocessed audio; the preprocessed audio is denoised using an audio separation model and a residual denoising diffusion model to obtain a denoised audio; acoustic features are extracted from the denoised audio, and the dimensions of the acoustic features under different noise conditions are analyzed to obtain weight coefficients for each dimension; each dimension of the acoustic features is weighted based on the weight coefficients to obtain a feature matrix; and feature comparison is performed based on the feature matrix and a feature classification model to determine the user identity corresponding to the voice data. The audio separation model and the residual denoising diffusion model are used to perform voice denoising, remove environmental noise, improve the clarity of the voice signal, and enhance the robustness of the system. The feature classification model is used to separate the audio of multiple users, effectively reducing interference, thereby improving the voiceprint recognition rate. At the same time, the feature matrix also significantly improves the recognition efficiency and accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0066] Figure 1 A flow chart of a method for identity recognition based on voiceprint features provided by an embodiment of the present invention;

[0067] Figure 2 A schematic diagram of the audio sampling process provided by an embodiment of the present invention;

[0068] Figure 3 The audio separation flow chart of the U-Conformer model provided in an embodiment of the present invention;

[0069] Figure 4 Flowchart of audio noise reduction using the RDDM model provided by an embodiment of the present invention;

[0070] Figure 5 A schematic diagram of the MFCC feature parameter extraction process provided by an embodiment of the present invention;

[0071] Figure 6 A schematic diagram of the extraction process of GFCC feature parameters provided by an embodiment of the present invention;

[0072] Figure 7 A flow chart for calculating weight coefficients provided in an embodiment of the present invention;

[0073] Figure 8 This is a structural block diagram of an identity recognition device based on voiceprint features provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0074] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0075] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0076] As can be seen from the background technology, voiceprint recognition, as a contactless technology for financial identity authentication, is convenient and improves security. However, its recognition accuracy is currently limited due to insufficient audio feature extraction and lack of effective processing of noisy environmental noise.

[0077] Therefore, an embodiment of the present invention provides an identity recognition method and device based on voiceprint features, which pre-processes the acquired voice data to obtain pre-processed audio; uses an audio separation model and a residual denoising diffusion model to perform denoising on the pre-processed audio to obtain denoised audio; extracts acoustic features from the denoised audio, analyzes the dimensions of the acoustic features under different noises, and obtains weight coefficients for each dimension; performs weighted processing on the dimensions of the acoustic features based on the weight coefficients to obtain a feature matrix; performs feature comparison based on the feature matrix and the feature classification model to determine the user identity corresponding to the voice data. The audio separation model and the residual denoising diffusion model perform voice denoising, remove environmental noise, improve the clarity of the voice signal, and enhance the robustness of the system. The feature classification model is used to separate the audio of multiple users, effectively reducing interference, thereby improving the voiceprint recognition rate. At the same time, the feature matrix also significantly improves the recognition efficiency and accuracy of the model.

[0078] It should be noted that the information (including but not limited to user device information, user personal information) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.

[0079] See also Figure 1 , which shows a flow chart of an identity recognition method based on voiceprint features provided by an embodiment of the present invention, the method comprising:

[0080] Step S101: pre-process the acquired voice data to obtain pre-processed audio.

[0081] In the specific setting process of step S101, voice data of the user waiting for identity recognition is collected. Since the collected voice data often has problems such as blanks and noise, the obtained voice data is preprocessed to obtain preprocessed audio.

[0082] The specific pretreatment process is as follows (process A1 to process A4):

[0083] Process A1: Set the sampling frequency and obtain voice data based on the sampling frequency.

[0084] It should be noted that since the sampling frequency directly affects the extraction and recognition accuracy of voiceprint features, the high-pass filter frequency band fH and the low-pass filter frequency band fL are set to optimize this process.

[0085] Combine Figure 2 The audio sampling process diagram shown in the figure shows that, when implementing process A1, an analog audio signal must first be acquired. This process uses a sampling device to convert the continuous analog signal into a discrete digital signal. Then, sampling is performed based on a predetermined sampling frequency, capturing the amplitude of the analog signal at each sampling moment. These sampled values ​​are then fixed by a holding circuit to form a series of digital values ​​(i.e., a sequence of 0s and 1s).

[0086] It can be understood that these digital sequences represent the amplitude information of the analog signal at different time points, providing discrete basic data for subsequent signal processing and analysis.

[0087] Process A2: performing a frame operation on the voice data and performing endpoint detection on the voice data after the frame operation.

[0088] It should be noted that endpoint detection, also known as voice activity detection (VAD), improves the density and information utilization of voiceprint feature extraction by identifying the valid voice parts in the audio data and removing intermittent silent segments. The endpoint detection process includes sound window framing, short-term zero-crossing rate analysis, and zero-crossing rate breakpoint detection. Specifically, the audio is first framed, then the short-term zero-crossing rate is used to detect the activity of each frame, and finally the zero-crossing rate breakpoint detection method is used to remove useless blanks and intermittent parts in the audio, ensuring that only valid voice information is retained. This process significantly improves the efficiency and accuracy of voiceprint feature extraction.

[0089] In the specific implementation of process A2, the input audio signal (i.e., voice data) is first windowed and framed at a fixed frequency. During the framing process, an overlapping segmentation method is used, whereby adjacent frames have some overlap to ensure smooth transitions and continuity of the voice signal. Next, the short-term energy and zero-crossing rate of each frame are analyzed. By examining these values, the activity level of the voice data is assessed, and the boundaries between active information segments and silent segments are determined.

[0090] Process A3: dividing the speech data after the frame operation into information segment speech data and silence segment speech data according to the detected breakpoints.

[0091] In the specific implementation process A3, based on the boundary between the effective information segment and the silence segment in the detected voice data, endpoint detection is used by combining the detection results of short-time energy and zero-crossing rate to divide the voice data after the framing operation into information segment voice data and silence segment voice data.

[0092] Process A4: Use a high-pass filter to perform pre-enhancement processing on the information segment speech data and the silence segment speech data to obtain pre-processed audio.

[0093] It should be noted that pre-enhancement processing processes the input speech signal to highlight the parts that reflect the speaker's voiceprint characteristics, especially enhancing the high-frequency components in the signal. This process processes the audio signal through a high-pass filter to better extract voiceprint characteristic parameters. This processing specifically targets speech information segments and silent segments, helping to further improve voiceprint feature extraction and ensure that the recognition system can more accurately capture the speaker's unique voiceprint characteristics.

[0094] Among them, the high-pass filter formula is shown in formula (1):

[0095] H(z)=1-a×z -1 (1)

[0096] In formula (1), a represents the enhancement coefficient. This high-pass filter emphasizes the high-frequency components of the signal.

[0097] It can be understood that the specific process of using a high-pass filter to pre-enhance the information segment speech data and the silent segment speech data is as shown in formula (2):

[0098] S2(n)=S(n)-a×S(na)(2)

[0099] In formula (2), S2(n) is the signal after pre-enhancement processing; S(n) is the information segment speech data and the silent segment speech data; a represents the enhancement coefficient; S(na) is the signal after time shifting and the enhancement effect obtained after adjustment, that is, the pre-processed audio.

[0100] Step S102: performing noise reduction processing on the pre-processed audio using the audio separation model and the residual denoising diffusion model to obtain noise-reduced audio.

[0101] It should be noted that the audio separation model is specifically the U-Conformer model. The residual denoising diffusion model is specifically the RDDM model. Among them, the U-Conformer model is a deep learning model that combines the U-Net structure and the Transformer. Its main idea is to use the encoding-decoding architecture of U-Net to extract multi-scale features, while introducing the Transformer module to capture global dependencies. This structure enables the model to effectively separate a single audio signal from the aliased signal. Secondly, the residual denoising diffusion model (RDDM) is a denoising deep learning model based on the diffusion process. The core idea of ​​this model is to gradually remove noise from the signal through the diffusion process, and then restore the clean signal.

[0102] In the specific implementation of step S102, the pre-processed audio is subjected to noise reduction processing by the audio separation model and the residual denoising diffusion model to obtain the noise-reduced audio. Figure 3 The audio separation flow chart of the U-Conformer model shown in the figure, the specific noise reduction process is as follows (process B1 to process B7):

[0103] Process B1: Calculate the short-time Fourier transform features and phase difference features of the preprocessed audio.

[0104] When implementing process B1, the preprocessed audio input (preprocessed audio is an aliased signal) is first subjected to a short-time Fourier transform (STFT) to extract its frequency domain features. Next, the phase difference features of this signal are calculated, including the sine phase difference (sinIPD) and cosine phase difference (cosIPD). These two features are collectively referred to as phase difference features (i.e., IPD features).

[0105] Process B2: The short-time Fourier transform features and phase difference features are normalized and input into the audio separation model to predict the amplitude spectrum mask of each sound source signal.

[0106] When implementing process B2, the short-time Fourier transform features and phase difference features are normalized and used as input to the audio separation model U-Conformer model. The U-Conformer model processes these features and predicts the amplitude spectrum mask (SMM) of each sound source signal.

[0107] Process B3: Determine the short-time Fourier transform feature of each sound source based on the amplitude spectrum mask and the short-time Fourier transform feature.

[0108] When implementing process B3, the STFT features (short-time Fourier transform features) of each sound source are calculated using the estimated amplitude spectrum mask (SMM) and the STFT features (short-time Fourier transform features) of the reference channel.

[0109] Process B4: Perform inverse Fourier transform on the short-time Fourier transform characteristics of each sound source to obtain the time domain signal of each sound source.

[0110] When implementing process B4, the frequency domain characteristics (i.e., short-time Fourier transform characteristics) of each sound source are converted back to time domain signals using inverse Fourier transform (ISTFT), thereby restoring each independent sound source signal.

[0111] It can be understood that process B1 to process B4 effectively separates each sound source from the aliased signal, providing a clearer signal for subsequent tasks such as speech recognition or voiceprint recognition.

[0112] Step B5: Draw a spectrogram based on the time domain signal of each sound source.

[0113] Combine Figure 4 The content shown in the RDDM model audio noise reduction flowchart is as follows: first, a spectrogram (time-frequency representation) is drawn based on the time domain signal of each input sound source, and the audio signal amplitude is standardized.

[0114] Process B6: Use the residual denoising diffusion model to perform noise reduction on the spectrum graph to obtain a processed spectrum graph.

[0115] When implementing process B6, the residual denoising diffusion model estimates the noise component in the spectrogram through multiple diffusion steps, predicts and subtracts the noise residual in each denoising step, and uses conditional guidance to ensure that the speech content is preserved to obtain the processed spectrogram.

[0116] The denoising process includes normalization, which normalizes the audio samples to the range [-1, 1]. It also includes STFT and logarithmic magnitude, which involves dividing the audio into overlapping short-time frames. Each frame is Fourier transformed to obtain a complex spectrum (amplitude + phase). The logarithm of the magnitude spectrum is taken to compress the dynamic range. A multi-step iterative process, called the residual denoising diffusion model, gradually removes noise. The model predicts the noise residual and generates a clean magnitude spectrum through an inverse diffusion process. The denoised logarithmic magnitude spectrum is output, which is the processed spectrogram.

[0117] Process B7: Convert the processed spectrogram back to a time domain signal to obtain a noise-reduced frequency spectrum.

[0118] When implementing process B7, the processed spectrogram is first restored to its original logarithmic amplitude range through inverse normalization and exponential transformation. To restore the linear amplitude, an exponential inverse operation is performed. Next, a phase multiplexing method is used to directly use the phase information of the original noisy frequency spectrum without modification.

[0119] The processed spectrogram is then converted back to the time domain using an inverse short-time Fourier transform (ISTFT) to restore the audio. To prevent numerical overflow, the amplitude is clipped during the restoration process to ensure the stability of the calculation results. Ultimately, after these processes, the resulting audio signal is denoised.

[0120] Step S103: extracting acoustic features from the noise reduction audio, analyzing the dimensions of the acoustic features under different noises, and obtaining weight coefficients of the dimensions.

[0121] It should be noted that currently commonly used acoustic layer features include Mel filterbank (MFBank), Mel-frequency cepstral coefficients (MFCC), Gammatone filter cepstral coefficients (GFCC), and shifted cepstrum (SDC). In practical applications, these features can effectively improve the performance of speech recognition systems.

[0122] It's understandable that voiceprint feature extraction is an essential step in achieving identity recognition from voice data. By extracting the user's voiceprint, it can be converted into characteristic parameters that machines can distinguish. Both MFCC and GFCC parameters are designed based on the characteristics of human hearing, offering unique advantages in feature extraction and effectively improving the accuracy of voiceprint recognition.

[0123] In the specific implementation of step S103 , Mel-frequency cepstral coefficient features and Gammatone frequency cepstral coefficient features are extracted from the noise reduction audio, and each dimension of the acoustic features under different noises is analyzed to obtain a weight coefficient of each dimension.

[0124] The specific extraction and analysis process is as follows (Process C1 to Process C6):

[0125] Process C1: Extract Mel-frequency cepstral coefficient features from the denoised audio using a triangular filter.

[0126] It should be noted that the relationship between frequency and Mel frequency is expressed by the Mel scale formula, which is formula (3):

[0127] Mel(f)=2595lg(1+f / 700)(3)

[0128] In formula (3), f represents the frequency of the noise reduction frequency (in Hertz), and Mel(f) is the corresponding Mel frequency. The Mel frequency scale simulates the human ear's frequency perception. It increases the frequency resolution in the low-frequency region while reducing the resolution in the high-frequency region, which is more consistent with the human ear's auditory characteristics.

[0129] It's important to note that the Mel-delta filter in spectral signal processing filters the spectrum to extract features at the Mel scale, namely the Mel-frequency cepstral coefficients (MFCCs). These coefficients represent the characteristics of the speech signal at the Mel frequency level. MFCC features play a crucial role in audio processing. By analyzing these coefficients, the system can effectively extract the acoustic features of speech, turning them into characteristic parameters that can be understood by machines and used for further analysis.

[0130] See for example Figure 5 The schematic diagram of the MFCC feature parameter extraction process is shown. The MFCC feature parameter extraction process includes: first, inputting the noise reduction frequency, performing a fast Fourier transform (FFT) on it to obtain frequency domain information, and obtaining the energy distribution X(i,k) on the spectrum; then taking the square of its modulus to obtain the spectral line energy E(i,k). Then, using the Mel filter bank, the spectrum is converted to the Mel frequency scale that conforms to the human hearing characteristics. After that, the energy in the Mel filter is calculated, and the logarithmic energy S(m) of each filter output is taken. Finally, the discrete cosine transform (DCT) is applied to extract the low-dimensional cepstral coefficients to obtain the final MFCC feature parameters. The specific formula is as follows (Formula (4)):

[0131] (4)

[0132] In formula (4), Represents the feature parameter, M represents the number of filters, n represents the order of the MFCC coefficient, and m represents the index of the Mel filter bank, with a value of 1<=m<=M.

[0133] These parameters effectively reflect the acoustic characteristics of speech and are widely used in fields such as speech recognition and voiceprint recognition.

[0134] Process C2: Gammatone frequency cepstral coefficient features are extracted from the noise reduction frequency through a gammatone filter.

[0135] It should be noted that the Gammatone filter is a filter based on the cochlear structure, and its time domain expression is as follows (Formula (5)):

[0136] (5)

[0137] In formula (5), ; ; A is the gain of the filter; is the center frequency of the filter; is a step function; is the offset phase (usually taken as 0); n is the order of the filter; N is the number of filters; is the attenuation factor, which determines the attenuation speed of the filter's impulse response. The relationship between the attenuation factor and the center frequency is as follows (Formula (6)):

[0138] (6)

[0139] In formula (6), is the equivalent rectangular bandwidth.

[0140] In addition, the equivalent rectangular bandwidth is related to the center frequency The relationship can also be expressed as follows (Formula (7)):

[0141] (7)

[0142] In formula (7), is the center frequency of the filter; is the equivalent rectangular bandwidth.

[0143] See also Figure 6 The figure shows the process of extracting GFCC feature parameters. The steps of GFCC (Gammatone Cepstral Coefficient) parameter extraction include: first, filtering the signal with a gammatone filter, and then calculating the spectrum of the filtered signal. Next, logarithmic processing is performed to compress the dynamic range, and discrete cosine transform (DCT) is applied to extract low-dimensional cepstral coefficients, finally obtaining the GFCC feature parameters.

[0144] These parameters have important applications in speech recognition and voiceprint recognition.

[0145] Process C3: Fusing the Mel-frequency cepstral coefficient features and the gamma-tone frequency cepstral coefficient features to obtain acoustic features.

[0146] It should be noted that in order to improve the noise resistance of the feature parameters, the Mel-frequency cepstral coefficient feature and the gamma-tone frequency cepstral coefficient feature are fused to obtain the acoustic feature, as shown in formula (8).

[0147] M mix = [(C1,C2,...,C n ),(G1,G2,...,G n )](8)

[0148] In formula (8), (C1, C2, ..., C n ) are the coefficients of MFCC feature parameters, (G1,G2,...,G n ) are the coefficients of the GFCC feature parameters. By combining these two feature parameters, we can leverage their respective strengths and enhance the robustness of the feature, especially its recognition capability in noisy environments. This hybrid approach can, to a certain extent, reduce the sensitivity of a single feature parameter to noise, thereby improving the system's anti-interference ability and recognition accuracy.

[0149] Process C4: Perform extreme value normalization on the acoustic features, calculate the intra-class sum of square differences of each dimensional feature of the acoustic features under different noise conditions, and calculate the inter-class sum of square differences of each dimensional feature of the acoustic features under different noise conditions.

[0150] Combine Figure 7 The weight coefficient calculation flow chart shown in the figure shows that in the specific implementation process C4, the contribution of the acoustic features is first evaluated by the addition and subtraction method. Unnecessary dimensional components are removed according to the contribution, thereby extracting the mixed audio features. The average contribution function of the addition and subtraction method is as follows (such as formula (9)):

[0151] (9)

[0152] In formula (9), R(i) represents the contribution; p(i,j) represents the recognition rate when the i-th to j-th order are used as speech feature parameters in the identity recognition system.

[0153] Then, the extracted mixed audio features are subjected to extreme value normalization, that is, Min-Max normalization is performed on the column vector, and the feature values ​​are mapped to the [0, 1] interval (as shown in the following formula (10)).

[0154] (10)

[0155] In formula (10), represents the standardized eigenvalue; The feature value representing the original mixed audio; Represents the original value of the kth feature dimension in the jth frame; k represents the temporary index of the feature; F represents the total number of feature dimensions; M represents the total number of audio frames; i represents the feature dimension index to be standardized; j represents the frame index.

[0156] After extreme value normalization, the features of each dimension are summed in the frame direction, and the mean of each dimensional feature for a specific user is calculated, thereby calculating the intra-class sum of square differences of each dimensional feature under different noises; at the same time, the mean of each dimensional feature for multiple users is calculated, thereby calculating the inter-class sum of square differences of each dimensional feature under different noises.

[0157] Process C5: Divide the sum of square differences within each class of features of each dimension by the sum of square differences between each class of features of each dimension to obtain the initial weight coefficient.

[0158] In the specific implementation process C5, the sum of square differences within each class of the features of each dimension is divided by the sum of square differences between each class of the features of each dimension, and the F ratio is weighted to obtain the initial weight coefficient.

[0159] Process C6: Standardize the initial weight coefficients to obtain the weight coefficients of each dimension.

[0160] In the specific implementation process C6, the initial weight coefficient is subjected to Z-Score normalization processing to obtain the weight coefficient of each dimension.

[0161] Step S104: performing weighted processing on each dimension of the acoustic feature based on the weight coefficient to obtain a feature matrix.

[0162] In the specific implementation of step S104, the dimensions of the acoustic features are sequentially concatenated according to the obtained F-ratio weighting coefficient to obtain the final F-GFCC feature matrix, as shown in formula (11).

[0163] (11)

[0164] In formula (11), H represents the F-ratio weighted feature matrix after weighted fusion, with a dimension of FxM; M is the total number of audio frames; represents the weighted acoustic feature vector obtained in the previous step, and represents the feature vector of the i-th frame audio after weighting by the F ratio, where the value of i is 1<=i<=M.

[0165] Step S105: performing feature comparison based on the feature matrix and the feature classification model to determine the user identity corresponding to the voice data.

[0166] In the specific implementation of step S105, the feature matrix is ​​first input into the residual diffusion denoising model for denoising; then the feature matrix after denoising is input into the feature classification model to extract the voiceprint features; finally, the voiceprint features are calculated and compared with the reserved user voice features, and the user identity corresponding to the voice data is determined based on the comparison results.

[0167] It should be noted that the feature classification model (ResNet50 model) is pre-trained based on sample feature parameters.

[0168] ResNet uses the Rectified Linear Unit (ReLU) activation function. The ResNet50 network architecture lacks a fully connected layer at the top, instead randomly initializing the network's weights. The network input shape is (128, None, 1), a single-channel tensor with a width of 128 and a variable height. Finally, the network implements feature classification by adding fully connected and dense layers. Adam is used for optimization, which dynamically adjusts machine learning parameters by estimating the first and second moments of the gradient. After the network is constructed, ResNet50 training begins to obtain a feature classification model.

[0169] It's understandable that when loading the model, the layer before the classification layer is loaded to obtain voice feature data. Next, sufficient features are extracted from the feature matrix after noise reduction processing and compared with the existing features in the voiceprint library (i.e., the reserved user voice features) to ensure their representativeness and effectiveness.

[0170] Specifically, for the speech input by different users, respective features are extracted and then similarity calculation is performed.

[0171] The similarity is calculated using the cosine similarity method. The closer the cosine value is to 1, the more similar the two feature vectors are. By taking the dot product of the two feature vectors and calculating their cross product, we can obtain the cosine similarity of the two feature matrices, which is the similarity between the two user voices.

[0172] In specific applications, after obtaining the similarity of the input voices, they are compared with a preset threshold to confirm whether the input voices come from the same user. For example, if the confidence threshold is set to 0.8, if the similarity is higher than this threshold, the input voices are considered to come from the same user, thus confirming the user identity corresponding to the voice data.

[0173] In an embodiment of the present invention, voice data is converted into a format suitable for RDDM model processing through operations such as normalization, short-time Fourier transform (STFT) and logarithmic amplitude, so as to cope with various voice interferences in the user's noisy environment. The RDDM model is innovatively applied to perform voice noise reduction, remove environmental noise, improve the clarity of the voice signal, and enhance the robustness of the system. In addition, the U-Conformer model is used to separate the audio of multiple users, effectively reducing interference from other users, thereby improving the voiceprint recognition rate. Compared with the traditional MFCC feature extraction method, the present invention combines GFCC spectrum information and adopts the subtraction method and F-ratio weighted processing to construct a new feature vector, which significantly improves the recognition efficiency and accuracy of the model. Through these innovative technologies, the present invention can significantly improve the user's voiceprint recognition performance in a noisy environment.

[0174] Corresponding to the above embodiment of the present invention, a method for identifying an individual based on voiceprint features is provided. Figure 8 , shows a structural block diagram of an identity recognition device based on voiceprint features provided by an embodiment of the present invention.

[0175] The device includes: a pre-processing unit 801 , a noise reduction unit 802 , an analysis unit 803 , a weighted processing unit 804 and a comparison unit 805 .

[0176] The preprocessing unit 801 is used to preprocess the acquired voice data to obtain preprocessed audio.

[0177] The noise reduction unit 802 is configured to perform noise reduction processing on the pre-processed audio using an audio separation model and a residual denoising diffusion model to obtain a noise-reduced audio.

[0178] The analyzing unit 803 is configured to extract acoustic features from the noise reduction audio, analyze the dimensions of the acoustic features under different noise conditions, and obtain a weight coefficient for each dimension.

[0179] The weighting processing unit 804 is configured to perform weighting processing on each dimension of the acoustic feature based on a weight coefficient to obtain a feature matrix.

[0180] The comparison unit 805 is used to perform feature comparison based on the feature matrix and the feature classification model to determine the user identity corresponding to the voice data.

[0181] The comparison unit 805 is specifically used to: input the feature matrix into the residual diffusion denoising model for denoising; input the feature matrix after denoising into the feature classification model to extract voiceprint features; calculate the voiceprint features and compare them with the reserved user voice features, and determine the user identity corresponding to the voice data based on the comparison results.

[0182] In an embodiment of the present invention, voice data is converted into a format suitable for RDDM model processing through operations such as normalization, short-time Fourier transform (STFT) and logarithmic amplitude, so as to cope with various voice interferences in the user's noisy environment. The RDDM model is innovatively applied to perform voice noise reduction, remove environmental noise, improve the clarity of the voice signal, and enhance the robustness of the system. In addition, the U-Conformer model is used to separate the audio of multiple users, effectively reducing interference from other users, thereby improving the voiceprint recognition rate. Compared with the traditional MFCC feature extraction method, the present invention combines GFCC spectrum information and adopts the subtraction method and F-ratio weighted processing to construct a new feature vector, which significantly improves the recognition efficiency and accuracy of the model. Through these innovative technologies, the present invention can significantly improve the user's voiceprint recognition performance in a noisy environment.

[0183] Combine Figure 8 The content shown, the pre-processing unit 801, includes: a setting module, a detection module, a division module and a pre-processing module.

[0184] The setting module is used to set the sampling frequency and obtain voice data based on the sampling frequency.

[0185] The detection module is used to perform a framing operation on the voice data and perform endpoint detection on the voice data after the framing operation.

[0186] The division module is used to divide the voice data after the frame operation into information segment voice data and silence segment voice data according to the detected breakpoints.

[0187] The preprocessing module is used to perform pre-enhancement processing on the information segment voice data and the silence segment voice data using a high-pass filter to obtain preprocessed audio.

[0188] Combine Figure 8 The content shown, the noise reduction unit 802, includes: a first calculation module, a prediction module, a determination module, a transformation module, a drawing module, a noise reduction processing module and a conversion module.

[0189] The first calculation module is used to calculate the short-time Fourier transform characteristics and phase difference characteristics of the preprocessed audio.

[0190] The prediction module is used to normalize the short-time Fourier transform features and phase difference features, input them into the audio separation model, and predict the amplitude spectrum mask of each sound source signal.

[0191] The determination module is used to determine the short-time Fourier transform feature of each sound source based on the amplitude spectrum mask and the short-time Fourier transform feature.

[0192] The transformation module is used to perform inverse Fourier transform on the short-time Fourier transform characteristics of each sound source to obtain the time domain signal of each sound source.

[0193] The drawing module is used to draw a spectrum diagram based on the time domain signal of each sound source.

[0194] The noise reduction processing module is used to perform noise reduction processing on the spectrum graph using the residual denoising diffusion model to obtain a processed spectrum graph.

[0195] The conversion module is used to convert the processed spectrum graph back into a time domain signal to obtain a noise-reduced frequency signal.

[0196] Combine Figure 8 The content shown, the analysis unit 803, includes: a first extraction module, a second extraction module, a fusion module, a second calculation module, a third calculation module and a standardization processing module.

[0197] The first extraction module is used to extract Mel-frequency cepstral coefficient features from the noise reduction audio frequency using a triangular filter.

[0198] The second extraction module is used to extract gamma-tone frequency cepstral coefficient features from the noise reduction frequency through a gamma-tone filter.

[0199] The fusion module is used to fuse the Mel-frequency cepstral coefficient features and the gamma-tone frequency cepstral coefficient features to obtain acoustic features.

[0200] The second calculation module is used to perform extreme value normalization processing on the acoustic features, calculate the intra-class sum of square differences of each dimensional feature of the acoustic features under different noises, and calculate the inter-class sum of square differences of each dimensional feature of the acoustic features under different noises.

[0201] The third calculation module is used to divide the sum of square differences within each class of features of each dimension by the sum of square differences between each class of features of each dimension to obtain an initial weight coefficient.

[0202] The standardization processing module is used to standardize the initial weight coefficients to obtain the weight coefficients of each dimension.

[0203] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0204] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0205] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for identity recognition based on voiceprint features, characterized in that: The method comprises: Preprocessing the acquired voice data to obtain preprocessed audio; Performing noise reduction processing on the preprocessed audio using an audio separation model and a residual denoising diffusion model to obtain a noise-reduced audio; Extracting acoustic features from the noise reduction frequency, analyzing dimensions of the acoustic features under different noise conditions, and obtaining weight coefficients for the dimensions; Performing weighted processing on each dimension of the acoustic feature based on the weight coefficient to obtain a feature matrix; A feature comparison is performed based on the feature matrix and the feature classification model to determine the user identity corresponding to the voice data.

2. The method according to claim 1, characterized in that The preprocessing of the acquired voice data to obtain preprocessed audio includes: Setting a sampling frequency, and acquiring voice data based on the sampling frequency; Performing a framing operation on the voice data, and performing endpoint detection on the voice data after the framing operation; Dividing the voice data after the framing operation into information segment voice data and silence segment voice data according to the detected breakpoints; The information segment voice data and the silence segment voice data are pre-enhanced using a high-pass filter to obtain pre-processed audio.

3. The method according to claim 1, characterized in that The method of performing noise reduction processing on the pre-processed audio using the audio separation model and the residual denoising diffusion model to obtain the noise-reduced audio includes: Calculating short-time Fourier transform features and phase difference features of the preprocessed audio; Normalizing the short-time Fourier transform features and the phase difference features, and inputting the results into an audio separation model to predict an amplitude spectrum mask of each sound source signal; Determining a short-time Fourier transform feature of each sound source based on the amplitude spectrum mask and the short-time Fourier transform feature; Perform inverse Fourier transform on the short-time Fourier transform characteristics of each sound source to obtain the time domain signal of each sound source; plotting a spectrogram based on the time domain signal of each sound source; Performing denoising processing on the spectrum graph using a residual denoising diffusion model to obtain a processed spectrum graph; Convert the processed spectrogram back to the time domain signal to obtain the noise-reduced frequency spectrum.

4. The method according to claim 1, wherein The extracting of acoustic features from the noise reduction frequency, analyzing each dimension of the acoustic features under different noise conditions, and obtaining a weight coefficient of each dimension includes: Extracting Mel-frequency cepstral coefficient features from the noise reduction frequency using a triangular filter; extracting gamma-tone frequency cepstral coefficient features from the noise reduction frequency using a gamma-tone filter; Fusing the Mel-frequency cepstral coefficient feature and the gamma-tone frequency cepstral coefficient feature to obtain an acoustic feature; Performing extreme value normalization processing on the acoustic features, calculating the intra-class sum of square differences of each dimensional feature of the acoustic features under different noise conditions, and calculating the inter-class sum of square differences of each dimensional feature of the acoustic features under different noise conditions; Divide the sum of square differences within each class of the features of each dimension by the sum of square differences between each class of the features of each dimension to obtain an initial weight coefficient; The initial weight coefficients are normalized to obtain weight coefficients of each dimension.

5. The method according to claim 1, wherein The performing feature comparison based on the feature matrix and the feature classification model to determine the user identity corresponding to the voice data includes: Inputting the characteristic matrix into the residual diffusion denoising model for denoising; The feature matrix after noise reduction is input into the feature classification model to extract the voiceprint features; The voiceprint feature is calculated and compared with the reserved user voice feature, and the user identity corresponding to the voice data is determined based on the comparison result.

6. An identity recognition device based on voiceprint features, characterized in that: The device comprises: A preprocessing unit, configured to preprocess the acquired voice data to obtain preprocessed audio; a noise reduction unit, configured to perform noise reduction processing on the preprocessed audio using an audio separation model and a residual denoising diffusion model to obtain a noise-reduced audio; an analysis unit, configured to extract acoustic features from the noise reduction frequency, analyze each dimension of the acoustic features under different noise conditions, and obtain a weight coefficient for each dimension; a weighting processing unit, configured to perform weighted processing on each dimension of the acoustic feature based on the weight coefficient to obtain a feature matrix; The comparison unit is used to perform feature comparison based on the feature matrix and the feature classification model to determine the user identity corresponding to the voice data.

7. The device according to claim 6, characterized in that The pre-processing unit comprises: A setting module, configured to set a sampling frequency and obtain voice data based on the sampling frequency; A detection module, configured to perform a framing operation on the voice data and perform endpoint detection on the voice data after the framing operation; A division module, configured to divide the voice data after the framing operation into information segment voice data and silence segment voice data according to the detected breakpoints; The preprocessing module is used to perform pre-enhancement processing on the information segment voice data and the silence segment voice data using a high-pass filter to obtain preprocessed audio.

8. The device according to claim 6, characterized in that The noise reduction unit comprises: A first calculation module, configured to calculate a short-time Fourier transform feature and a phase difference feature of the preprocessed audio; A prediction module, configured to normalize the short-time Fourier transform features and the phase difference features, input the normalized features into an audio separation model, and predict an amplitude spectrum mask for each sound source signal; a determination module, configured to determine a short-time Fourier transform feature of each sound source based on the amplitude spectrum mask and the short-time Fourier transform feature; A transformation module is used to perform inverse Fourier transform on the short-time Fourier transform characteristics of each sound source to obtain the time domain signal of each sound source; A drawing module, configured to draw a spectrum diagram based on the time domain signal of each sound source; A noise reduction processing module, configured to perform noise reduction processing on the spectrum graph using a residual denoising diffusion model to obtain a processed spectrum graph; The conversion module is used to convert the processed spectrum graph back into a time domain signal to obtain a noise-reduced frequency signal.

9. The device according to claim 6, characterized in that The analysis unit comprises: A first extraction module is used to extract Mel-frequency cepstral coefficient features from the noise reduction frequency using a triangular filter; A second extraction module is configured to extract gamma-tone frequency cepstral coefficient features from the noise reduction frequency using a gamma-tone filter; A fusion module, configured to fuse the Mel-frequency cepstral coefficient feature and the gamma-tone frequency cepstral coefficient feature to obtain an acoustic feature; a second calculation module, configured to perform extreme value normalization processing on the acoustic features, calculate the within-class sum of square differences of each dimensional feature of the acoustic features under different noise conditions, and calculate the between-class sum of square differences of each dimensional feature of the acoustic features under different noise conditions; A third calculation module is used to divide the intra-class sum of square differences of the features of each dimension by the inter-class sum of square differences of the features of each dimension to obtain an initial weight coefficient; The normalization processing module is used to perform normalization processing on the initial weight coefficient to obtain the weight coefficient of each dimension.

10. The device according to claim 6, characterized in that The comparison unit is specifically configured to: input the feature matrix into a residual diffusion denoising model for denoising; input the denoised feature matrix into a feature classification model to extract voiceprint features; The voiceprint feature is calculated and compared with the reserved user voice feature, and the user identity corresponding to the voice data is determined based on the comparison result.