Voiceprint recognition method and device based on multiband analysis

The voiceprint recognition method, which employs multi-band analysis and dynamic noise suppression, solves the problems of complex noise and low recognition accuracy across devices, enabling low-latency real-time identity authentication on embedded devices and improving the robustness and cross-device adaptability of voiceprint recognition.

CN120932653APending Publication Date: 2025-11-11MINAMI ACOUSTICS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510877113.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing voiceprint recognition technologies have low recognition accuracy in complex noisy environments and cross-device scenarios, and it is difficult to achieve low-latency real-time processing in embedded devices, resulting in insufficient environmental robustness, feature discrimination and cross-device generalization ability.

Method used

A multi-band analysis method is adopted, which dynamically divides the frequency bands through a learnable filter bank. Combined with an attention fusion network and a contrastive training strategy, multimodal feature extraction and model training are performed. Streaming processing and edge optimization are used for deployment to achieve dynamic noise suppression and cross-device robustness.

Benefits of technology

It improves the accuracy of voiceprint recognition in noisy environments and cross-device recognition precision, reduces the false recognition rate, and achieves low-latency real-time identity authentication on low-resource devices, supporting scenarios such as smart homes and mobile payments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932653A_ABST
    Figure CN120932653A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice recognition, in particular to a voiceprint recognition method and device based on multiband analysis, and the method comprises the steps: data preparation and preprocessing, dynamic band division and feature extraction, model training and optimization, real-time reasoning and deployment, and evaluation and iteration. Compared with the prior art that voiceprint feature extraction is carried out by adopting fixed frequency band division, frequency band division is rigid, and a complex noise environment and cross-equipment frequency response difference cannot be adapted, the scheme dynamically optimizes the frequency band center frequency and bandwidth through a learnable filter bank, and improves the voiceprint feature extraction efficiency. In training, a frequency band (such as a fundamental frequency harmonic wave and a formant region) with strong adaptive focusing discrimination is reversely propagated in combination with a loss function, and meanwhile, a frequency band attention mechanism is introduced to suppress low signal-to-noise ratio sub-band interference, so that the error rate of voiceprint recognition in a noise environment is reduced, the cross-device scene recognition precision is improved, and the recognition efficiency is improved. And the robustness of a complex scene is obviously enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice recognition technology, and in particular to a voiceprint recognition method and apparatus based on multi-band analysis. Background Technology

[0002] In existing voiceprint recognition technologies, the feature extraction process typically relies on a pre-defined fixed frequency band division method, such as a filter bank based on the Mel scale or a uniform distribution.

[0003] While these methods perform reasonably well in quiet environments, in complex noisy scenarios, fixed frequency bands struggle to dynamically adapt to varying noise distributions and differences in device frequency response characteristics, leading to a significant decrease in recognition accuracy. For example, low-frequency environmental noise (such as air conditioner hum) and high-frequency channel distortion (such as mobile phone compression) can interfere with the effectiveness of fixed frequency band features. Furthermore, traditional methods often employ single frequency domain features (such as MFCC or LPC), lacking joint modeling of temporal speech dynamics and phase details, resulting in a high false recognition rate in scenarios with similar speakers or short speech.

[0004] Meanwhile, existing technologies lack effective solutions to the problem of cross-device voiceprint offset. The distribution of speech features collected from the same speaker on different devices (such as mobile phones and microphones) varies greatly, limiting the generalization ability of the model. On the other hand, due to model complexity and computational resource limitations, existing methods are difficult to achieve low-latency real-time processing in embedded devices, which restricts the application of voiceprint recognition technology in low-power scenarios such as smart homes and mobile payments.

[0005] These issues collectively constitute the bottlenecks of existing voiceprint recognition technology in terms of environmental robustness, feature distinguishability, cross-device generalization, and deployment efficiency, and urgently require breakthroughs through technological innovation. Summary of the Invention

[0006] To overcome the problems mentioned in the background art, the present invention proposes a voiceprint recognition method and device based on multi-band analysis.

[0007] The technical solution of this invention is: a voiceprint recognition method based on multi-band analysis, comprising the following steps:

[0008] S11: Data preparation and preprocessing, optimizing speech quality through standardized input and dynamic noise reduction;

[0009] S12: Dynamic frequency band division and feature extraction, using a learnable filter bank to dynamically divide frequency bands and extract multimodal features;

[0010] S13: Model training and optimization, using attention fusion networks and contrastive training strategies to train robust models;

[0011] S14: Real-time inference and deployment, using streaming processing and edge optimization for model deployment;

[0012] S15: Evaluation and iteration, based on online feedback and continuous learning model adaptive iterative updates.

[0013] As a preferred method, when optimizing speech quality through standardized input and dynamic noise reduction, the specific methods include:

[0014] S21: Unified audio format, resampling the original speech to 16kHz, adjusting the quantization precision to 16bit, and compatible with real-time streaming input and offline audio files;

[0015] S22: Automatic gain control, which balances the amplitude of speech through dynamic range compression algorithm, so that the peak amplitude of all input speech is within the range of [-1,1];

[0016] S23: Speech activity detection, based on the WebRTC open source model, detects effective speech segments, marks silent segments and noise segments and truncates them, retaining only the human voice activity range;

[0017] S24: Dynamic noise suppression. First, the noise power spectrum is estimated according to the frequency band, and the spectral reduction coefficient is dynamically adjusted according to the signal-to-noise ratio of each frequency band. For the low frequency band, aggressive noise reduction is adopted with a spectral reduction coefficient of 3, and for the high frequency band, conservative processing is adopted with a spectral reduction coefficient of 1.

[0018] Preferably, when using a learnable filter bank to dynamically divide frequency bands and extract multimodal features, the specific steps include:

[0019] S31: Initialize the learnable filter, preset 16 sub-band filters based on Mel scale or uniform scale, set the initial parameters to traditional empirical values, obtain the filter bank, and implement the filter bank as a neural network layer;

[0020] S32: Dynamic frequency band division optimization generates 16 sub-band time-frequency signals from the pre-processed speech input filter bank, and updates the filter center frequency and bandwidth through backpropagation of the loss function during the training process, so that the sub-band distribution automatically focuses on the frequency band with strong discrimination.

[0021] S33: Parallel extraction of multimodal features, extracting multimodal features from the original signal respectively;

[0022] S34: Sub-band noise reduction enhancement, which uses the channel attention module to perform noise reduction weighting on the frequency domain features of each sub-band;

[0023] S35: Feature fusion and output. The obtained multimodal features are concatenated across modalities and processed by bidirectional LSTM to output speaker embedding vectors containing contextual information.

[0024] Preferably, when performing parallel extraction of multimodal features, the specific steps include:

[0025] A11: Frequency domain feature extraction. Calculate the improved MFCC for each sub-band spectrum and count the sub-band energy entropy to obtain 40-dimensional features for each sub-band. The improved MFCC includes 13-dimensional MFCC + first-order difference + second-order difference + entropy.

[0026] A12: Temporal feature extraction. Input the original waveform into a RawNet3 network containing SincNet convolutional layers to extract frame-level temporal features.

[0027] A13: Phase feature extraction, calculate the instantaneous phase derivative from the complex spectrum of each sub-band to generate phase change mode features.

[0028] Preferably, when performing noise-resistant weighting on the frequency domain features of each sub-band through the channel attention module, the channel attention module adds a compression-excitation module after the frequency domain features of each sub-band. The compression-excitation module includes global average pooling, fully connected layer learning weights, and feature recalibration. Global average pooling is used to compress the sub-band features into channel description vectors, fully connected layer learning weights is used to generate noise suppression coefficients for each frequency band, and feature recalibration is used to dynamically scale the sub-band features according to the weights to suppress low signal-to-noise ratio frequency bands.

[0029] Preferably, when 16 sub-band filters are preset based on the Mel-scale or uniform-scale, and the initial parameters are set to traditional empirical values ​​to obtain a filter bank, and the filter bank is implemented as a neural network layer, the response formula of each sub-band filter is:

[0030]

[0031] Among them, H m (f k ) represents the m-th filter at frequency f. k The response value at f c,m Let B be the center frequency of the m-th filter. m f is the bandwidth of the m-th filter. k These are discrete frequency points.

[0032] Preferably, when training a robust model using an attention fusion network and a contrastive training strategy, the specific steps include:

[0033] S41: Attention mechanism fusion, input the multi-subband feature matrix into the multi-head attention layer, calculate the correlation weight between different subbands, and dynamically weight and fuse the subband features through attention weights to generate a global context-aware joint representation;

[0034] S42: Bidirectional LSTM temporal modeling, inputting the attention-fused features into the bidirectional LSTM network to capture long-term dependencies between speech frames, and taking the hidden state of the last time step of the LSTM or performing average pooling on all time steps to obtain the speaker embedding vector.

[0035] S43: Contrastive learning constraints: calculate cosine similarity for different sub-band features of the same speaker to maximize similarity; minimize similarity for features of different speakers and force features of different frequency bands to be close together in the embedding space to enhance the robustness of the model to partial frequency band missingness.

[0036] As a preferred approach, when training a robust model using an attention fusion network and a contrastive training strategy, angular spacing is introduced into the classification layer to increase the angular difference between the embedding vectors of different speakers and improve inter-class discriminability. The principle behind this is as follows:

[0037]

[0038] Where e is the natural exponential function, s is the scaling factor, and θ y Let m be the target angle, and cos(θ) be the interval parameter. y +m) is the cosine value after interval adjustment, ∑j≠ye scos θ i This is the sum of scores for non-target classes;

[0039] A domain classifier is added, and the domain features are obfuscated through a gradient inversion layer to reduce cross-device voiceprint offset. The principle formula is as follows:

[0040]

[0041] Where D is the domain classifier, f(x) is the voiceprint feature, and E is the expected value.

[0042] As a preferred method, when training a robust model using an attention fusion network and a contrastive training strategy, it is also possible to randomly select noise samples from a noise library and superimpose them onto clean speech with an SNR of 0-20dB, and randomly mask some subbands with a 30% probability, forcing the model to use the remaining subbands for recognition, thereby enhancing generalization ability.

[0043] Preferably, when deploying models using streaming processing and edge optimization, the specific methods include:

[0044] S51: Streaming processing workflow, the speech stream is divided into 20ms frames, the features of each sub-band are extracted in parallel by multiple threads, and only the most recent 1 second speech frame is cached. Historical data is automatically eliminated, and the speaker feature library is compressed to 64 dimensions using PCA and retrieved in milliseconds through the approximate nearest neighbor algorithm.

[0045] S52: Edge optimization process, including model lightweighting, cross-platform adaptation and low power consumption strategy. Among them, ITN8 is used for model lightweighting and low contribution subbands and neurons are removed. The low power consumption strategy is to use VAD to control activation and adjust the CPU and GPU frequencies according to the load.

[0046] A voiceprint recognition device based on multi-band analysis includes:

[0047] The data preprocessing module is used to unify the voice format, extract the human voice segment using VAD, and dynamically reduce noise to suppress low-frequency noise and high-frequency distortion.

[0048] The dynamic frequency band module is used to automatically optimize the frequency band division, extract multimodal features, and perform noise-resistant weighting through a learnable filter bank;

[0049] The model training module is used to train a cross-device robust voiceprint embedding model by jointly performing attention fusion, LSTM temporal modeling, contrastive learning and noise enhancement.

[0050] The real-time deployment module is used to achieve cross-platform, low-power operation of the model by employing streaming frame processing and lightweight models.

[0051] The iterative optimization module is used to monitor the recognition effect online, dynamically adjust the threshold, incrementally learn new data, and update the frequency band division strategy.

[0052] The beneficial effects of this invention are:

[0053] 1. Compared to existing technologies that use fixed frequency band division for voiceprint feature extraction, which suffer from rigid frequency band division and inability to adapt to complex noise environments and cross-device frequency response differences, this solution dynamically optimizes the center frequency and bandwidth of the frequency band through a learnable filter bank. During training, it adaptively focuses on highly discriminative frequency bands (such as fundamental harmonics and formant regions) by incorporating backpropagation of the loss function. Simultaneously, it introduces a frequency band attention mechanism to suppress low signal-to-noise ratio sub-band interference. This solution reduces the error rate of voiceprint recognition in noisy environments, improves recognition accuracy in cross-device scenarios, and significantly enhances robustness in complex scenarios.

[0054] 2. In contrast to existing technologies that rely on single frequency domain features (such as MFCC or LPC), leading to incomplete speaker recognition, this solution proposes a multimodal feature parallel extraction mechanism. It combines frequency domain features (improved MFCC + sub-band energy entropy), time domain features (RawNet3 waveform modeling), and phase features (instantaneous phase derivative) to comprehensively capture pronunciation details, vocal tract characteristics, and spatiotemporal dynamic changes. By fusing cross-modal features through bidirectional LSTM and multi-head attention, a context-aware speaker embedding vector is generated. This addresses the shortcomings of traditional methods, such as large intra-class differences and insufficient inter-class discriminability due to single features, thus reducing the false recognition rate in scenarios with similar speakers (such as twins).

[0055] 3. Compared to existing technologies that lack joint optimization strategies for cross-device and noise interference, resulting in large performance fluctuations of the voiceprint model across different acquisition devices such as mobile phones and microphones, this solution designs an adversarial training and dynamic noise injection mechanism. It uses a gradient inversion layer to obfuscate the device classifier, forcing enhanced device independence of voiceprint features. Simultaneously, it randomly superimposes real environmental noise with an SNR of 0-20dB and simulates frequency band gaps by masking sub-bands with a 30% probability. This solution reduces the EER of cross-device voiceprint recognition.

[0056] 4. Addressing the challenge of real-time deployment in embedded devices due to the high computational complexity and large memory footprint of existing voiceprint recognition systems, this solution proposes streaming frame-segmentation processing and lightweight edge optimization techniques. It employs 20ms frame-segmentation parallel computation, PCA compression of the embedding vector to 64 dimensions, and combines INT8 quantization and subband pruning to reduce the model size to 25% of its original size. Furthermore, it dynamically controls computational resources based on VAD, achieving power consumption below 0.1W during sleep mode. This solution enables end-to-end latency of less than 100ms and memory usage below 50MB for voiceprint recognition on ARM Cortex-M4 chips, supporting real-time identity authentication in low-resource scenarios such as smart locks and wearable devices, while reducing power consumption compared to traditional solutions. Attached Figure Description

[0057] Figure 1 The diagram shown is a flowchart of the voiceprint recognition method based on multi-band analysis of the present invention.

[0058] Figure 2 The diagram shown is a schematic representation of the structure of the voiceprint recognition device based on multi-band analysis of the present invention. Detailed Implementation

[0059] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0060] Please see Figure 1-2 The present invention provides an embodiment: a voiceprint recognition method and apparatus based on multi-band analysis. The voiceprint recognition method based on multi-band analysis includes the following steps:

[0061] Step 1: Data Preparation and Preprocessing

[0062] Optimizing speech quality through standardized input and dynamic noise reduction specifically includes:

[0063] The audio format is standardized by resampling the original speech to 16kHz and adjusting the quantization precision to 16bit, ensuring compatibility with real-time streaming input and offline audio files. Automatic gain control balances speech amplitude through a dynamic range compression algorithm, keeping the peak amplitude of all input speech within the range of [-1,1]. Speech activity detection detects effective speech segments based on the WebRTC open-source model, marking silent and noisy segments and truncating them, retaining only the active human voice range. Dynamic noise suppression first estimates the noise power spectrum by frequency band and then dynamically adjusts the spectral reduction coefficient based on the signal-to-noise ratio of each frequency band. Aggressive noise reduction with a spectral reduction coefficient of 3 is used for low-frequency bands, while conservative processing with a spectral reduction coefficient of 1 is used for high-frequency bands.

[0064] Step 2: Dynamic frequency band division and feature extraction

[0065] The use of learnable filter banks to dynamically divide frequency bands and extract multimodal features includes:

[0066] The learnable filter is initialized by pre-setting 16 sub-band filters based on Mel-scale or uniform scale, with initial parameters set to traditional empirical values, resulting in a filter bank, which is then implemented as a neural network layer. Dynamic frequency band division optimization involves inputting the preprocessed speech into the filter bank to generate 16 sub-band time-frequency signals. During training, the filter center frequency and bandwidth are updated via backpropagation of the loss function, automatically focusing the sub-band distribution on highly discriminative frequency bands. Parallel multimodal feature extraction is performed, extracting multimodal features from the original signal. Specifically, this includes frequency domain feature extraction, calculating the improved MFCC for each sub-band spectrogram and calculating the sub-band energy entropy to obtain 40-dimensional features for each sub-band. The improved MFCC includes 13-dimensional MFCC + first-order difference + second-order difference. Differential + Entropy; Temporal Feature Extraction: Input the original waveform into a RawNet3 network containing SincNet convolutional layers to extract frame-level temporal features; Phase Feature Extraction: Calculate the instantaneous phase derivative from the complex spectrum of each sub-band to generate phase change pattern features; Sub-band Noise Reduction Enhancement: Apply noise reduction weights to the frequency domain features of each sub-band through a channel attention module; Feature Fusion and Output: Perform cross-modal concatenation on the obtained multimodal features and process the frame-level feature sequence through bidirectional LSTM to output a speaker embedding vector containing contextual information. Specifically, when 16 sub-band filters are preset based on Mel-scale or uniform scale, with initial parameters set to traditional empirical values, and the filter bank is implemented as a neural network layer, the response formula for each sub-band filter is:

[0067]

[0068] Among them, H m (f k ) represents the m-th filter at frequency f. k The response value at fc,m Let B be the center frequency of the m-th filter. m f is the bandwidth of the m-th filter. k These are discrete frequency points;

[0069] Step 3: Model Training and Optimization

[0070] A robust model is trained using an attention fusion network and a contrastive training strategy, specifically including:

[0071] Attention mechanism fusion involves inputting multi-subband feature matrices into a multi-head attention layer, calculating the association weights between different subbands, and dynamically weighting and fusing subband features through attention weights to generate a globally context-aware joint representation. Bidirectional LSTM temporal modeling involves inputting the attention-fused features into a bidirectional LSTM network to capture long-term dependencies between speech frames, and obtaining the speaker embedding vector by taking the hidden state of the last time step of the LSTM or by performing average pooling over all time steps. Contrastive learning constraints involve calculating cosine similarity for different subband features of the same speaker to maximize similarity, minimizing similarity for features of different speakers, and forcing features of different frequency bands to be close together in the embedding space to enhance the model's robustness to missing frequency bands.

[0072] Step 4: Real-time Inference and Deployment

[0073] Model deployment employs streaming processing and edge optimization, specifically including:

[0074] The streaming processing workflow divides the speech stream into 20ms frames, extracts features from each sub-band in parallel using multiple threads, and caches only the most recent 1-second speech frame, automatically discarding historical data. The speaker feature library is compressed to 64 dimensions using PCA and retrieved in milliseconds using an approximate nearest neighbor algorithm. The edge optimization workflow includes model lightweighting, cross-platform adaptation, and low-power strategies. Specifically, ITN8 is used for model lightweighting, and low-contribution sub-bands and neurons are removed. The low-power strategy uses VAD to control activation and adjusts the CPU and GPU frequencies according to the load.

[0075] Step 5: Evaluation and Iteration

[0076] It is based on online feedback and continuous learning model for adaptive iterative updates.

[0077] Specifically, when applying noise-resistant weighting to the frequency domain features of each sub-band through the channel attention module, the channel attention module adds a compression-activation module after the frequency domain features of each sub-band. The compression-activation module includes global average pooling, fully connected layer learning weights, and feature recalibration. Global average pooling is used to compress the sub-band features into channel description vectors, fully connected layer learning weights is used to generate noise suppression coefficients for each frequency band, and feature recalibration is used to dynamically scale the sub-band features according to the weights to suppress low signal-to-noise ratio frequency bands.

[0078] In training a robust model using an attention fusion network and a contrastive training strategy, angular margin is introduced into the classification layer to increase the angular difference between the embedding vectors of different speakers, thereby improving inter-class discriminability. The underlying principle is as follows:

[0079]

[0080] Where e is the natural exponential function, s is the scaling factor, and θ y Let m be the target angle, and cos(θ) be the interval parameter. y +m) is the cosine value after interval adjustment, ∑j≠ye scos θ i This is the sum of scores for non-target classes;

[0081] A domain classifier is added, and the domain features are obfuscated through a gradient inversion layer to reduce cross-device voiceprint offset. The principle formula is as follows:

[0082]

[0083] Where D is the domain classifier, f(x) is the voiceprint feature, and E is the expected value;

[0084] It also includes randomly selecting noise samples from a noise library and superimposing them onto clean speech with an SNR of 0-20dB, and randomly masking some subbands with a 30% probability, forcing the model to use the remaining subbands for recognition, thereby enhancing generalization ability.

[0085] A voiceprint recognition device based on multi-band analysis includes:

[0086] The data preprocessing module is used to unify the voice format, extract the human voice segment using VAD, and dynamically reduce noise to suppress low-frequency noise and high-frequency distortion.

[0087] The dynamic frequency band module is used to automatically optimize the frequency band division, extract multimodal features, and perform noise-resistant weighting through a learnable filter bank;

[0088] The model training module is used to train a cross-device robust voiceprint embedding model by jointly performing attention fusion, LSTM temporal modeling, contrastive learning and noise enhancement.

[0089] The real-time deployment module is used to achieve cross-platform, low-power operation of the model by employing streaming frame processing and lightweight models.

[0090] The iterative optimization module is used to monitor the recognition effect online, dynamically adjust the threshold, incrementally learn new data, and update the frequency band division strategy.

[0091] Example 1: Smart Home Voiceprint Door Lock System

[0092] Application scenario: Used for voiceprint unlocking of smart door locks, requiring high-precision and low-latency identity authentication in a home environment (including low-frequency noise from air conditioners and high-frequency noise from children).

[0093] Implementation steps

[0094] Data preprocessing:

[0095] Unified audio format: The voice captured by the door lock microphone is uniformly 16kHz / 16bit, supporting offline voice (such as pre-recorded commands) and real-time streaming input. Dynamic noise reduction: In the low-frequency band (0-500Hz), a spectral reduction coefficient β=3 is used to suppress the humming sound of the air conditioner; in the high-frequency band (above 4kHz), β=1 is used to preserve pronunciation details such as hissing sounds. VAD truncation: Based on the WebRTC model, silent segments are removed, retaining only the valid voice segment of the "Open Sesame" command.

[0096] Dynamic frequency band division and feature extraction:

[0097] Learnable filter bank: Initializes 16 Mel-scale filters, and after end-to-end training, the low-frequency bandwidth is reduced to 100Hz (focusing on the fundamental frequency), while the high-frequency bandwidth is extended to 500Hz (capturing sibilance). Multimodal features: Frequency domain: sub-band MFCC + entropy (40-dimensional); Time domain: RawNet3 extracts plosive features (256-dimensional); Phase: IPD captures lip movement differences (64-dimensional). Noise-resistant weighting: The SE module reduces the weight of the low-frequency sub-band to 0.2 to suppress residual air conditioning noise.

[0098] Model training and optimization:

[0099] Attention Fusion: Multi-head attention is used to associate the fundamental frequency (low frequency) and formant (mid frequency) to improve the distinguishability of "father-son voiceprints". Contrastive Learning: The similarity of the embedding vectors of different passwords for the same user is forced to be >0.85, while that for different users is <0.3. Data Augmentation: An infant's cry (SNR=10dB) is superimposed and 30% of the high-frequency subbands are randomly masked to simulate a real environment.

[0100] Real-time deployment:

[0101] Streaming processing: 20ms frame segmentation, multi-threaded feature extraction, response latency <200ms. Lightweight: INT8 quantized model size is 4MB, deployed on a door lock ARM chip (Cortex-A35). Low power consumption: VAD controlled activation, power consumption in sleep mode is 0.1W, battery life of 1 year with AA batteries.

[0102] Iterative optimization: Online updates: When a new user registers, the EWC algorithm is incrementally trained to avoid forgetting features of old users. Dynamic threshold: The similarity threshold is automatically adjusted based on the false recognition rate (FAR) (default 0.75).

[0103] Effect:

[0104] In a home environment, the recognition accuracy (EER) is reduced from 8.2% of the traditional method to 2.5%; the unlocking delay is <0.3 seconds, and the power consumption is reduced by 70%; it supports 50 user voiceprint libraries, and the cross-device (different door lock models) error is <1%.

[0105] Example 2: Mobile Financial Identity Authentication System

[0106] Application scenario: Used for voiceprint login in mobile banking apps, requiring support for cross-device (different mobile phone microphones) and noise interference in public places.

[0107] Implementation steps:

[0108] Data preprocessing:

[0109] Format compatibility: Supports 48kHz voice input from iOS / Android microphones, downsampled to 16kHz. AGC enhancement: Dynamically compresses amplitude to [-1,1] to avoid differences in recording between near and far fields. Noise suppression: In subway scenarios, low-frequency (0-300Hz) β=3 suppresses vibration noise, while high-frequency (above 3kHz) β=1 preserves high-frequency voice details.

[0110] Dynamic frequency band and characteristics:

[0111] Filter optimization: During training, the high-frequency filter is shifted to 8kHz to adapt to the high-frequency attenuation characteristics of mobile phone microphones. Multimodal features: Frequency domain: Subband energy entropy distinguishes dialect differences; Time domain: SincNet captures syllable boundaries; Phase: IPD enhances the recording location invariance.

[0112] Noise reduction weighting: The SE module dynamically increases the weight of subbands with a signal-to-noise ratio >20dB (e.g., intermediate frequency 1-3kHz).

[0113] Model training:

[0114] Domain adversarial training: Add a mobile phone model classifier (e.g., iPhone vs. Huawei), and after gradient inversion, the cross-device EER decreases from 12% to 4%. Angular margin loss: s=32, m=0.5, increasing the angular difference in embedding vectors between different users. Noise injection: Randomly superimpose mall noise (SNR=0-20dB) to enhance robustness in noisy environments.

[0115] Mobile deployment:

[0116] Streaming processing: 20ms frame rate, feature extraction and matching process in under 100ms. Model compression: PCA compresses embedding vectors to 64 dimensions, ANN retrieval library uses less than 10MB of memory. Cross-platform compatibility: Converted to CoreML model on iOS, adapted to NNAPI on Android, 30FPS.

[0117] Continuous learning:

[0118] Security Update: Upon detecting an attack sample (audio replay), all user models are simultaneously updated in the cloud. User Feedback: Users flag misidentified results, triggering local model fine-tuning.

[0119] Effect:

[0120] The cross-phone recognition error rate (EER) has been reduced from 15% to 3.8%; the recognition rate remains >97% in noisy environments (subway / shopping mall); it supports millions of concurrent users, and the server cost for a single authentication is <0.001 CPU cores per second.

[0121] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A voiceprint recognition method based on multi-band analysis, characterized in that: Includes the following steps: S11: Data preparation and preprocessing, optimizing speech quality through standardized input and dynamic noise reduction; S12: Dynamic frequency band division and feature extraction, using a learnable filter bank to dynamically divide frequency bands and extract multimodal features; S13: Model training and optimization, using attention fusion networks and contrastive training strategies to train robust models; S14: Real-time inference and deployment, using streaming processing and edge optimization for model deployment; S15: Evaluation and iteration, based on online feedback and continuous learning model adaptive iterative updates.

2. The voiceprint recognition method based on multi-band analysis according to claim 1, characterized in that: Optimizing speech quality through standardized input and dynamic noise reduction specifically includes: S21: Unified audio format, resampling the original speech to 16kHz, adjusting the quantization precision to 16bit, and compatible with real-time streaming input and offline audio files; S22: Automatic gain control, which balances the amplitude of speech through dynamic range compression algorithm, so that the peak amplitude of all input speech is within the range of [-1,1]; S23: Speech activity detection, based on the WebRTC open source model, detects effective speech segments, marks silent segments and noise segments and truncates them, retaining only the human voice activity range; S24: Dynamic noise suppression. First, the noise power spectrum is estimated according to the frequency band, and the spectral reduction coefficient is dynamically adjusted according to the signal-to-noise ratio of each frequency band. For the low frequency band, aggressive noise reduction is adopted with a spectral reduction coefficient of 3, and for the high frequency band, conservative processing is adopted with a spectral reduction coefficient of 1.

3. The voiceprint recognition method based on multi-band analysis according to claim 2, characterized in that: When using learnable filter banks to dynamically divide frequency bands and extract multimodal features, the specific steps include: S31: Initialize the learnable filter, preset 16 sub-band filters based on Mel scale or uniform scale, set the initial parameters to traditional empirical values, obtain the filter bank, and implement the filter bank as a neural network layer; S32: Dynamic frequency band division optimization generates 16 sub-band time-frequency signals from the pre-processed speech input filter bank, and updates the filter center frequency and bandwidth through backpropagation of the loss function during the training process, so that the sub-band distribution automatically focuses on the frequency band with strong discrimination. S33: Parallel extraction of multimodal features, extracting multimodal features from the original signal respectively; S34: Sub-band noise reduction enhancement, which uses the channel attention module to perform noise reduction weighting on the frequency domain features of each sub-band; S35: Feature fusion and output. The obtained multimodal features are concatenated across modalities and processed by bidirectional LSTM to output speaker embedding vectors containing contextual information.

4. The voiceprint recognition method based on multi-band analysis according to claim 3, characterized in that: When performing parallel extraction of multimodal features, the specific steps include: A11: Frequency domain feature extraction. Calculate the improved MFCC for each sub-band spectrum and count the sub-band energy entropy to obtain 40-dimensional features for each sub-band. The improved MFCC includes 13-dimensional MFCC + first-order difference + second-order difference + entropy. A12: Temporal feature extraction. Input the original waveform into a RawNet3 network containing SincNet convolutional layers to extract frame-level temporal features. A13: Phase feature extraction, calculate the instantaneous phase derivative from the complex spectrum of each sub-band to generate phase change mode features.

5. The voiceprint recognition method based on multi-band analysis according to claim 4, characterized in that: When applying noise-resistant weighting to the frequency domain features of each sub-band through the channel attention module, the channel attention module adds a compression-activation module after the frequency domain features of each sub-band. The compression-activation module includes global average pooling, fully connected layer learning weights, and feature recalibration. Global average pooling is used to compress the sub-band features into channel description vectors, fully connected layer learning weights is used to generate noise suppression coefficients for each frequency band, and feature recalibration is used to dynamically scale the sub-band features according to the weights to suppress low signal-to-noise ratio frequency bands.

6. The voiceprint recognition method based on multi-band analysis according to claim 5, characterized in that: When 16 sub-band filters are preset based on Mel-scale or uniform-scale, and the initial parameters are set to traditional empirical values ​​to obtain a filter bank, and the filter bank is implemented as a neural network layer, the response formula of each sub-band filter is: Among them, H m (f k ) represents the m-th filter at frequency f. k The response value at f c,m Let B be the center frequency of the m-th filter. m f is the bandwidth of the m-th filter. k These are discrete frequency points.

7. The voiceprint recognition method based on multi-band analysis according to claim 6, characterized in that: When training a robust model using attention fusion networks and contrastive training strategies, the specific steps include: S41: Attention mechanism fusion, input the multi-subband feature matrix into the multi-head attention layer, calculate the correlation weight between different subbands, and dynamically weight and fuse the subband features through attention weights to generate a global context-aware joint representation; S42: Bidirectional LSTM temporal modeling, inputting the attention-fused features into the bidirectional LSTM network to capture long-term dependencies between speech frames, and taking the hidden state of the last time step of the LSTM or performing average pooling on all time steps to obtain the speaker embedding vector. S43: Contrastive learning constraints: calculate cosine similarity for different sub-band features of the same speaker to maximize similarity; minimize similarity for features of different speakers and force features of different frequency bands to be close together in the embedding space to enhance the robustness of the model to partial frequency band missingness.

8. The voiceprint recognition method based on multi-band analysis according to claim 7, characterized in that: When training a robust model using attention fusion networks and contrastive training strategies, angular spacing is introduced into the classification layer to increase the angular difference between the embedding vectors of different speakers and improve inter-class discriminability. It also adds a domain classifier and obfuscates domain features through a gradient inversion layer to reduce cross-device voiceprint offset.

9. The voiceprint recognition method based on multi-band analysis according to claim 8, characterized in that: When deploying models using streaming processing and edge optimization, the specific steps include: S51: Streaming processing workflow, the speech stream is divided into 20ms frames, the features of each sub-band are extracted in parallel by multiple threads, and only the most recent 1 second speech frame is cached. Historical data is automatically eliminated, and the speaker feature library is compressed to 64 dimensions using PCA and retrieved in milliseconds through the approximate nearest neighbor algorithm. S52: Edge optimization process, including model lightweighting, cross-platform adaptation and low power consumption strategy. Among them, ITN8 is used for model lightweighting and low contribution subbands and neurons are removed. The low power consumption strategy is to use VAD to control activation and adjust the CPU and GPU frequencies according to the load.

10. A voiceprint recognition device based on multi-band analysis, used in the voiceprint recognition method based on multi-band analysis as described in any one of claims 1-9, characterized in that: include: The data preprocessing module is used to unify the voice format, extract the human voice segment using VAD, and dynamically reduce noise to suppress low-frequency noise and high-frequency distortion. The dynamic frequency band module is used to automatically optimize the frequency band division, extract multimodal features, and perform noise-resistant weighting through a learnable filter bank; The model training module is used to train a cross-device robust voiceprint embedding model by jointly performing attention fusion, LSTM temporal modeling, contrastive learning and noise enhancement. The real-time deployment module is used to achieve cross-platform, low-power operation of the model by employing streaming frame processing and lightweight models. The iterative optimization module is used to monitor the recognition effect online, dynamically adjust the threshold, incrementally learn new data, and update the frequency band division strategy.