Implementation method and device of multi-channel voiceprint recognition system
Through the combination of a multi-microphone ring array and an end-to-end neural network, high-precision voiceprint recognition is achieved in complex acoustic environments, solving the accuracy problem in multi-person conversation scenarios, optimizing the registration time for new users, adapting to complex environmental changes, and meeting the real-time requirements of smart vehicles and remote meetings.
Patent Information
- Application Number
- CN202510822880.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-19
AI Technical Summary
Existing voiceprint recognition technology has low recognition accuracy in complex acoustic environments, the multi-channel system has insufficient matching accuracy in multi-person conversation scenarios, and it takes a long time to register new users, making it difficult to meet the real-time and accuracy requirements of smart cars, remote conferencing and other fields.
It uses a multi-microphone ring array for data collection, achieves clock synchronization through the PTP protocol, combines an end-to-end neural network and a speaker perception separation network, dynamically optimizes training strategies, integrates high-resolution acoustic features and multi-channel phase difference features, supports lightweight deployment and incremental learning, and quickly adapts to environmental changes.
Improve voice quality and recognition accuracy in complex noisy environments, increase the accuracy of multi-person conversation separation, reduce new user registration time, reduce latency and memory usage, adapt to sudden noise and channel changes, and meet real-time and robustness requirements.
Smart Images

Figure CN120673765A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice recognition technology, and in particular to a method and device for implementing a multi-channel voiceprint recognition system. Background Art
[0002] With the rapid development of biometric technology, voiceprint recognition, as a contactless, convenient and efficient method of identity authentication, has shown broad application prospects in financial payment, smart home, security monitoring and other fields. However, the practical application of existing voiceprint recognition technology in complex acoustic environments still faces many technical bottlenecks.
[0003] Traditional voiceprint recognition systems are primarily based on single-channel audio processing technology, and their performance is severely limited by ambient noise and acoustic interference. In practical application scenarios, such as noisy public spaces or reverberant indoor environments, system recognition accuracy can significantly decrease. To address this issue, multi-channel voiceprint recognition technology has emerged. It uses microphone arrays to collect spatial audio information and beamforming technology to enhance the target sound source signal. However, existing technical solutions still have significant shortcomings in signal synchronization accuracy, noise suppression effectiveness, and feature fusion strategies.
[0004] More critically, existing solutions treat speech separation and voiceprint recognition as two separate modules. This disconnected design results in a target speaker matching accuracy of less than 70% when handling multi-person conversations. Furthermore, new user registration requires a full retraining of the model, which can take over 10 minutes and severely limits the system's practicality and scalability.
[0005] These technical bottlenecks collectively limit the large-scale application of voiceprint recognition systems in complex scenarios. Existing technical solutions often struggle to meet real-time, accuracy, and environmental adaptability requirements, particularly in areas like smart vehicles, remote conferencing, and public security. Therefore, a new multi-channel voiceprint recognition technology solution that can overcome these limitations is urgently needed. Summary of the Invention
[0006] In order to overcome the problems raised in the above background technology, the present invention proposes a method and device for implementing a multi-channel voiceprint recognition system.
[0007] The technical solution of the present invention is: a method for implementing a multi-channel voiceprint recognition system, comprising the following steps:
[0008] S11: Multi-channel data acquisition and synchronization, using a multi-microphone ring array for sound data acquisition and clock synchronization via the PTP protocol;
[0009] S12: Signal preprocessing and enhancement: Input multi-channel raw audio, generate target speech through an end-to-end neural network, and use a speaker-aware separation network combined with pre-registered voiceprint embedding to separate the target speech;
[0010] S13: Feature extraction and fusion: extract high-resolution logarithmic Mel-spectrogram from the enhanced speech, calculate the phase difference matrix between multiple channels, and finally fuse the high-resolution acoustic features with the multi-channel phase difference spatial features;
[0011] S14: Model training, based on branch-attention network combined with multi-channel data, dynamically optimizes training strategy;
[0012] S15: Real-time deployment, achieving low-latency inference within 10ms through lightweight models and heterogeneous computing;
[0013] S16: Adaptive optimization, supporting rapid adaptation to new users and environments through incremental learning and dynamic parameter adjustment.
[0014] Preferably, when inputting multi-channel original audio, generating target speech through an end-to-end neural network, and using a speaker-aware separation network in combination with pre-registered voiceprint embedding to separate the target speech, the method specifically includes:
[0015] S21: Deep beamforming: This method feeds synchronously collected multi-channel raw signals into an end-to-end neural network, which automatically learns noise distribution and speech characteristics. It then directly generates the target speaker's enhanced speech through time-frequency masking and waveform reconstruction, and outputs a pure single-channel signal that retains the target speaker's core characteristics.
[0016] S22: Multi-person conversation separation: First, based on the short-term energy mutation and spectral entropy value, it is determined whether there are multiple people speaking at the same time, triggering multi-person conversation separation, calling the pre-registered voiceprint embedding as the reference of the target speaker, inputting the enhanced speech into the speaker perception separation network, generating a time-frequency mask for the target speaker, and applying the mask to the mixed speech to separate the independent audio stream of each speaker and output it.
[0017] Preferably, when the synchronously collected multi-channel original signals are input into the end-to-end neural network, the architecture of the end-to-end neural network includes:
[0018] A11: Input layer, used to receive multi-channel raw waveform data and input the microphone position coordinates and pre-registered features of the target speaker;
[0019] A12: Spatial encoder, used to extract inter-channel spatial relationships and time-frequency features. It includes a spatially aware convolutional layer and a graph neural network. The spatially aware convolutional layer uses the convolution kernel to parameterize the microphone spacing and orientation and output inter-channel correlation features. The graph neural network is used to model the microphone array as a graph structure.
[0020] A13: Cross-channel interaction module, used to dynamically fuse multi-channel information and suppress interference. It includes a deformable attention mechanism and a conditional gating mechanism. The deformable attention mechanism is used to dynamically adjust the receptive field and focus on the direction of the target sound source. The conditional gating mechanism is used to dynamically activate different branches based on the input signal-to-noise ratio (SNR). High SNR activates lightweight branches, while low SNR activates complex branches.
[0021] A14: Speech separation network, used to separate the target speaker's speech and supports multi-person scenarios. It includes a speaker-aware mask generator and a multi-scale residual network. The speaker-aware mask generator processes voiceprint embedding information and cross-channel features to generate a time-frequency mask for the target speaker.
[0022] A15: Output layer, used to reconstruct the target speech waveform through the decoder;
[0023] A16: Dynamic optimization module, used to adapt to environmental and hardware changes in real time. It includes an online meta-learning controller and a hardware-aware scheduler. The online meta-learning controller dynamically adjusts network weights based on current environmental characteristics. The hardware-aware scheduler selects an operating mode based on the computing power of the device. These modes include high-performance mode and energy-saving mode. In high-performance mode, all modules are enabled, while in energy-saving mode, the multi-scale residual network is skipped.
[0024] Preferably, when the pre-registered voiceprint embedding is used as a reference for the target speaker and the enhanced speech is input to the speaker-aware separation network, the architecture of the speaker-aware separation network includes:
[0025] A21: Input layer, used to synchronously input multi-channel speech, target voiceprint embedding, lip movement video stream and environmental parameters;
[0026] A22: Multimodal fusion module, used to combine acoustic, spatial, and visual information to improve separation accuracy. It includes a spatial-voiceprint joint attention module and a visual-assisted alignment module. The spatial-voiceprint joint attention module processes voiceprint embeddings and multi-channel phase differences to obtain dynamic weights and focus on the direction of the target speaker. The visual-assisted alignment module aligns lip movements with speech timing across the Transformer, constraining the separated speech to have consistent lip movement rhythm.
[0027] A23: Dynamic Conditioning Module, which dynamically adjusts the network structure based on the input scenario. This module includes context-aware gating and speaker-adaptive routing. Context-aware gating activates network branches of varying complexity based on SNR, RT60, and the number of speakers. Speaker-adaptive routing activates the zero-shot separation branch when an unknown speaker is detected, generating voiceprint embeddings using temporary reference speech.
[0028] A24: Hierarchical separation module, used for phased and refined speech separation. It includes a coarse separation layer, a fine-tuning layer, and a residual restoration network. The coarse separation layer generates a global time-frequency mask and outputs the initially separated speech. The fine-tuning layer performs frequency-band partitioning. For data below 2kHz, speech continuity is restored based on phase difference and spatial features. For data above 2kHz, voiceprint embedding is used to enhance speech details. The residual restoration network outputs detailed, enhanced speech based on the coarsely separated speech.
[0029] A25: Output layer, used to output the target speaker's clean speech, real-time voiceprint embedding, and separation confidence;
[0030] A26: Online adaptive module, used to dynamically adapt to the environment and unknown speakers, including an incremental meta-learning controller and a noise spectrum tracker. The incremental meta-learning controller is used to output the network weight fine-tuning amount based on a small amount of speech data in the new scene, and the noise spectrum tracker is used to update the noise power spectrum in real time and dynamically adjust the mask generation threshold.
[0031] Preferably, when extracting a high-resolution logarithmic Mel spectrum from the enhanced speech, calculating the phase difference matrix between multiple channels, and finally fusing the high-resolution acoustic features with the multi-channel phase difference spatial features, the method specifically includes:
[0032] S31: Acoustic feature extraction. First, pre-emphasize the enhanced single-channel speech. Then, perform spectrum calculation using a short-time Fourier transform and a Mel filter bank. The output of the Mel filter is logarithmized to obtain an 80-dimensional logarithmic Mel spectrum. Power-law compression is applied to each frame of the Mel spectrum to enhance details in low-energy frequency bands.
[0033] S32: Spatial feature modeling. First, the short-time Fourier transform of the multi-channel original signal is calculated to obtain the complex spectrum, and the phase difference matrix between the channels is calculated. Then, a lightweight convolutional network is used to compress and encode the features.
[0034] S33: Joint feature generation, which fuses high-resolution acoustic features with multi-channel phase difference spatial features to obtain joint features.
[0035] Preferably, when the high-resolution acoustic features and the multi-channel phase difference spatial features are fused to obtain the joint features, the method specifically includes:
[0036] S41: Feature concatenation, concatenating the 80-dimensional acoustic features and 32-dimensional spatial features of each frame to form a 112-dimensional joint feature vector;
[0037] S42: Temporal alignment, ensuring that the frame numbers of acoustic features and spatial features are consistent. If there is a deviation due to processing delay, linear interpolation alignment is used;
[0038] S43: Feature normalization: First, the feature mean of the entire speech segment is subtracted by cepstral mean subtraction to eliminate channel offset. Then, variance normalization is performed to normalize the feature variance to unity per frame to suppress the influence of environmental noise.
[0039] S44: Sliding window context integration, which concatenates the joint features of 5 adjacent frames to form a 560-dimensional context window to capture dynamic speech characteristics.
[0040] As a preference, when dynamically optimizing the training strategy based on the branch-attention network combined with multi-channel data, the following are specifically included:
[0041] S51: In the pre-training phase, we used millions of talking heads videos, converted them into simulated multi-channel data, and used Pyroomacoustics to add reverberation and inject multiple types of noise. The parameters were set to an initial learning rate of 1e-3, a batch size of 256, and 50 epochs of training.
[0042] S52: In the fine-tuning stage, the target manufacturer's data is used as the dataset, the weights of the first three layers of the ResNet branch are fixed, the parameters are set to a learning rate of 1e-4, a batch size of 64, and 10-20 training rounds.
[0043] As a preference, when dynamically optimizing the training strategy based on the branch-attention network combined with multi-channel data, the architecture of the branch-attention network is:
[0044] A31: Input processing layer, where the 112-dimensional joint features of each microphone channel are input into an independent ResNet branch to extract channel-specific deep features;
[0045] A32: Branch feature extraction layer, each ResNet branch contains 4 residual blocks and outputs a 512-dimensional feature vector;
[0046] A33: Cross-channel attention fusion layer: This layer concatenates the 512-dimensional features of the six channels into a matrix, inputs the matrix into the multi-head attention layer, calculates the correlation weights between the channels, enhances the features of the channel where the target speaker is located, and suppresses the channels dominated by noise and interference sources.
[0047] A34: Global feature generation layer, which takes the average of the weighted channel features and outputs a 512-dimensional global voiceprint feature.
[0048] As a preference, when dynamically optimizing the training strategy based on the branch-attention network combined with multi-channel data, the principle formula of the loss function adopted by the branch-attention network is:
[0049]
[0050] in, is the loss function, e is the natural exponential function, s is the scaling factor, θ yi is the target class angle, m is the angular interval, cos(θ yi +m) is the cosine value after interval adjustment, is the sum of the scores of non-target classes.
[0051] As a preferred method, when supporting rapid adaptation of new users and environments through incremental learning and dynamic parameter adjustment, it specifically includes:
[0052] A41: Incremental learning adaptation. When a new user registers, the backbone network is frozen after multi-channel voice is collected, and only the fully connected layer is fine-tuned. EWC regularization is used to constrain the parameter update direction to prevent performance degradation for existing users.
[0053] A42: Environmental adaptation: Based on real-time noise statistics, it dynamically adjusts the CMS and MVN normalization parameters to optimize feature distribution stability, adapt to burst noise, and improve recognition robustness in low signal-to-noise ratio scenarios.
[0054] The implementation device of the channel voiceprint recognition system includes:
[0055] The data acquisition and synchronization module is used to collect audio signals through a multi-microphone array, implement multi-channel clock synchronization in conjunction with the PTP protocol, and integrate environmental sensors to assist in sound source localization.
[0056] The signal processing and enhancement module is used to suppress noise and generate target speech through an end-to-end neural network, and to separate multi-speaker speech streams based on voiceprint embedding and time-frequency masking technology;
[0057] The feature processing module processes the enhanced data, generates an 80-dimensional high-resolution logarithmic Mel spectrum, enhances low-energy frequency band details, calculates the multi-channel phase difference matrix, and stitches acoustic and spatial features to eliminate environmental interference through CMS and MVN.
[0058] The model training and inference module is used to train models based on branch-attention networks and multi-channel data, and achieves real-time voiceprint recognition within 10ms with lightweight deployment;
[0059] Adaptive optimization module, used to adapt to new users through incremental learning and dynamically adjust noise parameters to adapt to environmental changes;
[0060] The exception handling and interaction module is used to detect forgery attacks and sudden noise, trigger emergency strategies and feedback user interaction instructions.
[0061] Beneficial effects of the present invention:
[0062] 1. Traditional multi-channel voiceprint recognition systems rely on fixed beamforming algorithms (such as MVDR) and independent clock synchronization modules, which have problems such as large synchronization errors, manual parameter adjustment required for noise suppression, and poor generalization. This solution uses the PTP protocol to achieve hardware-level clock synchronization, and combines end-to-end neural networks to automatically learn noise distribution and sound source spatial location, dynamically generating beamforming weights. This solution can improve voice quality in complex noise scenarios without manual intervention, while significantly reducing the interference of synchronization errors on sound source positioning, bringing the accuracy and stability of far-field voice enhancement to a new level;
[0063] 2. In existing technologies for multi-speaker separation, voiceprint registration and speech separation modules are usually designed independently, resulting in a low degree of match between the separation results and the target voiceprint, especially in scenarios with high voiceprint similarity. This solution innovatively introduces a multimodal speaker-aware separation network, combining voiceprint embedding, multi-channel phase difference, and lip movement video timing alignment to generate a target-oriented time-frequency mask. This design improves the accuracy of multi-speaker separation and reduces the voiceprint matching error rate. It also supports zero-sample separation of unknown speakers, significantly expanding the system's applicable scenarios.
[0064] 3. Traditional methods rely on single features or static feature fusion strategies, resulting in poor environmental robustness and high model computational complexity, making it difficult to meet real-time requirements. This solution fuses high-resolution acoustic and spatial features, combined with INT8 quantization and a heterogeneous computing architecture, to reduce recognition error rates in noisy scenarios while achieving end-to-end low-latency inference. Furthermore, dynamic normalization and sliding window context integration mechanisms improve system stability under sudden noise and channel variations, reducing memory usage to a quarter of that of traditional models, making it ideal for embedded device deployment. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 Shown is a flowchart of a method for implementing a multi-channel voiceprint recognition system of the present invention;
[0066] Figure 2 Shown is a schematic diagram of the structure of the implementation device of the multi-channel voiceprint recognition system of the present invention. DETAILED DESCRIPTION
[0067] The present invention will be further described below with reference to the accompanying drawings and examples.
[0068] See also Figure 1-2 The present invention provides an embodiment: a method for implementing a multi-channel voiceprint recognition system, the method for implementing a multi-channel voiceprint recognition system comprising the following steps:
[0069] Step 1: Multi-channel data acquisition and synchronization
[0070] Use a multi-microphone ring array to collect sound data and achieve clock synchronization through the PTP protocol;
[0071] Step 2: Signal preprocessing and enhancement
[0072] The target speech is generated by inputting multi-channel raw audio through an end-to-end neural network. The target speech is then separated using a speaker-aware separation network combined with pre-registered voiceprint embedding. Specifically, the following steps are performed:
[0073] Deep beamforming inputs synchronously collected multi-channel raw signals into an end-to-end neural network, which automatically learns noise distribution and speech features. It then directly generates the enhanced speech of the target speaker through time-frequency masking and waveform reconstruction, outputting a pure single-channel signal that retains the core characteristics of the target speaker. Multi-person conversation separation first determines whether multiple people are speaking simultaneously based on short-term energy mutations and spectral entropy values, triggering multi-person conversation separation. The pre-registered voiceprint embedding is used as a reference for the target speaker, and the enhanced speech is input into the speaker perception separation network. A time-frequency mask for the target speaker is generated and applied to the mixed speech to separate and output each speaker's independent audio stream.
[0074] The end-to-end neural network architecture includes: an input layer for receiving multi-channel raw waveform data and inputting the position coordinates of the microphone and the pre-registered features of the target speaker;
[0075] Spatial encoder, used to extract spatial relationships and time-frequency features between channels, including spatial perception convolution layer and graph neural network. The spatial perception convolution layer is used to parameterize the microphone spacing and direction of the convolution kernel and output the inter-channel correlation features. The graph neural network is used to model the microphone array as a graph structure; Cross-channel interaction module, used to dynamically fuse multi-channel information and suppress interference, including deformable attention mechanism and conditional gating mechanism. The deformable attention mechanism is used to dynamically adjust the receptive field and focus on the direction of the target sound source. The conditional gating mechanism is used to dynamically activate different branches based on the input SNR, where high SNR enables lightweight branches and low SNR enables complex branches; Speech separation network, used to separate the target speech. Speaker voice, supporting multi-person scenarios, including a speaker-aware mask generator and a multi-scale residual network. The speaker-aware mask generator is used to process voiceprint embedding information and cross-channel features to generate a time-frequency mask for the target speaker; the output layer is used to reconstruct the target speech waveform through the decoder; the dynamic optimization module is used to adapt to environmental and hardware changes in real time, including an online meta-learning controller and a hardware-aware scheduler. The online meta-learning controller is used to dynamically adjust network weights based on current environmental characteristics, and the hardware-aware scheduler is used to select an operating mode based on the device computing power, including high-performance mode and energy-saving mode. In high-performance mode, all modules are enabled, and in energy-saving mode, the multi-scale residual network is skipped.
[0076] The architecture of the speaker awareness separation network includes:
[0077] A21: Input layer, used to synchronously input multi-channel speech, target voiceprint embedding, lip movement video stream and environmental parameters; multimodal fusion module, used to combine acoustic, spatial and visual information to improve separation accuracy, including spatial-voiceprint joint attention and visual-assisted alignment module. Spatial-voiceprint joint attention is used to process voiceprint embedding and multi-channel phase difference to obtain dynamic weights and focus on the direction of the target speaker. The visual-assisted alignment module is used to align lip movement and speech timing across the modal Transformer to constrain the separated speech to be consistent with the lip movement rhythm; dynamic conditioning module, used to dynamically adjust the network structure according to the input scenario, specifically including environmental perception gating and speaker adaptive routing. Environmental perception gating is used to activate network branches of different complexities according to SNR, RT60 and the number of speakers. Speaker adaptive routing is used to activate the zero-sample separation branch when an unknown speaker is detected, using temporary reference Speech generation voiceprint embedding; hierarchical separation module, used for staged refined speech separation, including a coarse separation layer, a fine-tuning layer and a residual repair network. The coarse separation layer is used to generate a global time-frequency mask and output the preliminarily separated speech. The fine-tuning layer is used for frequency band divide-and-conquer processing. For data less than 2kHz, the speech continuity is repaired based on phase difference and spatial features. For data greater than 2kHz, the speech details are enhanced by relying on voiceprint embedding. The residual repair network is used to output the enhanced speech according to the coarsely separated speech; the output layer is used to output the clean speech of the target speaker, real-time voiceprint embedding and separation confidence; the online adaptive module is used to dynamically adapt to the environment and unknown speakers, including an incremental meta-learning controller and a noise spectrum tracker. The incremental meta-learning controller is used to output the network weight fine-tuning amount based on a small amount of speech data in the new scene. The noise spectrum tracker is used to update the noise power spectrum in real time and dynamically adjust the mask generation threshold.
[0078] Step 3: Feature extraction and fusion
[0079] The high-resolution logarithmic Mel spectrum is extracted from the enhanced speech, and the phase difference matrix between multiple channels is calculated. Finally, the high-resolution acoustic features are fused with the multi-channel phase difference spatial features. Specifically, the acoustic feature extraction is performed. First, the enhanced single-channel speech is pre-emphasized. Then, the spectrum is calculated through short-time Fourier transform and Mel filter group. Then, the output of the Mel filter is logarithmic to obtain an 80-dimensional logarithmic Mel spectrum. Power-law compression is applied to each frame of the Mel spectrum to enhance the details of the low-energy frequency band. Spatial feature modeling is performed. First, the short-time Fourier transform is calculated for the original multi-channel signal to obtain a complex spectrum, and the phase difference matrix between channels is calculated. Then, a lightweight convolutional network is used for feature extraction. Compression and encoding; Joint feature generation, fusing high-resolution acoustic features with multi-channel phase difference spatial features to obtain joint features. Specifically, the 80-dimensional acoustic features and 32-dimensional spatial features of each frame are spliced together to form a 112-dimensional joint feature vector, ensuring that the number of frames of the acoustic features and spatial features is consistent. If there is a deviation due to processing delay, linear interpolation alignment is used. First, the feature mean of the entire speech segment is subtracted by the cepstral mean subtraction method to eliminate channel offset. Then, the feature variance is normalized to the unit range by frame through the variance normalization method to suppress the influence of environmental noise. The joint features of 5 adjacent frames are spliced to form a 560-dimensional context window to capture dynamic speech characteristics.
[0080] Step 4: Model training
[0081] Based on the branch-attention network combined with multi-channel data, the training strategy is dynamically optimized, including:
[0082] In the pre-training phase, we used millions of talking head videos, converted them into simulated multi-channel data, and used Pyroomacoustics to add reverberation and inject multiple types of noise. The initial learning rate was set to 1e-3, the batch size was 256, and training was performed for 50 epochs. In the fine-tuning phase, we used target manufacturer data as the dataset, fixed the weights of the first three layers of the ResNet branch, set the learning rate to 1e-4, the batch size to 64, and trained for 10-20 epochs.
[0083] Step 5: Real-time deployment
[0084] Achieve low-latency inference within 10ms through lightweight models and heterogeneous computing;
[0085] Step 6: Adaptive Optimization
[0086] Incremental learning and dynamic parameter adjustment support rapid adaptation to new users and environments, including:
[0087] Incremental learning adaptation: When a new user registers, the backbone network is frozen after collecting multi-channel voice, and only the fully connected layer is fine-tuned. EWC regularization is used to constrain the parameter update direction to prevent performance degradation for old users; environmental adaptation: Based on real-time noise statistics, the CMS and MVN normalization parameters are dynamically adjusted to optimize the stability of feature distribution, adapt to sudden noise, and improve recognition robustness in low signal-to-noise ratio scenarios.
[0088] Among them, when dynamically optimizing the training strategy based on the branch-attention network combined with multi-channel data, the architecture of the branch-attention network is:
[0089] In the input processing layer, the 112-dimensional joint features of each microphone channel are input into an independent ResNet branch to extract channel-specific deep features;
[0090] Branch feature extraction layer, each ResNet branch contains 4 residual blocks and outputs a 512-dimensional feature vector;
[0091] The cross-channel attention fusion layer concatenates the 512-dimensional features of the six channels into a matrix, which is then fed into the multi-head attention layer to calculate the correlation weights between the channels. This enhances the features of the channel where the target speaker is located and suppresses channels dominated by noise and interference sources.
[0092] The global feature generation layer takes the average of the weighted channel features and outputs a 512-dimensional global voiceprint feature.
[0093] Among them, when dynamically optimizing the training strategy based on the branch-attention network combined with multi-channel data, the principle formula of the loss function adopted by the branch-attention network is:
[0094]
[0095] in, is the loss function, e is the natural exponential function, s is the scaling factor, θ yi is the target class angle, m is the angular interval, cos(θ yi +m) is the cosine value after interval adjustment, is the sum of the scores of non-target classes.
[0096] The implementation device of the channel voiceprint recognition system includes:
[0097] The data acquisition and synchronization module is used to collect audio signals through a multi-microphone array, implement multi-channel clock synchronization in conjunction with the PTP protocol, and integrate environmental sensors to assist in sound source localization.
[0098] The signal processing and enhancement module is used to suppress noise and generate target speech through an end-to-end neural network, and to separate multi-speaker speech streams based on voiceprint embedding and time-frequency masking technology;
[0099] The feature processing module processes the enhanced data, generates an 80-dimensional high-resolution logarithmic Mel spectrum, enhances low-energy frequency band details, calculates the multi-channel phase difference matrix, and stitches acoustic and spatial features to eliminate environmental interference through CMS and MVN.
[0100] The model training and inference module is used to train models based on branch-attention networks and multi-channel data, and achieves real-time voiceprint recognition within 10ms with lightweight deployment;
[0101] Adaptive optimization module, used to adapt to new users through incremental learning and dynamically adjust noise parameters to adapt to environmental changes;
[0102] The exception handling and interaction module is used to detect forgery attacks and sudden noise, trigger emergency strategies and feedback user interaction instructions.
[0103] Example 1: Intelligent vehicle-mounted voiceprint recognition system
[0104] Application scenario: In a high-speed vehicle, the driver needs to use voiceprint recognition for identity authentication and control of vehicle systems (such as navigation, air conditioning, etc.). Due to interference from engine noise, wind noise, and passenger conversations, traditional single-channel systems have difficulty accurately identifying the target speaker.
[0105] Implementation steps:
[0106] Hardware configuration:
[0107] A circular array of six omnidirectional microphones is installed in the center of the vehicle roof, spaced 15 cm apart. It supports a 48kHz sampling rate. Hardware-level clock synchronization is achieved via the PTPv2 protocol over in-vehicle Ethernet, with a synchronization error of ≤0.05ms. An integrated DSP chip (TI TDA4VM) is used for real-time signal preprocessing.
[0108] Signal processing and enhancement:
[0109] Deep beamforming: Uses an improved Conv-TasNet network (8-layer time-domain convolution, 256 channels), inputs a 6-channel raw signal, and outputs a single-channel speech with enhanced driver direction. Dynamic noise suppression: Dynamically adjusts network parameters based on real-time vehicle speed (obtained via the CAN bus), maintaining speech clarity (PESQ ≥ 3.5) even in wind noise at 120 km / h.
[0110] Feature extraction and fusion:
[0111] Acoustic features: 80-dimensional log-mel spectrum (25ms frame length, 10ms frame shift), with power-law compression (exponent 0.3) enhancing low-energy frequency bands. Spatial features: 6-channel inter-channel phase difference (MPSD) matrices are calculated and compressed using three layers of lightweight convolution (kernel 3×3, output 32 dimensions). Feature fusion: Concatenated into a 112-dimensional joint feature, integrating 5 frames of context (560 dimensions) using a sliding window.
[0112] Model training and deployment:
[0113] Branch-Attention Network: Six independent ResNet-18 branches (each with four residual blocks) extract channel features, followed by weighted fusion of cross-channel attention (8 heads). Training Strategy: Pre-training uses simulated vehicle data (including engine, wind, and tire noise), and fine-tuning uses 200 hours of real-world driving recordings. Lightweight Deployment: The model is quantized to INT8 (with <0.5% accuracy loss) and deployed on NVIDIA Jetson Xavier, achieving an end-to-end latency of 8ms.
[0114] Adaptive Optimization:
[0115] Incremental learning: When a new driver registers, we collect one minute of speech and freeze the backbone network. Only the classification head is fine-tuned (learning rate 1e-4) with an EWC regularization coefficient of λ = 0.3. Environmental adaptation: We update the noise mean and variance (CMS / MVN) in real time and dynamically adjust feature normalization parameters.
[0116] Test results:
[0117] Noise scenario (65dB engine noise): EER = 2.3%, False Rejection Rate (FRR) < 1%. New user registration time: 45 seconds (traditional methods take 10 minutes). Extreme scenario (120km / h high-speed wind noise): Voice command recognition accuracy ≥ 95%.
[0118] Example 2: Multi-speaker recognition system in conference scenarios
[0119] Application scenario: In a remote conference with 8 participants, the system needs to separate and identify the voiceprint of each speaker in real time, while supporting the rapid registration of new participants.
[0120] Implementation steps:
[0121] Hardware configuration:
[0122] Eight omnidirectional microphones are distributed around the conference table (30cm apart), supporting Wi-Fi 6's TWT mechanism for low-power synchronization (error <0.1ms). An integrated RGB camera (1080P@30fps) captures lip movements.
[0123] Signal processing and separation:
[0124] Multimodal Separation Network (SpEx++): Input: 8-channel speech + lip movement video (256-dimensional features extracted by EfficientNet). Coarse separation layer: GCC-PHAT locates the sound source direction and generates a global time-frequency mask. Fine-tuning layer: Phase difference repair is used for frequencies <2kHz, and voiceprint embedding is used for enhancement in frequencies >2kHz. Zero-shot separation: Voiceprint embedding is generated from a temporary reference speech (5 seconds) to separate unknown speakers.
[0125] Feature extraction and fusion: Acoustic features: 128-dimensional linear Mel-spectrogram (0-16kHz), with dynamic range compression (DRC) to enhance speech detail. Spatial features: 64-dimensional phase difference features (combined across 8 channels), with topological relationships modeled using a graph neural network (GNN). Multimodal fusion: Cross-modal Transformer alignment of voiceprints and lip movement timing (number of attention heads = 8).
[0126] Model training and deployment:
[0127] Training data: A mixture of simulated conference room audio (reverberation time 0.3-1.5s) and real-world conference recordings (100 hours). Network architecture: A hierarchical separation network (12M parameters), outputting separated speech and real-time voiceprint embedding. Deployment optimization: The model was pruned to 3.5W power consumption and runs on an Intel Movidius Myriad XVPU.
[0128] Dynamic optimization mechanism:
[0129] Incremental meta-learning: New participants register through 5 seconds of speech, and the meta-learning controller generates adaptation parameters within 10 seconds (learning rate 1e-5). Noise spectrum tracking: Real-time estimation of the noise power spectrum and dynamic adjustment of the mask threshold (update frequency 10Hz).
[0130] Test results:
[0131] 8 people speaking simultaneously: Speech separation signal-to-noise ratio (SI-SNR) = 12.7dB, voiceprint recognition accuracy 95.6%. New user registration: 5 seconds of reference speech, recognition accuracy 88% (traditional methods require ≥ 30 seconds). Burst noise (such as shuffling papers): The system automatically switches to anti-interference mode, with EER fluctuation ≤ 1%. Power consumption and latency: Total power consumption is 3.5W, and end-to-end processing latency is 15ms (including lip movement video processing).
[0132] The embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the scope of knowledge of those skilled in the art without departing from the spirit of the present invention.
Claims
1. A method for implementing a multi-channel voiceprint recognition system, characterized by: The following steps are involved: S11: Multi-channel data acquisition and synchronization, using a multi-microphone ring array for sound data acquisition and clock synchronization via the PTP protocol; S12: Signal preprocessing and enhancement: Input multi-channel raw audio, generate target speech through an end-to-end neural network, and use a speaker-aware separation network combined with pre-registered voiceprint embedding to separate the target speech; S13: Feature extraction and fusion: extract high-resolution logarithmic Mel-spectrogram from the enhanced speech, calculate the phase difference matrix between multiple channels, and finally fuse the high-resolution acoustic features with the multi-channel phase difference spatial features; S14: Model training, based on branch-attention network combined with multi-channel data, dynamically optimizes training strategy; S15: Real-time deployment, achieving low-latency inference within 10ms through lightweight models and heterogeneous computing; S16: Adaptive optimization, supporting rapid adaptation to new users and environments through incremental learning and dynamic parameter adjustment.
2. The method for implementing the multi-channel voiceprint recognition system according to claim 1, characterized in that: When inputting multi-channel raw audio, generating the target speech through an end-to-end neural network, and using a speaker-aware separation network combined with pre-registered voiceprint embedding to separate the target speech, the following are specifically involved: S21: Deep beamforming: This method feeds synchronously collected multi-channel raw signals into an end-to-end neural network, which automatically learns noise distribution and speech characteristics. It then directly generates the target speaker's enhanced speech through time-frequency masking and waveform reconstruction, and outputs a pure single-channel signal that retains the target speaker's core characteristics. S22: Multi-person conversation separation: First, based on the short-term energy mutation and spectral entropy value, it is determined whether there are multiple people speaking at the same time, triggering multi-person conversation separation, calling the pre-registered voiceprint embedding as the reference of the target speaker, inputting the enhanced speech into the speaker perception separation network, generating a time-frequency mask for the target speaker, and applying the mask to the mixed speech to separate the independent audio stream of each speaker and output it.
3. The method for implementing the multi-channel voiceprint recognition system according to claim 2, characterized in that: When the synchronously collected multi-channel original signals are input into the end-to-end neural network, the architecture of the end-to-end neural network includes: A11: Input layer, used to receive multi-channel raw waveform data and input the microphone position coordinates and pre-registered features of the target speaker; A12: Spatial encoder, used to extract inter-channel spatial relationships and time-frequency features. It includes a spatially aware convolutional layer and a graph neural network. The spatially aware convolutional layer uses the convolution kernel to parameterize the microphone spacing and orientation and output inter-channel correlation features. The graph neural network is used to model the microphone array as a graph structure. A13: Cross-channel interaction module, used to dynamically fuse multi-channel information and suppress interference. It includes a deformable attention mechanism and a conditional gating mechanism. The deformable attention mechanism is used to dynamically adjust the receptive field and focus on the direction of the target sound source. The conditional gating mechanism is used to dynamically activate different branches based on the input signal-to-noise ratio (SNR). High SNR activates lightweight branches, while low SNR activates complex branches. A14: Speech separation network, used to separate the target speaker's speech and supports multi-person scenarios. It includes a speaker-aware mask generator and a multi-scale residual network. The speaker-aware mask generator processes voiceprint embedding information and cross-channel features to generate a time-frequency mask for the target speaker. A15: Output layer, used to reconstruct the target speech waveform through the decoder; A16: Dynamic optimization module, used to adapt to environmental and hardware changes in real time. It includes an online meta-learning controller and a hardware-aware scheduler. The online meta-learning controller dynamically adjusts network weights based on current environmental characteristics. The hardware-aware scheduler selects an operating mode based on the device computing power. The operating modes include high-performance mode and energy-saving mode. In high-performance mode, all modules are enabled, while in energy-saving mode, the multi-scale residual network is skipped.
4. The method for implementing the multi-channel voiceprint recognition system according to claim 3, characterized in that: When the pre-registered voiceprint embedding is used as the target speaker reference and the enhanced speech is input to the speaker-aware separation network, the architecture of the speaker-aware separation network includes: A21: Input layer, used to synchronously input multi-channel speech, target voiceprint embedding, lip movement video stream and environmental parameters; A22: Multimodal fusion module, used to combine acoustic, spatial, and visual information to improve separation accuracy. It includes a spatial-voiceprint joint attention module and a visual-assisted alignment module. The spatial-voiceprint joint attention module processes voiceprint embeddings and multi-channel phase differences to obtain dynamic weights and focus on the direction of the target speaker. The visual-assisted alignment module aligns lip movements with speech timing across the Transformer, constraining the separated speech to have consistent lip movement rhythm. A23: Dynamic Conditioning Module, which dynamically adjusts the network structure based on the input scenario. This module includes context-aware gating and speaker-adaptive routing. Context-aware gating activates network branches of varying complexity based on SNR, RT60, and the number of speakers. Speaker-adaptive routing activates the zero-shot separation branch when an unknown speaker is detected, generating voiceprint embeddings using temporary reference speech. A24: Hierarchical separation module, used for phased and refined speech separation. It includes a coarse separation layer, a fine-tuning layer, and a residual restoration network. The coarse separation layer generates a global time-frequency mask and outputs the initially separated speech. The fine-tuning layer performs frequency-band partitioning. For data below 2kHz, speech continuity is restored based on phase difference and spatial features. For data above 2kHz, voiceprint embedding is used to enhance speech details. The residual restoration network outputs detailed, enhanced speech based on the coarsely separated speech. A25: Output layer, used to output the target speaker's clean speech, real-time voiceprint embedding, and separation confidence; A26: Online adaptive module, used to dynamically adapt to the environment and unknown speakers, including an incremental meta-learning controller and a noise spectrum tracker. The incremental meta-learning controller is used to output the network weight fine-tuning amount based on a small amount of speech data in the new scene, and the noise spectrum tracker is used to update the noise power spectrum in real time and dynamically adjust the mask generation threshold.
5. The method for implementing the multi-channel voiceprint recognition system according to claim 4, characterized in that: Extracting high-resolution logarithmic mel-spectrograms from the enhanced speech, calculating the phase difference matrix between multiple channels, and finally fusing high-resolution acoustic features with multi-channel phase difference spatial features, specifically includes: S31: Acoustic feature extraction. First, pre-emphasize the enhanced single-channel speech. Then, perform spectrum calculation using a short-time Fourier transform and a Mel filter bank. The output of the Mel filter is logarithmized to obtain an 80-dimensional logarithmic Mel spectrum. Power-law compression is applied to each frame of the Mel spectrum to enhance details in low-energy frequency bands. S32: Spatial feature modeling. First, the short-time Fourier transform of the multi-channel original signal is calculated to obtain the complex spectrum, and the phase difference matrix between the channels is calculated. Then, a lightweight convolutional network is used to compress and encode the features. S33: Joint feature generation, which fuses high-resolution acoustic features with multi-channel phase difference spatial features to obtain joint features.
6. The method for implementing the multi-channel voiceprint recognition system according to claim 5, characterized in that: When the high-resolution acoustic features and multi-channel phase difference spatial features are fused to obtain the joint features, the specific features include: S41: Feature concatenation, concatenating the 80-dimensional acoustic features and 32-dimensional spatial features of each frame to form a 112-dimensional joint feature vector; S42: Temporal alignment to ensure that the frame numbers of acoustic features and spatial features are consistent. If there is a deviation due to processing delay, linear interpolation alignment is used; S43: Feature normalization: First, the feature mean of the entire speech segment is subtracted by cepstral mean subtraction to eliminate channel offset. Then, variance normalization is performed to normalize the feature variance to unity per frame to suppress the influence of environmental noise. S44: Sliding window context integration, which concatenates the joint features of 5 adjacent frames to form a 560-dimensional context window to capture dynamic speech characteristics.
7. The method for implementing the multi-channel voiceprint recognition system according to claim 6, characterized in that: When dynamically optimizing the training strategy based on the branch-attention network combined with multi-channel data, it specifically includes: S51: In the pre-training phase, we used millions of talking heads videos, converted them into simulated multi-channel data, and used Pyroomacoustics to add reverberation and inject multiple types of noise. The parameters were set as follows: initial learning rate 1e-3, batch size 256, and training for 50 rounds. S52: In the fine-tuning stage, the target manufacturer's data is used as the dataset, the weights of the first three layers of the ResNet branch are fixed, the parameters are set to a learning rate of 1e-4, a batch size of 64, and 10-20 training rounds.
8. The method for implementing the multi-channel voiceprint recognition system according to claim 7, characterized in that: When dynamically optimizing the training strategy based on the branch-attention network combined with multi-channel data, the architecture of the branch-attention network is as follows: A31: Input processing layer, where the 112-dimensional joint features of each microphone channel are input into an independent ResNet branch to extract channel-specific deep features; A32: Branch feature extraction layer, each ResNet branch contains 4 residual blocks and outputs a 512-dimensional feature vector; A33: Cross-channel attention fusion layer: This layer concatenates the 512-dimensional features of the six channels into a matrix, inputs the matrix into the multi-head attention layer, calculates the correlation weights between the channels, enhances the features of the channel where the target speaker is located, and suppresses the channels dominated by noise and interference sources. A34: Global feature generation layer, which takes the average of the weighted channel features and outputs a 512-dimensional global voiceprint feature.
9. The method for implementing the multi-channel voiceprint recognition system according to claim 8, characterized in that: When supporting rapid adaptation of new users and environments through incremental learning and dynamic parameter adjustment, it specifically includes: A41: Incremental learning adaptation. When a new user registers, the backbone network is frozen after multi-channel voice is collected, and only the fully connected layer is fine-tuned. EWC regularization is used to constrain the parameter update direction to prevent performance degradation for existing users. A42: Environmental adaptation: Based on real-time noise statistics, it dynamically adjusts the CMS and MVN normalization parameters to optimize feature distribution stability, adapt to burst noise, and improve recognition robustness in low signal-to-noise ratio scenarios.
10. A device for implementing a multi-channel voiceprint recognition system, used in the method for implementing a multi-channel voiceprint recognition system according to any one of claims 1 to 9, characterized in that: include: The data acquisition and synchronization module is used to collect audio signals through a multi-microphone array, implement multi-channel clock synchronization in conjunction with the PTP protocol, and integrate environmental sensors to assist in sound source localization. The signal processing and enhancement module is used to suppress noise and generate target speech through an end-to-end neural network, and to separate multi-speaker speech streams based on voiceprint embedding and time-frequency masking technology; The feature processing module processes the enhanced data, generates an 80-dimensional high-resolution logarithmic Mel spectrum, enhances low-energy frequency band details, calculates the multi-channel phase difference matrix, and stitches acoustic and spatial features to eliminate environmental interference through CMS and MVN. The model training and inference module is used to train models based on branch-attention networks and multi-channel data, and achieves real-time voiceprint recognition within 10ms with lightweight deployment; The adaptive optimization module is used to adapt to new users through incremental learning and dynamically adjust noise parameters to adapt to environmental changes.
Citation Information
Cited By
Voice data recognition method and system based on AI voice algorithm
CN121237092A
Voice noise reduction method based on reasoning optimization and Bluetooth earphone
CN121506162A
Voice noise reduction method based on reasoning optimization and bluetooth earphone
CN121506162B