Multi-channel voice processing method and system

By building a multi-channel parallel processing architecture and multi-stage optimization training, dynamically integrating multi-channel spatial features, the problem of insufficient robustness and adaptability of multi-channel speech processing methods in the existing technology in complex acoustic environments is solved, and efficient multi-channel speaker verification is achieved.

CN120600031APending Publication Date: 2025-09-05ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510909250.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The existing multi-channel speech processing methods are insufficient in complex acoustic environments, and single-channel pre-trained models are difficult to effectively utilize multi-microphone array spatial information, resulting in enhanced error propagation and limited performance.

Method used

By building a multi-channel parallel processing architecture, a multi-channel information interaction module and a cross-channel attention mechanism are introduced, and multi-stage optimization training is carried out in combination with multi-head attention pooling and additive angle cosine loss, a multi-channel vocalprint feature extraction model is generated, and multi-channel spatial features are dynamically integrated and feature distinction is improved.

Benefits of technology

It significantly improves the accuracy and robustness of multi-channel speaker verification, reduces the error rate, enhances the system's adaptability in complex acoustic environments, and is better than the existing technology performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600031A_ABST
    Figure CN120600031A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-channel voice processing method and system, and belongs to the field of artificial intelligence and voice signal processing. Comprising the following steps: acquiring a multi-channel audio signal and constructing a single-channel pre-training model of SSL; based on the multi-channel audio signal, performing structure optimization on the single-channel pre-training model of the SSL to obtain a multi-channel voiceprint feature extraction pre-training model; performing multi-stage joint optimization training, performing fine tuning on the multi-channel voiceprint feature extraction pre-training model in combination with AAM loss, and generating a multi-channel voice processing model; and when a to-be-processed multi-channel audio signal is received, processing the to-be-processed multi-channel audio signal through the multi-channel voice processing model, and outputting a high-discrimination multi-channel voiceprint feature. The invention aims to improve the accuracy and robustness of speaker verification in a multi-channel scene, remarkably reduce the error rate and improve the adaptability of a system to a complex acoustic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and speech signal processing, and more particularly to a multi-channel speech processing method and system. Background Art

[0002] In recent years, speech or speaker recognition in far-field microphone environments has faced significant interference from reverberation and noise. To alleviate this problem, multi-channel speech processing technology has become an important means of improving system robustness by leveraging the spatial information provided by microphone arrays to distinguish target sound sources from interference sources. In particular, multi-channel technology is crucial for personalized services and identity authentication in speaker verification (SV).

[0003] Current multi-channel speaker verification methods can be divided into three main categories: the first is multi-channel preprocessing, which performs front-end processing on the input signal through beamforming or speech enhancement techniques, and then feeds the processed signal into a single-channel model. Although this approach has a high degree of modularity, its performance is limited due to the weak correlation between the enhancement target and downstream tasks (such as speech recognition). Errors generated during the enhancement process may also propagate to subsequent models. The second category is customized multi-channel models, which are implemented by training a dedicated network architecture from scratch. However, this approach requires a large amount of training data and has poor adaptability to changes in the number of channels, limiting its flexibility in practical applications. The third category is single-channel model extension, which independently processes data from each channel based on a powerful single-channel model (such as a pre-trained self-supervised learning SSL model) and then integrates data from different channels through a cross-channel information fusion module. This approach has shown superior performance when performing information fusion at the frame level. However, existing SSL models (such as WavLM and HuBERT) are typically designed only for a single channel. Directly applying them to multi-channel data may overlook important spatial information.

[0004] Traditional methods often combine cascaded multi-channel enhancement with single-channel SSL models (such as the CHiME-7 competition proposal). However, spatial information is only utilized in the front-end processing, and enhancement errors can affect the SSL model's feature extraction. Furthermore, existing multi-channel SSL extension models (such as Spatial HuBERT) are not publicly available and are not optimized for speaker verification tasks.

[0005] Therefore, how to provide a multi-channel voice processing method and system is a problem that those skilled in the art need to solve urgently. Summary of the Invention

[0006] In view of this, the present invention provides a multi-channel speech processing method and system, which aims to improve the accuracy and robustness of speaker verification in multi-channel scenarios, significantly reduce the error rate and enhance the system's adaptability to complex acoustic environments, and solve the problem that existing single-channel pre-trained speech models are difficult to effectively utilize the spatial information of multi-microphone arrays.

[0007] In order to achieve the above object, the present invention provides the following technical solutions:

[0008] A multi-channel speech processing method comprises the following steps:

[0009] Acquire multi-channel audio signals and build a single-channel pre-trained model for SSL;

[0010] Based on the multi-channel audio signal, the single-channel pre-trained model of the SSL is structurally optimized to obtain a multi-channel voiceprint feature extraction pre-trained model;

[0011] Perform multi-stage joint optimization training, fine-tune the multi-channel voiceprint feature extraction pre-training model in combination with AAM loss, and generate a multi-channel speech processing model;

[0012] When a multi-channel audio signal to be processed is received, the multi-channel audio signal to be processed is processed by the multi-channel speech processing model, and a highly distinguishable multi-channel voiceprint feature is output.

[0013] Furthermore, the single-channel pre-training model of SSL is structurally optimized to obtain a multi-channel voiceprint feature extraction pre-training model, including:

[0014] By replicating the CNN and Transformer encoders of a single-channel pre-trained model, multiple parallel processing branches are formed. Each branch is responsible for processing the audio signal of one channel, and multi-channel features are extracted in parallel for multi-channel audio signals.

[0015] After parallel processing, a multi-channel information interaction module is introduced, and cross-channel audio signal fusion is performed through a cross-channel attention mechanism while maintaining the temporal integrity of single-channel features;

[0016] The head-attention pooling (MHFA) mechanism is used to reduce the dimension and optimize the fused features to generate voiceprint embeddings.

[0017] Obtain a multi-channel voiceprint feature extraction pre-training model.

[0018] Furthermore, during the parallel extraction of multi-channel features, the parameters of each branch remain shared to ensure that the model can learn universal voiceprint features.

[0019] Furthermore, the multi-channel information interaction module includes an MM module and an MS module;

[0020] Among them, the MM module also includes a transform-average-concatenation TAC module and a cross-frame joint attention COATT module.

[0021] Furthermore, the transform-average-splicing TAC module is specifically implemented as follows:

[0022] For speech data from C different channels, let the feature of the cth channel be expressed as Where T is the number of time frames and D is the feature dimension. For the TAC module, the following formula is used to realize the information interaction between channels:

[0023]

[0024] Where C is the total number of channels, f FC (·; θ) is the fully connected layer and nonlinear activation, is the second layer full connection parameter, MT is the global channel summary feature, represents the channel enhancement feature, is the TAC modular output, FC(x;α)=σ(xW+b), α={W,b}, σ(·) is the nonlinear activation function, || represents the concatenation operation, LN(·) represents the layer normalization operation, and ψ is f FC The set of parameters in a feedforward neural network, including the weight matrix and bias terms

[0025] Furthermore, the cross-frame joint attention COATT module includes:

[0026] By summarizing the channel information, a channel summary embedding S is generated, and the cross-frame joint attention mechanism is used to contextualize the channel summary and each channel feature to achieve feature interaction and enhancement between channels.

[0027] Furthermore, the specific implementation of the MS module includes two methods. The first method is to perform weighted averaging on the characteristics of each channel, and the weights are automatically learned and optimized based on the training data. The formula is:

[0028]

[0029] Among them, α c is the weight of the cth channel, satisfying

[0030] The second method is to use only the features of the first channel as output.

[0031] Furthermore, the head attention pooling MHFA mechanism includes:

[0032] Use multiple attention heads to model different dimensions of features and capture richer feature information;

[0033] Among them, the fused features are first subjected to dimensionality reduction processing, and then the features are mapped to different semantic spaces through attention pooling operation; finally, the MHFA mechanism generates robust voiceprint embedding through the combination of multi-head attention.

[0034] Furthermore, a multi-stage joint optimization training is performed, and the multi-channel voiceprint feature extraction pre-training model is fine-tuned in combination with the AAM loss to generate a multi-channel speech processing model, including:

[0035] By forcibly adding a fixed angle interval between the feature vector and the classification hyperplane on the basis of the traditional Softmax loss, the feature vectors of different categories are made more dispersed in the feature space, generating a multi-channel speech processing model.

[0036] On the other hand, the present invention provides a multi-channel speech processing system, comprising:

[0037] Data acquisition module: acquires multi-channel audio signals and builds a single-channel pre-training model for SSL;

[0038] Model construction module: Based on the multi-channel audio signal, the single-channel pre-trained model of the SSL is structurally optimized to obtain a multi-channel voiceprint feature extraction pre-trained model;

[0039] Model training module: performs multi-stage joint optimization training, fine-tunes the multi-channel voiceprint feature extraction pre-training model in combination with AAM loss, and generates a multi-channel speech processing model;

[0040] Output module: When receiving a multi-channel audio signal to be processed, the multi-channel audio signal to be processed is processed by the multi-channel speech processing model, and a highly distinguishable multi-channel voiceprint feature is output.

[0041] As can be seen from the above technical solution, compared with existing technologies, the present invention provides a multi-channel speech processing method and system. By alternating between intra-channel processing and cross-channel information exchange, the system ultimately fuses the results into a single-channel output. This invention achieves superior performance on the MultiSV dataset, providing an efficient and flexible solution for multi-channel speaker verification. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0043] Figure 1 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0045] The purpose of this invention is to address the difficulty of existing single-channel pre-trained speech models in effectively utilizing the spatial information of multi-microphone arrays by providing a multi-channel voiceprint feature extraction pre-training model. This model utilizes an alternating channel fusion mechanism to dynamically integrate multi-channel spatial features while retaining the powerful representational capabilities of the pre-trained model. This solves the error propagation problem caused by traditional cascade processing methods and significantly improves the performance of multi-channel speaker verification systems in complex acoustic environments.

[0046] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0047] See also Figure 1 The embodiment of the present invention discloses a multi-channel speech processing method, comprising the following steps:

[0048] Acquire multi-channel audio signals and build a single-channel pre-trained model for SSL;

[0049] Based on the multi-channel audio signal, the single-channel pre-trained model of the SSL is structurally optimized to obtain a multi-channel voiceprint feature extraction pre-trained model;

[0050] Perform multi-stage joint optimization training, fine-tune the multi-channel voiceprint feature extraction pre-training model in combination with AAM loss, and generate a multi-channel speech processing model;

[0051] When a multi-channel audio signal to be processed is received, the multi-channel audio signal to be processed is processed by the multi-channel speech processing model, and a highly distinguishable multi-channel voiceprint feature is output.

[0052] In the process of model establishment, firstly, based on the single-channel pre-training model of self-supervised learning SSL, a multi-channel parallel processing architecture is constructed by copying the convolutional neural network (CNN) and Transformer encoder modules of the original model, and feature extraction is performed on the audio signals of each channel respectively; then, after parallel processing, a multi-channel information interaction module (MM module and MS module) is introduced to achieve inter-channel information fusion through the cross-channel attention mechanism while maintaining the temporal integrity of the single-channel features; then, the multi-head attention pooling (MHFA) mechanism is used to reduce the dimensionality and optimize the fused features to generate a highly robust voiceprint embedding; finally, a multi-stage joint optimization training strategy is used, combined with the additive angular cosine loss (AAM) to fine-tune the model and improve the discriminability of the multi-channel voiceprint features. The present invention aims to improve the accuracy and robustness of speaker verification in multi-channel scenarios, significantly reduce the error rate and enhance the system's adaptability to complex acoustic environments.

[0053] Specifically, first, the single-channel pre-training model of the multi-channel voiceprint feature extraction pre-training model SSL constructs a multi-channel parallel processing architecture by copying the CNN and Transformer encoder modules of the original model to realize independent feature extraction of each channel audio signal; then, the MM module and MS module are introduced in the parallel processing stage, and the cross-channel attention mechanism is used to complete the inter-channel information fusion while ensuring the temporal integrity of the single-channel features; then the MHFA mechanism is used to reduce the dimension and optimize the fused features to generate a highly robust voiceprint embedding; finally, through a multi-stage joint optimization training strategy, combined with AAM, the model is fine-tuned to effectively improve the discriminability of the multi-channel voiceprint features.

[0054] Specifically, the present invention includes the following key steps:

[0055] Build a multi-channel parallel processing architecture. First, the architecture of the single-channel pre-trained model based on SSL (such as WavLM Base+) is expanded, and its core feature extraction module (including CNN encoder and Transformer encoder) is copied to form multiple parallel processing branches. Each branch is responsible for processing the audio input signal of an independent channel to achieve parallel feature extraction of multi-channel data. In terms of architectural design, these parallel branches adopt a parameter sharing mechanism, that is, the CNN and Transformer modules of all branches use the same weight parameters. Through parameter sharing, the model can learn voiceprint feature representations with channel invariance; significantly reduce the number of parameters that need to be trained, improve model training efficiency; and maintain the powerful feature extraction capabilities that the pre-trained model has learned on single-channel tasks.

[0056] According to formula (a) (b) (c), the transform-average-concatenation module in the MM module is implemented. For speech data from C different channels, assuming that the feature representation of the cth channel is Where T is the number of time frames and D is the feature dimension. For the transform-average-concatenate module, the following formula is used to implement information interaction between channels:

[0057]

[0058] Where C is the total number of channels, f FC (·; θ) is a fully connected layer + nonlinear activation, is the second layer full connection parameter, MT is the global channel summary feature, represents the channel enhancement feature, is the TAC modular output, FC(x;α)=σ(xW+b), α={W,b}, σ(·) is the nonlinear activation function, || represents the concatenation operation, LN(·) represents the layer normalization operation, and ψ is f FC The set of parameters in the feedforward neural network, including the weight matrix and bias terms. The TAC module captures the common information between multiple channels through the global feature T, and splices it with the features of each channel. After nonlinear transformation and layer normalization, it enhances the expressiveness of each channel feature, thereby realizing cross-channel information interaction. The COATT module in the MM module realizes contextual interaction between channels through a cross-frame joint attention mechanism. Specifically, it generates a channel summary embedding S by summarizing the channel information, and uses the cross-frame joint attention mechanism to contextualize the channel summary and each channel feature, thereby realizing feature interaction and enhancement between channels. The implementation of the MS module includes two methods. The first is to perform weighted averaging of the channel features according to formula (d), and the weights can be automatically learned and optimized based on the training data. The formula is as follows:

[0059]

[0060] Among them, α c is the weight of the cth channel, satisfying H c represents the features of c channels, where T is the number of time frames and D is the feature dimension.

[0061] The second method is the “take first channel” method, which only uses the features of the first channel as output.

[0062] For the MHFA mechanism, the system will use multiple attention heads to model different dimensions of features separately, thereby capturing richer feature information. During the implementation process, the MHFA mechanism first performs dimensionality reduction on the fused features to reduce the feature dimension and computational complexity. Through the attention pooling operation, the features are mapped to different semantic spaces to obtain more representative feature representations. Ultimately, the MHFA mechanism generates a robust voiceprint embedding through the combination of multi-head attention, so that the voiceprint features can maintain good stability and distinguishability in the face of different channel conditions, noise interference and voice changes.

[0063] Fine-tune the model using AAM. Based on the traditional Softmax loss, a fixed angle interval is forced between the feature vector and the classification hyperplane, making the feature vectors of different categories more dispersed in the feature space, further improving the model's discriminative ability.

[0064] Specifically, the present invention addresses the problems of low spatial information utilization and severe error propagation in existing single-channel pre-trained speech models in multi-microphone array scenarios, and provides a multi-channel voiceprint feature extraction pre-training model. Through an innovative alternating channel fusion mechanism, this framework achieves dynamic integration of multi-channel spatial features while retaining the powerful representation ability of the pre-training model, effectively solving the error propagation problem caused by traditional cascade processing methods. Experiments show that this method has achieved the current optimal performance on the MultiSV multi-channel speaker verification dataset, significantly improving the robustness and recognition accuracy of the system in complex acoustic environments. In addition, the modular design enables it to flexibly adapt to pre-training models of different sizes, providing a universal solution for multi-channel speech processing tasks.

[0065] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0066] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multi-channel speech processing method, characterized in that: The steps include: Acquire multi-channel audio signals and build a single-channel pre-trained model for SSL; Based on the multi-channel audio signal, the single-channel pre-trained model of the SSL is structurally optimized to obtain a multi-channel voiceprint feature extraction pre-trained model; Perform multi-stage joint optimization training, fine-tune the multi-channel voiceprint feature extraction pre-training model in combination with AAM loss, and generate a multi-channel speech processing model; When a multi-channel audio signal to be processed is received, the multi-channel audio signal to be processed is processed by the multi-channel speech processing model, and a highly distinguishable multi-channel voiceprint feature is output.

2. A multi-channel speech processing method according to claim 1, characterized in that: The single-channel pre-training model of SSL is structurally optimized to obtain a multi-channel voiceprint feature extraction pre-training model, including: By replicating the CNN and Transformer encoders of a single-channel pre-trained model, multiple parallel processing branches are formed. Each branch is responsible for processing the audio signal of one channel, and multi-channel features are extracted in parallel for multi-channel audio signals. After parallel processing, a multi-channel information interaction module is introduced, and cross-channel audio signal fusion is performed through a cross-channel attention mechanism while maintaining the temporal integrity of single-channel features; The head-attention pooling (MHFA) mechanism is used to reduce the dimension and optimize the fused features to generate voiceprint embeddings. Obtain a multi-channel voiceprint feature extraction pre-training model.

3. A multi-channel speech processing method according to claim 2, characterized in that: During the parallel extraction of multi-channel features, the parameters of each branch remain shared to ensure that the model can learn universal voiceprint features.

4. A multi-channel speech processing method according to claim 2, characterized in that: The multi-channel information interaction module includes an MM module and an MS module; Among them, the MM module also includes a transform-average-concatenation TAC module and a cross-frame joint attention COATT module.

5. A multi-channel speech processing method according to claim 4, characterized in that: The specific implementation of the transform-average-splicing TAC module is as follows: For speech data from C different channels, let the feature of the cth channel be expressed as Where T is the number of time frames and D is the feature dimension. For the TAC module, the following formula is used to realize the information interaction between channels: Where C is the total number of channels, f FC (·; θ) is the fully connected layer and nonlinear activation, is the second layer full connection parameter, MT is the global channel summary feature, represents the channel enhancement feature, is the TAC modular output, FC(x;α)=σ(xW+b), α={W,b}, σ(·) is the nonlinear activation function, || represents the concatenation operation, LN(·) represents the layer normalization operation, and ψ is f FC The set of parameters in a feedforward neural network, including the weight matrix and bias terms.

6. A multi-channel speech processing method according to claim 4, characterized in that: The cross-frame joint attention COATT module includes: By summarizing the channel information, a channel summary embedding S is generated, and the cross-frame joint attention mechanism is used to contextualize the channel summary and each channel feature to achieve feature interaction and enhancement between channels.

7. A multi-channel speech processing method according to claim 4, characterized in that: The specific implementation of the MS module includes two methods. The first method is to perform weighted averaging on the characteristics of each channel. The weights are automatically learned and optimized based on the training data. The formula is: Among them, α c is the weight of the cth channel, satisfying The second method is to use only the features of the first channel as output.

8. A multi-channel speech processing method according to claim 2, characterized in that: The head attention pooling MHFA mechanism includes: Use multiple attention heads to model different dimensions of features and capture richer feature information; Among them, the fused features are first subjected to dimensionality reduction processing, and then the features are mapped to different semantic spaces through attention pooling operation; finally, the MHFA mechanism generates robust voiceprint embedding through the combination of multi-head attention.

9. The multi-channel speech processing method according to claim 1, characterized in that: Perform multi-stage joint optimization training, fine-tune the multi-channel voiceprint feature extraction pre-training model in combination with AAM loss, and generate a multi-channel speech processing model, including: By forcibly adding a fixed angle interval between the feature vector and the classification hyperplane on the basis of the traditional Softmax loss, the feature vectors of different categories are made more dispersed in the feature space, generating a multi-channel speech processing model.

10. A multi-channel speech processing system, characterized in that: include: Data acquisition module: acquires multi-channel audio signals and builds a single-channel pre-training model for SSL; Model construction module: Based on the multi-channel audio signal, the single-channel pre-trained model of the SSL is structurally optimized to obtain a multi-channel voiceprint feature extraction pre-trained model; Model training module: performs multi-stage joint optimization training, fine-tunes the multi-channel voiceprint feature extraction pre-training model in combination with AAM loss, and generates a multi-channel speech processing model; Output module: When receiving a multi-channel audio signal to be processed, the multi-channel audio signal to be processed is processed by the multi-channel speech processing model, and a highly distinguishable multi-channel voiceprint feature is output.