Voice separation method, device and system irrelevant to array geometry

Through virtual microphone estimation and spatial dictionary learning methods, the dependence of multi-channel speech separation technology on fixed array geometry and number of channels is solved, and efficient and accurate speech separation under different array configurations is achieved, adapting to complex acoustic environments and multi-speakers, improving separation performance and adaptability.

CN120452466APending Publication Date: 2025-08-08HUNAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510593680.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The dependence of existing multi-channel speech separation technology on fixed array geometry and number of channels makes it difficult to adapt to the non-structural and dynamic changes in microphone array layout in practical applications, resulting in a decrease in separation quality and robustness.

Method used

Using virtual microphone estimation and spatial dictionary learning methods, virtual channel signals that enhance spatial information density are generated, and combined with spectrum-time features and spatial direction features, speech separation is achieved through a hierarchical dual-path modeling network to adapt to different array geometric configurations.

Benefits of technology

It realizes efficient and accurate speech separation under different microphone array shapes and number of channels, has good generalization capabilities and real-time processing performance, and is suitable for a variety of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452466A_ABST
    Figure CN120452466A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice signal processing, and particularly provides a voice separation method, device and system irrelevant to array geometry. The method is suitable for various microphone array structures, adopts a virtual microphone estimation mechanism to generate a virtual channel signal for enhancing spatial information density, combines frequency spectrum-time features and spatial direction features, and extracts multi-modal representation through a spatial dictionary learning and attention fusion module. And the extracted features are further input into a layered double-path modeling network, and global dependency relationships are respectively modeled on a time axis and a frequency axis, so that high-precision separation of voices of multiple speakers is realized. The system has good array structure adaptivity, can adapt to channel number changes and array shape differences, and has good application value in scenes such as teleconferences, voice recognition front ends and vehicle-mounted voice processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of speech signal processing technology, and specifically relates to a multi-channel speech separation method, device and system applicable to various microphone array structures, and more particularly to a speech separation processing scheme with array geometry independence. Background Art

[0002] With the rapid development of technologies such as automatic speech recognition, voice assistants, remote communications, conferencing systems, and smart homes, speech separation technology, as a key component of front-end speech processing, plays a vital role in applications such as multi-speaker recognition, speech enhancement, and speech interaction. Its core goal is to isolate the clear speech of each speaker from mixed signals in complex environments with multiple speakers speaking simultaneously or in the presence of background noise, providing reliable input for subsequent recognition, understanding, and interaction tasks.

[0003] Currently, multi-channel speech separation technology has become an important direction for improving separation performance. It typically relies on microphone arrays to collect spatial information to enhance inter-speaker differentiation. However, existing mainstream methods generally assume a fixed array geometry and a fixed number of channels, such as linear or circular arrays. While these methods perform well in controlled environments, they face significant challenges in practical deployment.

[0004] Specifically, microphone array layouts in practical applications are often highly unstructured and uncertain. The installation positions of microphones vary significantly across devices, platforms, or spatial scenarios, potentially leading to asymmetric, irregular, or even dynamically changing structural characteristics in the array. For example, in products such as smart speakers, smart TVs, and in-vehicle devices, microphone layouts vary depending on the product design; and in conferencing systems or far-field sound pickup scenarios, the number, distribution, and orientation of microphones often cannot be uniformly set in advance. These factors make methods based on fixed geometric assumptions difficult to adapt to the real-world needs of changing channel counts or diverse array structures.

[0005] Furthermore, when the number of channels changes or microphone failure causes channel loss, the performance of traditional methods will significantly degrade, making it difficult to guarantee separation quality and system robustness. Therefore, achieving stable and efficient speech separation without relying on a pre-set array structure has become a critical technical issue that needs to be addressed in the current field of speech processing. Summary of the Invention

[0006] This invention aims to address the limitations of existing multi-channel speech separation methods, which rely on fixed array geometry and channel counts, and thus face practical limitations. It proposes a method and system for speech separation that is array-independent and employs unified spectral-spatial modeling. This method achieves efficient and accurate speech separation in a variety of microphone array geometries, demonstrating excellent generalization and real-time processing performance.

[0007] A method for speech separation that is independent of array geometry comprises the following steps:

[0008] Acquire a mixed speech signal collected by a multi-channel microphone array of arbitrary geometric structure, and perform short-time Fourier transform on the speech signal to obtain a complex spectrogram;

[0009] Perform interpolation between adjacent microphone channels to generate multiple virtual microphone signals to enhance spatial directional information density;

[0010] Based on the complex spectrograms of the original channel and the virtual channel, the spectrum-time features and spatial direction features are extracted and fused through the attention mechanism to obtain a unified multimodal feature tensor;

[0011] The fused features are input into a hierarchical dual-path modeling network to model long-range dependencies on the time axis and frequency axis respectively, and output plural spectrograms of multiple speakers;

[0012] An inverse short-time Fourier transform is performed on the complex spectrogram to generate time-domain speech signals of multiple speakers.

[0013] The step of generating a virtual microphone signal comprises:

[0014] Calculate the phase difference and amplitude response between adjacent microphones;

[0015] Based on the phase difference and amplitude response, a preset number of virtual microphone signals are generated between adjacent channels using a weighted interpolation method.

[0016] The extraction of the spatial direction features adopts a spatial dictionary learning module, which specifically includes:

[0017] Construct a complex vector at each time frame and frequency point and map it to a preset spatial dictionary;

[0018] Based on the directional response normalization and amplitude weighting mechanism, the multi-directional spatial response embedding vector is obtained.

[0019] The hierarchical dual-path modeling network includes:

[0020] The multi-level Patch Merging layer is used for feature downsampling;

[0021] The dual-path attention modeling module includes TConformer and FConformer submodules, which model along the time dimension and frequency dimension respectively;

[0022] The Patch Expanding layer is used for feature reconstruction and outputs multi-speaker decoding features.

[0023] Also includes:

[0024] The network model is trained based on the speaker permutation invariant training mechanism (PIT) and the scale-invariant signal-to-noise ratio (SI-SDR) loss function.

[0025] An array geometry-independent speech separation device, comprising:

[0026] A sound collection unit, used to collect mixed voice signals through a multi-channel microphone array;

[0027] a virtual microphone estimation unit, configured to generate a plurality of virtual microphone signals according to a microphone channel layout;

[0028] a spectrum and spatial feature extraction unit for extracting spectrum-time features and spatial direction features and performing feature fusion; a speech separation unit for modeling the fused features through a hierarchical dual-path network and outputting complex spectrograms of multiple speakers; and an audio restoration unit for performing an inverse Fourier transform on the spectrograms to generate separated time-domain speech signals;

[0029] The aforementioned method is used to achieve speech separation that is independent of array geometry.

[0030] The virtual microphone estimation unit comprises:

[0031] Phase difference calculation subunit;

[0032] Amplitude response calculation subunit;

[0033] The interpolation generation subunit is used to generate a virtual channel based on the phase and amplitude results.

[0034] The spectrum and spatial feature extraction unit includes:

[0035] Two-dimensional convolutional neural network for extracting spectral-temporal features;

[0036] Spatial dictionary learning module, used to extract spatial direction features;

[0037] The attention fusion module is used to fuse the two types of features.

[0038] A speech separation system comprises the aforementioned speech separation device, and further comprises:

[0039] Speaker tracking unit, used to extract voiceprint features based on the separated speech signal and perform identity recognition and location tracking;

[0040] The control and display interaction unit is used to drive the conference system terminal to complete multimodal interaction functions such as camera following, screen switching and speaker annotation based on the tracking results.

[0041] The speech separation method of the present invention comprises the following steps:

[0042] 1. Data acquisition step: Voice data is collected using a multi-channel microphone array. The geometry of the microphone array can be linear, circular, square, triangular, or other irregular shapes, with a variable number of channels, suitable for various practical application scenarios.

[0043] 2. Virtual Microphone Estimation: Based on the geometry of the microphone array and the positional information between adjacent microphones, an interpolation method is used to generate virtual microphone signals, increasing the spatial information density of the array. Specifically, the phase and amplitude of the virtual microphone signals are linearly interpolated from adjacent microphones to improve spatial resolution. This virtual microphone generation is independent of the array geometry and is applicable to various array configurations.

[0044] 3. Spectral and spatial feature extraction step: A two-dimensional convolutional neural network (2D-CNN) is used to extract the local spectrum-time features of each channel to obtain a spectral feature map. At the same time, the spatial dictionary learning (SDL) module is used to capture the spatial information between multi-channel signals. The SDL module projects the multi-channel complex spectrogram onto a hypersphere by constructing a learnable complex kernel matrix to generate a spatial embedding vector for more accurate spatial feature modeling. The combination of spectral and spatial features enables the model to fully utilize both time-frequency and spatial information, improving the accuracy of speech separation.

[0045] 4. Hierarchical dual-path architecture modeling steps: A hierarchical dual-path structure is introduced to model dependencies along the time axis and frequency axis respectively. First, the time-frequency dimension is gradually reduced through block merging operations to reduce computational complexity. Then, a dual-path encoder (such as the dual-path Conformer module) is used to capture the dependencies between features on the time axis and frequency axis respectively, enhancing the correlation between time-frequency features. Finally, a symmetric deconvolution operation is performed to restore the feature map to its original resolution, outputting a high-fidelity separated speech signal.

[0046] 5. Model Training and Validation: The processed data is fed into an array geometry-independent speech separation model based on unified spectral-spatial modeling for model training and validation. Leveraging this large-scale training data, the model learns complex spectral and spatial feature mappings, improving separation performance.

[0047] The present invention also provides an array geometry-independent speech separation device, comprising:

[0048] Sound collection unit: used to collect multi-channel voice signals through a multi-channel microphone array.

[0049] Virtual microphone estimation unit: used to generate virtual microphone signals to increase spatial information density.

[0050] Spectrum and spatial feature extraction unit: used to extract spectrum features and spatial features.

[0051] Speech separation unit: used to perform speech separation processing and output the separated speech signal.

[0052] This method, applicable to various microphone array configurations, employs a virtual microphone estimation mechanism to generate virtual channel signals with enhanced spatial information density. Combining spectral-temporal features with spatial directional features, it extracts multimodal representations through spatial dictionary learning and an attention fusion module. These features are then fed into a hierarchical dual-path modeling network to model global dependencies on both the time and frequency axes, enabling high-precision multi-speaker speech separation. The system exhibits excellent array structure adaptability, adapting to changes in channel count and array shape, and has promising applications in scenarios such as teleconferencing, speech recognition front-ends, and in-vehicle speech processing.

[0053] The advantages of the present invention are:

[0054] 1. Array geometry independence: This method is independent of fixed microphone array geometry and channel counts and is adaptable to microphone arrays of varying array shapes (e.g., linear, circular, and irregular) and channel counts. It demonstrates good generalization even in unseen array geometries and single-channel scenarios, meeting practical application requirements.

[0055] 2. Efficient Spatial Information Utilization: Virtual microphone estimation technology generates additional virtual microphone signals, enhancing the density of spatial information and improving spatial resolution. Combined with a spatial dictionary learning module, it accurately captures the spatial features between multi-channel signals, improving the accuracy of speech separation.

[0056] 3. Reduced computational complexity: The introduction of a hierarchical dual-path architecture reduces the time-frequency dimension through block merging, thereby reducing computational complexity. Furthermore, by modeling dependencies separately on both the time and frequency axes, the model improves computational efficiency while maintaining performance, making it suitable for real-time speech separation applications.

[0057] 4. Wide Range of Applications: This invention can be widely applied in various scenarios, including intelligent voice assistants, autonomous driving, remote conferencing, speech recognition, and smart homes. It can maintain excellent speech separation performance in complex acoustic environments and with multiple speakers, demonstrating high practical value.

[0058] Compared to existing technologies, this invention achieves consistent separation performance across diverse array geometries and channel counts by introducing virtual microphone estimation and spatial dictionary learning. The application of a hierarchical dual-path architecture effectively reduces computational complexity, meeting the demands of real-time processing. Overall, this invention significantly improves the accuracy, adaptability, and efficiency of speech separation, overcoming the limitations of existing technologies.

[0059] In summary, the present invention provides an array geometry-independent speech separation method and system with unified spectrum-space modeling, which can achieve efficient and accurate speech separation under various microphone array configurations and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The accompanying drawings are schematic illustrations of embodiments of the present invention and are used to assist in understanding the technical solutions of the present invention. The accompanying drawings and the specification constitute part of this application and are used to explain the present invention but do not constitute a limitation of the present invention.

[0061] Figure 1 An overall flow chart of an array geometry-independent speech separation method provided by an embodiment of the present invention;

[0062] Figure 2 Schematic diagram of the structure of a virtual microphone estimation module in an embodiment of the present invention;

[0063] Figure 3 Schematic diagram of the structure of the spectrum and spatial feature fusion process in an embodiment of the present invention;

[0064] Figure 4 A structural block diagram of a hierarchical dual-path modeling network according to an embodiment of the present invention;

[0065] Figure 5 Schematic diagram of the internal structure of the TFConformer module in an embodiment of the present invention;

[0066] Figure 6 A module structure diagram of an array geometry-independent speech separation device provided by an embodiment of the present invention;

[0067] Figure 7 This is a structural diagram of a voice processing system applied to a conference system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0068] In order to facilitate a full understanding of the technical solution of the present application, the proposed array geometry-independent speech separation method and system are described in detail below in conjunction with the accompanying drawings and multiple embodiments. The specific implementation process provided in this application is only a preferred embodiment. The relevant module configuration, network structure and modeling process can be replaced, deformed or expanded without deviating from the technical essence of this application, and should all be included in the scope of protection of this application. In addition, the functional modules and sub-algorithm processes involved in some implementation structures are highly combinable and can be adapted to a variety of actual deployment scenarios, such as remote conferencing, microphone array front-ends, speech recognition pre-processing, etc., and do not constitute undue restrictions.

[0069] To improve the adaptability and consistency of speech separation systems across different array configurations and channel counts, this application proposes a joint spectral-temporal-spatial modeling method based on virtual microphone estimation and spatial dictionary learning (SDL). This method achieves high-fidelity, low-complexity multi-speaker speech separation by enhancing spatial directional information density and fusing spectral-spatial features, without relying on a fixed array geometry.

[0070] The speech separation method provided in this application is applicable to linear array, circular array, directional array and irregular microphone array structures, and can maintain stable performance under different sound source numbers and signal-to-noise ratio conditions. Figure 1 Taking the process shown in the figure as an example, the main processing steps of this method are described in detail:

[0071] First embodiment:

[0072] This embodiment provides an array structure-independent speech separation method based on virtual microphone estimation and spatial dictionary learning, aiming to cope with the impact of changes in microphone array shape and channel number on speech separation performance in real scenarios. Figure 1 The process shown in FIG. 1 includes the following steps:

[0073] S1: Time-frequency transformation of input audio

[0074] First, we obtain the multi-channel mixed speech signal collected by any microphone array and perform a short-time Fourier transform (STFT) on it to convert it from the time domain to a complex spectrogram. The transformed input is represented as a three-dimensional tensor X∈C T×F×C , where T represents the number of frames, F represents the number of frequency channels, and C represents the number of microphone channels. This step provides the basis for subsequent spectrum modeling and spatial feature extraction.

[0075] S2: Virtual microphone signal generation (VME module)

[0076] To enhance spatial resolution and support inputs with varying channel counts, this method introduces a virtual microphone estimation module. This module interpolates between each pair of adjacent real microphone signals to generate $M$ virtual channel signals. The virtual microphones are distributed as follows:

[0077]

[0078] where N v Indicates the total number of virtual channels to be generated, N p = C is the number of adjacent real microphone pairs, n i is the number of virtual channels allocated to the i-th pair. This step ensures that the virtual microphones are evenly distributed in the array structure and adapts to various array configurations.

[0079] S3: Spectral and spatial feature extraction

[0080] This step mainly includes two submodules: spectral feature learning and spatial dictionary learning, which respectively model the structural information of the input audio from the frequency domain and spatial perspectives, forming Figure 3 The upper part of the flow chart shown.

[0081] Spectral feature extraction: A two-dimensional convolutional neural network is used to extract local spectro-temporal pattern features from the reference channel (usually the 0th channel) to enhance the ability to model the time-frequency structure of speech.

[0082] Spatial dictionary learning (SDL): Construct the complex vector x(t,f)∈C at time frame t and frequency point f from the M-channel complex spectrum M , and mapped to the spatial dictionary D∈C M×N On top, we get the embedding vector:

[0083]

[0084] a n (t,f)=x avg (t,f)·a n '(t,f)

[0085] where a n (t,f) is the spatial response value of the nth dictionary direction. Through normalization and amplitude weighting, this module can effectively model spatial difference information.

[0086] S4: Spectral-spatial feature fusion

[0087] In order to unify the spectrum and spatial information, this method Figure 3The attention mechanism fusion module is introduced at the position shown below to perform weighted integration of the two types of features. The fusion process uses a combination of local attention and global attention to automatically learn the relative importance of spectral features and spatial features, and finally obtain the fused tensor X∈R T×F×E , for subsequent separation network use.

[0088] S5: Hierarchical dual-path modeling (e.g. Figure 4 and Figure 5 shown)

[0089] like Figure 4 As shown in the figure, the fused spectral-spatial features are input into the hierarchical dual-path modeling module, which adopts a pyramidal structure design to reduce computational complexity and enhance cross-scale modeling capabilities. Specifically, it includes the following stages:

[0090] First, the input features are sequentially downsampled and merged into adjacent time-frequency blocks through multiple Patch Merging layers to achieve spatial resolution compression of the feature map. After each merging, a dual-path attention modeling module TFConformer (such as Figure 5 As shown), the module consists of two stacked layers.

[0091] like Figure 5 As shown, each TFConformer module consists of two key submodules: TConformer and FConformer, which model attention in the time (T) and frequency (F) dimensions, respectively. First, TConformer focuses on temporal dependencies across frames within a fixed frequency band; then, FConformer models structural resonances across frequency points within a fixed time frame. Each submodule uses residual connections and 2D convolutions for feature transformation, ensuring stable information transfer and flexible structure.

[0092] After the encoding process is complete, the model gradually restores the spatial resolution of the feature map through a symmetrically structured Patch Expanding layer. Finally, it connects to a two-dimensional convolutional (Conv2d) layer and a linear layer (Linear) to output the decoded speaker separation results.

[0093] S6: Inverse time-frequency transform output

[0094] Finally, the decoding module maps the high-order fusion features into complex spectrogram outputs of multiple speakers, as shown in Where K is the number of speakers. The frequency domain output is converted back to the time domain waveform through the inverse STFT (iSTFT) operation to complete the final separation.

[0095] This method uses virtual microphones to compensate for channel shortages, spatial dictionaries to model frequency differences, an integrated attention mechanism to capture multimodal features, and a hierarchical Conformer architecture to achieve joint time-frequency modeling. It has good versatility and practical adaptability, and exhibits excellent separation effects under multiple channel numbers and microphone array types.

[0096] To achieve the end-to-end optimal performance of the above-mentioned speech separation model, the network model described in this embodiment needs to be trained offline. The training process is shown in the figure and specifically includes the following steps:

[0097] S201: Data preparation and simulation construction

[0098] Training samples were constructed using a standard speech corpus, and various microphone array configurations (including but not limited to circular and linear arrays) were constructed through acoustic simulation. Various degrees of reverberation (with a reverberation time range of 0.1–1.0 seconds) and diffuse noise (with a signal-to-noise ratio range of 10–20 dB) were added during the simulation to improve the model's generalization in real-world scenarios. Multi-channel speech mixture data was generated by the array simulation tool and used for subsequent model training.

[0099] S202: Time-Frequency Transformation and Feature Construction

[0100] A short-time Fourier transform (STFT) is performed on the mixed speech signal for each channel, with a 32ms Hanning window and a 16ms frame shift. This yields a complex spectrogram tensor as the model input. To accommodate inputs with varying numbers of channels, the input channel dimensions must be standardized to ensure consistent data format.

[0101] S203: Model module initialization and structure configuration

[0102] Initialize the model's structural parameters, including the dictionary size N = 64 in the spatial dictionary learning module, the fusion feature dimension E = 64, and the window size of the Patch Merging layer set to [1×1, 2×2, 2×2]. The Conformer module follows the standard architecture, including a multi-head attention mechanism, convolutional blocks, and feedforward layers. All model parameters are assigned using the Xavier initialization method to ensure training stability.

[0103] S204: Loss function definition and optimization goal setting

[0104] We use the speaker permutation invariant training (PIT) mechanism commonly used in the field of speech separation and use SI-SDR (Scale-Invariant Signal-to-Distortion Ratio) as the main optimization objective. The loss function is defined as follows:

[0105]

[0106] Among them, s k represents the kth true speaker signal, represents the corresponding predicted output signal, represents the set of all permutations of all K speakers.

[0107] S205: Optimization Algorithms and Training Strategies

[0108] The Adam optimizer is used to perform back propagation optimization on the model, and the initial learning rate is set to 10 -3 The total number of training rounds is 100. During the training process, the learning rate is dynamically adjusted in combination with the cosine annealing strategy. At the same time, the SI-SDRi indicator is monitored on the validation set, and an early stopping mechanism is set to avoid overfitting.

[0109] S206: Evaluation Metrics and Model Selection

[0110] After each round of training, the model is evaluated under seen and unseen microphone array configurations, and the model with the best SI-SDRi performance in the validation set is finally selected as the deployment version.

[0111] Second embodiment:

[0112] In the first embodiment, a method for speech separation that is independent of array geometry has been proposed. To implement the specific functions of this method, this application further provides a corresponding device for speech separation that is independent of array geometry. This device can adapt to microphone arrays of various structures and has good versatility and deployment flexibility.

[0113] Since the device functions of this embodiment correspond one-to-one to the method steps in the first embodiment, the following description will mainly focus on the functional modules, and specific details can be found in the method section.

[0114] The array geometry-independent speech separation device provided by this application is as follows: Figure 6 Shown, including:

[0115] Sound Collection Unit: This module collects mixed speech signals using a multi-channel microphone array. The array can be linear, circular, or other custom geometric shapes. This module normalizes the collected signals based on their channel dimensions to ensure stability during subsequent processing.

[0116] Virtual Microphone Estimation Unit: This unit is used to generate supplementary virtual microphone signals based on the geometric layout of the microphone array and the positional relationship between each channel, thereby improving the spatial feature density. This unit further includes:

[0117] Phase difference calculation subunit: used to estimate the phase difference of adjacent channel signals based on the geometric position relationship between microphones;

[0118] Amplitude response calculation subunit: used to estimate the corresponding amplitude change according to the array shape and channel distance;

[0119] Virtual microphone generation subunit: combines phase and amplitude to synthesize virtual channel signals through weighted interpolation method to enhance the spatial representation capability of the array.

[0120] Spectral and spatial feature extraction unit: used to jointly encode the original multi-channel spectrogram and virtual channel data to extract the spectral structure and spatial response. This module may include:

[0121] Local spectrum-time feature extraction subunit: extracts the local time-frequency structure of the short-time spectrogram from the reference channel based on a two-dimensional convolutional network;

[0122] Spatial feature extraction subunit: uses a spatial dictionary learning module to extract directional spatial features from the complex differences between channels;

[0123] Feature fusion subunit: It uses the attention mechanism to perform weighted fusion of spectral and spatial features, and outputs a feature tensor of unified dimension for subsequent network processing.

[0124] Speech separation unit: used to model and decode the fused multi-dimensional feature map and output a clear speaker signal. This module includes:

[0125] Multi-level Patch Merging and TFConformer modules (time-frequency joint modeling), such as Figure 4 As shown;

[0126] Symmetrical Patch Expanding decoder structure, restoring the original resolution layer by layer;

[0127] Finally, the convolution layer and the linear mapping layer output the separated spectrogram of each speaker, which is reconstructed into a time domain waveform by combining iSTFT, as shown in the following figure: Figure 5 shown.

[0128] In addition, the device can adapt to the following array types:

[0129] Linear array structure: The device can calculate the spatial geometric relationship under the linear arrangement based on the channel spacing information and orientation characteristics;

[0130] Circular array structure: Supports constructing spatial features using the array radius and the angle difference of the channels on the arc. The microphone pointing direction is radiated from the center of the circle by default.

[0131] Through the collaborative work of the above modules, this device achieves adaptive speech separation capabilities under different microphone array structures, effectively enhancing the universality and robustness of the algorithm in real-world deployment scenarios.

[0132] Third embodiment

[0133] To adapt to actual remote conferencing, voice interaction, and multi-speaker management scenarios, this embodiment provides a conferencing system based on array geometry-independent speech separation. This system, while achieving highly robust speech separation, further integrates speaker tracking and control interaction functions to enhance the overall intelligence and interactive experience.

[0134] like Figure 7 As shown, this conference system includes:

[0135] Speech separation subsystem (composed of units 310–350):

[0136] This subsystem is used to complete the pre-processing of multi-channel speech and the separation of speaker speech. Its internal structure and functions have been described in detail in the second embodiment, and will not be repeated here. You can directly refer to the above content.

[0137] Speaker Tracking Unit 360:

[0138] The speaker tracking unit is used to perform identity recognition and position estimation based on the separated speech signal, and realize dynamic monitoring of the current speaker. This unit includes the following submodules:

[0139] Speech feature extraction subunit 361: extracts speech embedded information (such as voiceprint features) for identity recognition from the separated speech;

[0140] Identity matching and modeling subunit 362: realizes speaker identity confirmation based on static registration or dynamic clustering, and supports recognition of multiple people speaking alternately;

[0141] Trajectory positioning subunit 363: combines the timestamp with the array geometry to estimate the speaker's activity trajectory in the space and outputs its real-time position in the conference space.

[0142] The speaker tracking unit supports collaboration with external vision systems and can also independently achieve spatial modeling through acoustic signals.

[0143] Control and display interaction unit 370:

[0144] The control and presentation interaction unit is used to receive the identity and location information from the speaker tracking unit and drive the conference terminal to complete multimodal feedback, including but not limited to:

[0145] The camera automatically follows and focuses the picture;

[0146] Speaker highlighting in multi-screen displays;

[0147] Synchronize with the voice focus switching of the remote conference platform, etc.

[0148] This unit can work with conference room management systems or video encoding and decoding equipment through interfaces to enhance remote interaction effects and user experience.

Claims

1. A speech separation method that is independent of array geometry, characterized in that: The following steps are involved: Acquire a mixed speech signal collected by a multi-channel microphone array of arbitrary geometric structure, and perform short-time Fourier transform on the speech signal to obtain a complex spectrogram; Perform interpolation between adjacent microphone channels to generate multiple virtual microphone signals to enhance spatial directional information density; Based on the complex spectrograms of the original channel and the virtual channel, the spectrum-time features and spatial direction features are extracted and fused through the attention mechanism to obtain a unified multimodal feature tensor; The fused features are input into a hierarchical dual-path modeling network to model long-range dependencies on the time axis and frequency axis respectively, and output plural spectrograms of multiple speakers; An inverse short-time Fourier transform is performed on the complex spectrogram to generate time-domain speech signals of multiple speakers.

2. The method according to claim 1, characterized in that The step of generating a virtual microphone signal comprises: Calculate the phase difference and amplitude response between adjacent microphones; Based on the phase difference and amplitude response, a preset number of virtual microphone signals are generated between adjacent channels using a weighted interpolation method.

3. The method according to claim 1, characterized in that The extraction of the spatial direction features adopts a spatial dictionary learning module, which specifically includes: Construct a complex vector at each time frame and frequency point and map it to a preset spatial dictionary; Based on the directional response normalization and amplitude weighting mechanism, the multi-directional spatial response embedding vector is obtained.

4. The method according to claim 1, wherein The hierarchical dual-path modeling network includes: The multi-level Patch Merging layer is used for feature downsampling; The dual-path attention modeling module includes TConformer and FConformer submodules, which model along the time dimension and frequency dimension respectively; The Patch Expanding layer is used for feature reconstruction and outputs multi-speaker decoding features.

5. The method according to any one of claims 1 to 4, characterized in that Also includes: The network model is trained based on the speaker permutation invariant training mechanism (PIT) and the scale-invariant signal-to-noise ratio (SI-SDR) loss function.

6. An array geometry-independent speech separation device, characterized in that: include: A sound collection unit, used to collect mixed voice signals through a multi-channel microphone array; a virtual microphone estimation unit, configured to generate a plurality of virtual microphone signals according to a microphone channel layout; Spectrum and spatial feature extraction unit, used to extract spectrum-time features and spatial direction features, and perform feature fusion; A speech separation unit that models and fuses features through a hierarchical dual-path network and outputs complex spectrograms for multiple speakers; an audio restoration unit, configured to perform an inverse Fourier transform on the spectrogram to generate a separated time-domain speech signal; The method according to any one of claims 1 to 5 is used to achieve speech separation that is independent of array geometry.

7. The device according to claim 6, characterized in that The virtual microphone estimation unit comprises: Phase difference calculation subunit; Amplitude response calculation subunit; The interpolation generation subunit is used to generate a virtual channel based on the phase and amplitude results.

8. The device according to claim 6, characterized in that The spectrum and spatial feature extraction unit includes: Two-dimensional convolutional neural network for extracting spectral-temporal features; Spatial dictionary learning module, used to extract spatial direction features; The attention fusion module is used to fuse the two types of features.

9. A speech separation system, characterized in that: The device for separating speech according to any one of claims 6 to 8, further comprising: Speaker tracking unit, used to extract voiceprint features based on the separated speech signal and perform identity recognition and location tracking; The control and display interaction unit is used to drive the conference system terminal to complete multimodal interaction functions such as camera following, screen switching and speaker annotation based on the tracking results.

Citation Information

Cited By

  • Double-end voice separation method and system suitable for handheld device, handheld device and storage medium

    CN122201330A