A Multi-Source Audio Fusion and Noise Reduction Method for Smart Meetings

CN122575396APending Publication Date: 2026-08-14HUNAN XINTIAN INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

现有技术中,单个麦克风阵列的盲源分离方法能够在一定程度上分离混合声源,但受限于阵列孔径和空间采样能力,分离出的声源分量往往残留较多串扰,且无法有效整合来自不同阵列的互补信息

Benefits of technology

利用跨阵列声源分组内各初始声源分量的相位差异和幅值相关性计算融合权重矩阵,对组内所有初始声源分量进行加权叠加。该过程中,对相位差矩阵进行高斯核映射得到相位一致性权重,再与幅值相关系数矩阵逐元素相乘并进行归一化处理,生成自适应融合权重矩阵。这种加权策略充分考虑了多阵列信号间的瞬时相位对准程度和整体幅值变化相似性,能够自动抑制相位失配严重或相关性较弱的分量,突出一致性高、相关性强的高质量声源分量。由此得到的增强声源信号有效避免了多阵列叠加引起的频率选择性衰落和声染色现象,在保留声源空间信息和音色特征的同时,显著提升了信号的信噪比和声源分离度。通过基于语音声源信号构建语音活动检测图谱,并按该图谱对非语音声源信号执行形态学滤波操作,提取瞬态噪声脉冲和稳态噪声轮廓。采用开运算处理噪声信号片段,利用扁平结构元素先腐蚀再膨胀,准确剥离出突发性、短时高能量的瞬态脉冲成分;再对剩余部分进行闭运算,先膨胀再腐蚀,获得连续平滑的稳态噪声轮廓。这种形态学分离方式不依赖于噪声的统计假设,能够从非语音片段中同时分离出特性迥异的两种噪声成分,且分离过程对语音活动边界自适应。在此基础上,对应时间点对语音声源信号进行自适应阈值限幅脉冲抑制,仅削除瞬态噪声尖峰而不改变语音波形的其余部分,避免了传统限幅造成的语音畸变;利用稳态噪声轮廓的幅度谱对初步降噪语音进行谱减,并配合底噪阈值约束,有效去除背景噪声且抑制了谱减引起的音乐噪声残余。二者结合使得最终输出的多源融合降噪语音信号在消除瞬态干扰与连续背景噪声的同时,保持语音的自然度和清晰度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575396A_ABST
    Figure CN122575396A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-source audio fusion and noise reduction method for smart conferencing, relating to the field of audio signal processing technology. It includes: acquiring multi-channel raw audio streams synchronously collected by multiple microphone arrays in a smart conferencing scenario; performing blind source separation on each array and cross-array sound source matching based on arrival time difference to form cross-array sound source groups; calculating a fusion weight matrix based on the phase difference and amplitude correlation of each initial sound source component within the group; weighting and superimposing all initial sound source components within the group to generate an enhanced sound source signal corresponding to each physical sound source; identifying speech sound source signals and non-speech sound source signals from these signals; constructing a speech activity detection map based on the speech sound source signals; performing morphological filtering on the non-speech sound source signals according to the map to extract transient noise pulses and steady-state noise contours; and using the transient noise pulses to perform pulse suppression on the speech sound source signals and spectral subtraction on the steady-state noise contours to obtain a multi-source fused and denoised speech signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing technology, specifically to a multi-source audio fusion noise reduction method for smart conferencing. Background Technology

[0002] In intelligent conferencing scenarios, multiple microphone arrays are widely deployed to pick up sound sources from different locations. Speech signals from different speakers and environmental noise overlap in space, resulting in multi-channel audio streams containing complex acoustic interference. To obtain clear conference speech, noise reduction and sound source enhancement processing are required for the signals acquired by multiple arrays. Existing technologies, blind source separation methods using a single microphone array can separate mixed sound sources to some extent, but limited by array aperture and spatial sampling capabilities, the separated sound source components often retain significant crosstalk and cannot effectively integrate complementary information from different arrays. Common cross-array fusion methods simply superimpose the signals separated by each array or select based on energy magnitude. These methods ignore the phase inconsistencies and amplitude correlation differences that exist when different arrays reach the same physical sound source, leading to comb filtering effects or sound source distortion in the weighted superimposed signal, resulting in insufficient intelligibility and naturalness of the enhanced speech. In the noise reduction stage, traditional methods typically treat all non-speech segments detected in the speech signal as noise and perform uniform spectral subtraction or filtering, without distinguishing between transient and steady-state components within the noise. This crude processing method can cause transient impulse noise to leave a diffuse residue after suppression, affecting speech intelligibility. On the other hand, using fixed parameters during steady-state noise removal can easily cause speech damage or introduce musical noise. There is a practical need to solve the problem of high-fidelity fusion and enhancement of multiple sound source components under cross-array conditions, as well as the problem of refined separation and targeted suppression of transient impulse components and steady-state background components in non-speech noise. Summary of the Invention

[0003] The purpose of this invention is to provide a multi-source audio fusion noise reduction method for intelligent conferencing. By grouping sound sources across arrays and calculating the fusion weight matrix using phase differences and amplitude correlation, coherent fusion enhancement of the sound source components of multiple arrays is achieved. Furthermore, by performing morphological filtering based on speech activity detection maps on non-speech sound source signals, transient noise pulses and steady-state noise profiles are separated and then pulse suppression and spectral subtraction are performed respectively to obtain high-quality multi-source fused noise-reduced speech signals.

[0004] The objective of this invention can be achieved through the following technical solutions: This invention provides a multi-source audio fusion and noise reduction method for smart conferencing, comprising: acquiring multi-channel raw audio streams synchronously collected by multiple microphone arrays in a smart conferencing scenario; independently performing blind source separation on each microphone array to obtain the initial sound source component set corresponding to each array; and performing cross-array sound source matching based on the spatial coordinates of each microphone array and the time difference of arrival, associating the initial sound source components belonging to the same physical sound source in different arrays into cross-array sound source groups; for each cross-array sound source group, calculating the fusion weight matrix corresponding to the group based on the phase difference and amplitude correlation of each initial sound source component within the group, and using the fusion weight matrix... All initial sound source components within a group are weighted and superimposed to generate enhanced sound source signals corresponding to each physical sound source. Speech and non-speech source signals are identified from all enhanced sound source signals. A speech activity detection map is constructed based on the identified speech source signals. Morphological filtering is performed on the non-speech source signals according to the speech activity detection map to extract transient noise pulses and steady-state noise contours. The transient noise pulses are used to suppress corresponding time points in the speech source signals to obtain preliminary denoised speech signals. The steady-state noise contours are then used to perform spectral subtraction on the preliminary denoised speech signals to obtain multi-source fused denoised speech signals. By grouping sound sources across space using multiple arrays and adaptive fusion based on phase and amplitude correlation, redundant information between arrays is effectively utilized to suppress environmental interference and reverberation, resulting in enhanced sound source signals with high signal-to-noise ratios. Combining morphological noise decomposition and step-by-step denoising, speech clarity and naturalness are maintained while removing transient and steady-state noise.

[0005] In a preferred embodiment, the blind source separation process includes: constructing an observation signal matrix from the multi-channel raw audio streams obtained for each microphone array; performing centering and whitening preprocessing; executing independent component analysis iterations; updating the unmixing matrix using the negative entropy maximization criterion during iterations until convergence; and multiplying the final unmixing matrix by the preprocessed signal matrix to obtain the separated signal matrix, where each row of this matrix represents the initial sound source component of the array. This negative entropy-driven separation mechanism accurately reconstructs independent sound sources, providing high-quality source components for subsequent cross-array grouping.

[0006] Further preferably, the formation of cross-array sound source groups includes: extracting the arrival time feature of each initial sound source component, which is determined by calculating the peak position of the cross-correlation function between the component and each microphone in the array; for any two initial sound source components in different arrays, calculating their arrival time difference vector, and matching this vector with a theoretical arrival time difference lookup table pre-generated based on the array spatial coordinates and the candidate sound source position grid; when the Euclidean distance between the arrival time difference vector and the theoretical arrival time difference of a certain sound source position in the lookup table is less than a preset matching threshold, the two initial sound source components are marked as candidate corresponding components, and after completing pairwise matching of all arrays, a graph-theoretic connected component algorithm is used to aggregate all mutually matched components into cross-array sound source groups. This matching method based on spatial location information solves the sorting ambiguity in single-array sound source separation and improves the accuracy and robustness of sound source grouping in multi-array joint scenarios.

[0007] As a technical solution of the present invention, the preferred method for calculating the fusion weight matrix is ​​as follows: For each cross-array sound source group, the short-time Fourier transform representation of each initial sound source component within the group is extracted to obtain the complex spectrum in the time-frequency domain; the phase difference matrix and amplitude correlation coefficient matrix between every two components are calculated, wherein the elements of the phase difference matrix are the phase angle differences at the same time-frequency point, and the elements of the amplitude correlation coefficient matrix are the Pearson correlation coefficients of the amplitude sequence over the entire time-frequency domain; a Gaussian kernel mapping is performed on the phase difference matrix to obtain a weight map characterizing phase consistency; this weight map is multiplied element-wise with the amplitude correlation coefficient matrix, and the multiplication result is row-normalized to obtain the fusion weight matrix. This weight matrix synergistically utilizes the linear correlation between the degree of phase consistency and amplitude fluctuation, enabling coherent sound source components to obtain greater fusion gain when superimposed, suppressing incoherent interference and noise components, and further enhancing the fidelity of the output sound source.

[0008] In speech source recognition, the preferred approach is to extract Mel-frequency cepstral coefficient features and fundamental frequency features from each enhanced sound source signal and concatenate them to form a joint feature vector. This feature vector is then input into a pre-trained Gaussian mixture model to calculate the posterior probability of its Gaussian component belonging to the speech category. If the posterior probability is greater than the speech determination threshold, the enhanced sound source signal is marked as a speech source signal; otherwise, it is marked as a non-speech source signal. The construction of the speech activity detection atlas involves dividing the speech source signal into fixed-duration analysis frames. The initial label of speech activity is determined based on the ratio of short-time energy to zero-crossing rate in each frame. Median filtering is then used to smooth and eliminate isolated transitions, forming a binary indicator function whose value changes over time. This function accurately reflects the interval distribution of speech and non-speech activities.

[0009] A preferred implementation of morphological filtering is as follows: The non-speech source signal is multiplied point-by-point with the speech activity detection map to obtain noise signal segments existing only during non-speech activity periods; an opening operation using flattened structuring elements (erosion followed by dilation) is performed on this noise signal segment to extract pulse-type components as transient noise pulses; a closing operation (dilation followed by erosion) is performed on the remaining noise segments after the opening operation to obtain a continuous noise background as a steady-state noise profile. By leveraging the different characteristics of morphological operators, pulse-type transient noise can be effectively decoupled from persistent steady-state noise, laying the foundation for graded noise reduction.

[0010] The impulse suppression processing can be specifically configured as follows: the time position index of transient noise impulses is mapped to the time axis of the speech source signal to determine the target time region contaminated by transient noise; for each time domain waveform segment of the target time region, an adaptive threshold limiting algorithm is used for processing. The adaptive threshold is dynamically calculated based on the mean and standard deviation of the waveform amplitude of the non-noise regions before and after the target time region, and sampling points in the waveform whose amplitude exceeds the adaptive threshold are replaced with the threshold. Thus, while eliminating transient impulses, the natural transition and fine structure of the speech waveform are maintained. The spectral subtraction processing is preferably performed as follows: short-time Fourier transform is performed on the preliminary denoised speech signal and the steady-state noise profile respectively to obtain the corresponding amplitude spectrum and phase spectrum; the product of the steady-state noise profile amplitude spectrum and the over-subtraction factor is subtracted from the amplitude spectrum of the preliminary denoised speech signal frequency by frequency point, and the frequency points that are lower than the noise floor threshold after subtraction are set as the noise floor threshold; the denoised amplitude spectrum and the original phase spectrum are reconstructed by complex number, and synthesized by inverse short-time Fourier transform and overlapping addition method to obtain the multi-source fused denoised speech signal. This two-stage architecture, which first uses pulse limiting and then spectral reduction, employs differentiated processing strategies for different noise characteristics. It significantly suppresses complex noises such as keyboard tapping and steady-state background noise from air conditioning in meeting scenarios, resulting in highly intelligible and natural-sounding voice output.

[0011] The beneficial effects of this invention are: A fusion weight matrix is ​​calculated using the phase differences and amplitude correlations of the initial sound source components within a cross-array sound source group. All initial sound source components within the group are then weighted and superimposed. In this process, a Gaussian kernel mapping is applied to the phase difference matrix to obtain phase consistency weights, which are then multiplied element-wise with the amplitude correlation coefficient matrix and normalized to generate an adaptive fusion weight matrix. This weighting strategy fully considers the instantaneous phase alignment and overall amplitude variation similarity among the multi-array signals, automatically suppressing components with severe phase mismatch or weak correlation, and highlighting high-quality sound source components with high consistency and strong correlation. The resulting enhanced sound source signal effectively avoids frequency-selective fading and sound coloration caused by multi-array superposition, significantly improving the signal-to-noise ratio and sound source separation while preserving the spatial information and timbre characteristics of the sound sources. A speech activity detection atlas is constructed based on the speech source signals, and morphological filtering is performed on non-speech source signals according to this atlas to extract transient noise pulses and steady-state noise contours. This method employs opening operations to process noise signal segments, utilizing flat structuring elements to first erode and then dilate, accurately extracting sudden, short-duration, high-energy transient pulse components. Then, closing operations are performed on the remaining portion, first dilating and then eroding, to obtain a continuous and smooth steady-state noise profile. This morphological separation method does not rely on statistical assumptions about noise and can simultaneously separate two noise components with distinct characteristics from non-speech segments, with the separation process adapting to the boundaries of speech activity. Based on this, adaptive threshold-limited pulse suppression is applied to the speech source signal at corresponding time points, removing only transient noise spikes without altering the rest of the speech waveform, avoiding speech distortion caused by traditional thresholding. Spectral subtraction is performed on the initial denoised speech using the amplitude spectrum of the steady-state noise profile, combined with a noise floor threshold constraint, effectively removing background noise and suppressing residual musical noise caused by spectral subtraction. The combination of these two methods allows the final multi-source fused denoised speech signal to eliminate transient interference and continuous background noise while maintaining the naturalness and clarity of the speech. Attached Figure Description

[0012] The invention will now be further described with reference to the accompanying drawings.

[0013] Figure 1 This is a flowchart of a multi-source audio fusion and noise reduction method for intelligent conferencing; Figure 2 This is a flowchart of cross-array sound source grouping fusion enhancement processing; Figure 3 This is a flowchart of the speech activity detection processing based on joint features and Gaussian mixture model; Detailed Implementation The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0014] See Figure 1 This invention provides a multi-source audio fusion and noise reduction method for smart conferencing, comprising the following steps: acquiring multi-channel raw audio streams synchronously collected by multiple microphone arrays in a smart conferencing scenario; performing blind source separation on each array and cross-array sound source matching based on arrival time difference to form cross-array sound source groups; calculating the fusion weight matrix corresponding to each cross-array sound source group based on the phase difference and amplitude correlation of each initial sound source component within the cross-array sound source group; and using the fusion weight matrix to weight and superimpose all initial sound source components within the group to generate enhanced sound for each physical sound source. The system identifies speech and non-speech source signals from all enhanced sound source signals and constructs a speech activity detection map based on the speech source signals. Morphological filtering is then performed on the non-speech source signals according to the speech activity detection map to extract transient noise pulses and steady-state noise contours. The transient noise pulses are used to suppress corresponding time points in the speech source signals to obtain a preliminary denoised speech signal. Finally, the steady-state noise contours are used to perform spectral subtraction on the preliminary denoised speech signal to obtain a multi-source fused denoised speech signal.

[0015] In practical implementation, multiple microphone arrays synchronously acquire multi-channel raw audio streams in a smart meeting scenario. Each microphone array contains multiple microphone units, and each channel in the multi-channel raw audio stream corresponds to the audio signal acquired by one microphone unit. Blind source separation is performed independently on each microphone array to obtain the initial set of sound source components for each array. For a single microphone array, the multi-channel raw audio stream acquired by that array is constructed into an observation signal matrix. The number of rows in the observation signal matrix equals the number of microphone units in the microphone array, and the number of columns equals the number of sampling points in the audio signal. The observation signal matrix is ​​centered by subtracting the mean of all elements in each row, making the mean of each row zero. The centered observation signal matrix is ​​then preprocessed with whitening to obtain a preprocessed signal matrix. This whitening preprocessing is achieved through eigenvalue decomposition, ensuring that the covariance matrix of the preprocessed signal matrix is ​​an identity matrix.

[0016] Independent component analysis (ICA) iterative operations are performed on the preprocessed signal matrix. In each iteration, the unmixing matrix is ​​updated based on the negative entropy maximization criterion until convergence, yielding the final unmixing matrix. The number of rows and columns in the unmixing matrix is ​​equal to the number of microphone units in the microphone array. The unmixing matrix is ​​updated in each iteration using the following formula:

[0017] in, This represents the updated unmixing matrix; This represents the unmixing matrix before the update; This represents the learning rate, which is 0.01. The reason for setting the learning rate to 0.01 is that, under this value, the unmixing matrix can be observed to tend towards a stable solution with a stable speed and reliable convergence during iteration. Represents the identity matrix with the same dimension as the unmixing matrix; This represents the number of columns in the preprocessed signal matrix, which is equivalent to the total number of time-domain sampling points. For time indexing; This indicates that the unmixed matrix before the update and the preprocessed signal matrix will be compared at the 1st... The instantaneous vector of the separated signal is obtained by multiplying the column vectors at each time point; Represents a nonlinear function. The calculation method is to For each element, the hyperbolic tangent function value is calculated, using the nonlinear function form: , Represents input variables; superscript This indicates the transpose operation. The convergence criterion for the unmixed matrix is ​​that the Frobenius norm between the updated and unmixed matrices is less than a preset convergence threshold, which is set to... .

[0018] The final demixing matrix is ​​multiplied by the preprocessed signal matrix to obtain the separated signal matrix. The number of rows in the separated signal matrix is ​​equal to the number of rows in the final demixing matrix, and the number of columns in the separated signal matrix is ​​equal to the number of columns in the preprocessed signal matrix. Each row in the separated signal matrix represents an initial sound source component. All initial sound source components are combined to form the set of initial sound source components corresponding to the microphone array.

[0019] In practical implementation, for each microphone array, a multi-source matching operation based on the time difference of arrival (TDOA) is performed on the initial sound source component set according to the spatial coordinates of the microphone array. This associates the initial sound source components belonging to the same physical sound source in different microphone arrays, forming cross-array sound source groups. The arrival time feature of each initial sound source component is extracted from the initial sound source component set of each microphone array. The arrival time feature is determined by calculating the peak position of the cross-correlation function between each initial sound source component and each microphone unit in the microphone array. For an initial sound source component in the microphone array, the cross-correlation function between the signal of the initial sound source component and the original audio signal acquired by each microphone unit in the microphone array is calculated. The time delay value corresponding to the peak position of the cross-correlation function in the time delay domain is the arrival time of the initial sound source component to that microphone unit. The arrival times of the same initial sound source component arriving at all microphone units in the same microphone array are combined to form the arrival time feature vector of the initial sound source component under that microphone array.

[0020] For any two initial sound source components from any two different microphone arrays, calculate the arrival time difference vector between these two initial sound source components. The arrival time difference vector is obtained by subtracting the arrival times of each of the two initial sound source components from the arrival times of each microphone unit in their respective microphone arrays. Match the arrival time difference vector with a pre-defined theoretical arrival time difference lookup table, which is pre-calculated based on the spatial coordinates of all microphone arrays and a grid of candidate sound source positions. The candidate sound source position grid covers the entire spatial area of ​​the smart meeting scenario with a pre-defined spatial resolution, and each grid point in the grid represents a candidate sound source position. For each candidate sound source position, based on the spatial coordinates of each microphone unit in all microphone arrays and the spatial coordinates of the candidate sound source position, calculate the theoretical arrival time of the sound source from that candidate sound source position to each microphone unit using the theoretical value of the sound propagation speed. Then, calculate the theoretical arrival time difference between all microphone units and store these theoretical arrival time differences as entries in the theoretical arrival time difference lookup table. When the Euclidean distance between the arrival time difference vector and the theoretical arrival time difference of a sound source location in the theoretical arrival time difference lookup table is less than a preset matching threshold, the two initial sound source components are marked as candidate components of the same physical sound source. The matching threshold is set as an integer multiple of the time difference caused by the distance corresponding to one sampling interval at the speed of sound propagation, typically two times.

[0021] After pairwise matching of all microphone arrays, the connected component algorithm in graph theory is used to aggregate all initial sound source components that are mutually marked as candidate components into cross-array sound source groups. Specifically, an undirected graph is constructed with each initial sound source component as a node and the matching relationship between two initial sound source components as edges. All connected components are found in the undirected graph, and all initial sound source components contained in each connected component constitute a cross-array sound source group.

[0022] In specific implementation, please refer to Figure 2 For each established cross-array sound source group, the short-time Fourier transform representation of each initial sound source component within that group is extracted to obtain the time-frequency domain complex spectrum of each initial sound source component. The short-time Fourier transform employs a preset window function and frame shift parameters. The window function is a Hamming window with a window length of 1024 sampling points and a frame shift of 256 sampling points. The number of points in the discrete Fourier transform of the short-time Fourier transform is consistent with the window length. The time-frequency domain complex spectrum of an initial sound source component is a two-dimensional matrix. The row coordinate indices correspond to the frequency indices, and the column coordinate indices correspond to the time frame indices. The matrix elements are complex numbers; the real part of the complex number represents the cosine component amplitude at the corresponding frequency and time frame, and the imaginary part represents the sine component amplitude at the corresponding frequency and time frame.

[0023] Calculate the phase difference matrix and amplitude correlation coefficient matrix between every two initial sound source components based on the time-frequency domain complex spectrum. For any two initial sound source components, denote one as the first sound source component and the other as the second sound source component. Extract the complex elements corresponding to the same frequency index and time frame index from the time-frequency domain complex spectra of the first and second sound source components. Calculate the difference between the phase angles of the complex elements of the first and second sound source components to obtain the phase angle difference values ​​at the same frequency and time frame indices. Traverse all frequency and time frame indices, arranging the phase angle differences at each time frequency point according to the corresponding frequency and time frame indices to form a phase difference matrix. The total number of rows in the phase difference matrix equals the total number of frequency indices in the short-time Fourier transform, and the total number of columns in the phase difference matrix equals the total number of time frames. The amplitude correlation coefficient matrix between the first and second sound source components is calculated as follows: Amplitude information is extracted from the time-frequency domain complex spectrum of the first and second sound source components, respectively, to obtain the time-frequency domain amplitude matrices of the first and second sound source components. The elements of the time-frequency domain amplitude matrices are the moduli of the complex elements. The amplitude sequences of the first and second sound source components are extracted from the entire time frame sequence according to the same frequency index position, forming two amplitude vectors with a length equal to the total number of time frames. The Pearson correlation coefficient between these two amplitude vectors is calculated to obtain the amplitude correlation coefficient corresponding to the aforementioned frequency index position. All frequency indices are traversed to obtain a column vector, where each element corresponds to the amplitude correlation coefficient of a frequency index. This column vector is the amplitude correlation coefficient matrix between the first and second sound source components. The total number of rows in the amplitude correlation coefficient matrix equals the total number of frequency indices, and the total number of columns in the amplitude correlation coefficient matrix is ​​1.

[0024] The phase difference matrix and amplitude correlation coefficient matrix calculated for each pair of initial sound source components within the cross-array sound source group are input into the fusion weight calculator. The fusion weight calculator performs the following processing: Gaussian kernel mapping is applied to the phase difference matrix between each pair of initial sound source components to obtain the phase consistency weight between these two initial sound source components. The calculation expression for the Gaussian kernel mapping is:

[0025] in, Indicates the first The initial sound source component and the first The initial sound source components are at the frequency index Phase consistency weight at the location, and These are all index numbers of the initial sound source components within the cross-array sound source group. and The values ​​range from 1 to the total number of initial sound source components within the group, and ; Indicates the first The initial sound source component and the first The initial sound source components are at the frequency index The phase angle difference at a given point is determined by the corresponding frequency index in the phase difference matrix. The phase angle difference across all time frame indices is averaged to obtain the result. This represents the bandwidth parameter of the Gaussian kernel. The bandwidth parameter is set to 0.5. The basis for setting the bandwidth parameter to 0.5 is to evaluate the weight decay curves corresponding to different phase difference values ​​under this value, so that the phase difference is within the range of 0 to 1. Within the specified range, the phase consistency weight can transition smoothly and retain effective discrimination. This represents an exponential function with the natural constant e as its base.

[0026] The fusion weight calculator multiplies the phase consistency weights between each pair of initial sound source components calculated above with the corresponding elements in the amplitude correlation coefficient matrix element by element. The corresponding elements in the amplitude correlation coefficient matrix refer to the elements of the first pair. The initial sound source component and the first The amplitude correlation coefficient matrix between the initial sound source components is in the frequency index. The element values ​​at each position are used to construct an unnormalized coefficient matrix. The row indices of the unnormalized coefficient matrix correspond to frequency indices, and the column indices correspond to the initial source component indices. The value in each column of the unnormalized coefficient matrix is ​​the sum of the products of the phase consistency weight and amplitude correlation coefficient of a given initial source component and all other initial source components at the corresponding frequency index. The unnormalized coefficient matrix is ​​then row-normalized so that the sum of all column elements in each row of the normalized matrix equals 1. The matrix obtained after row-normalization is the fusion weight matrix. The total number of rows in the fusion weight matrix equals the total number of frequency indices, and the total number of columns in the fusion weight matrix equals the total number of initial source components within the cross-array source group.

[0027] In the specific implementation, the short-time Fourier transform representations of all initial sound source components within the cross-array sound source group are stacked into a three-dimensional tensor. The three dimensions of the three-dimensional tensor correspond to the frequency index, time frame index, and sound source component index, respectively. The frequency index dimension of the three-dimensional tensor is equal to the total number of frequency indices in the short-time Fourier transform, the time frame index dimension is equal to the total number of time frames in the short-time Fourier transform, and the sound source component index dimension is equal to the total number of initial sound source components within the cross-array sound source group. A broadcast multiplication operation is performed on the fusion weight matrix along the sound source component dimensions of the three-dimensional tensor. The specific method of the broadcast multiplication operation is as follows: the fusion weight matrix corresponding to the frequency index... Harmony Source Component Index The weight values ​​and frequency indices in the 3D tensor are: Time frame index is The sound source component index is The complex elements of each time-frequency point are multiplied together, and the weight values ​​of the same frequency index and the same sound source component index in the fusion weight matrix are applied to the complex elements of all time frame indices under that frequency index and sound source component index. After completing the broadcast multiplication operation, a weighted three-dimensional tensor is obtained, and the dimension of the weighted three-dimensional tensor is the same as that of the original three-dimensional tensor.

[0028] A summation operation is performed along the sound source component dimension on the weighted 3D tensor, summing the values ​​at all sound source component indices corresponding to the same frequency and time frame indices to obtain a two-dimensional fused time-frequency representation. The row indices of the fused time-frequency representation correspond to the frequency indices, the column indices correspond to the time frame indices, and the matrix elements are complex numbers. An inverse short-time Fourier transform (ISFT) is performed on the fused time-frequency representation. The parameters of the IFT are the same as those used for the short-time Fourier transform (SFT). The output of the IFT is an enhanced sound source signal in the time domain. The number of sampling points for the enhanced sound source signal is determined by the number of time frames and the frame shift parameter of the fused time-frequency representation. This enhanced sound source signal is then used as the output signal of the physical sound source corresponding to the cross-array sound source group.

[0029] In specific implementation, please refer to Figure 3For each enhanced sound source signal, Mel-frequency cepstral coefficient features and fundamental frequency features are extracted. The enhanced sound source signal is processed by framing, using a Hamming window as the window function with a window length of 25 milliseconds and a frame shift of 10 milliseconds, resulting in a series of analysis frames. For each analysis frame, Mel-frequency cepstral coefficient features are calculated as follows: after applying a Hamming window to the sampling points within the analysis frame, a discrete Fourier transform is calculated to obtain a linear spectrum; the linear spectrum is then passed through a Mel filter bank consisting of 26 triangular filters, covering a frequency range from 20 Hz to the Nyquist frequency of the speech signal; the logarithmic energy output of each triangular filter is calculated, resulting in a 26-dimensional logarithmic energy vector; a discrete cosine transform is performed on this 26-dimensional logarithmic energy vector, and the first 13 coefficients of the discrete cosine transform result are used to construct the Mel-frequency cepstral coefficient feature vector of that analysis frame, with a dimension of 13.

[0030] Fundamental frequency features are extracted for each enhanced sound source signal using the autocorrelation function method. The enhanced sound source signal is divided into frames with a frame length of 40 milliseconds and a frame shift of 10 milliseconds. For each frame, the autocorrelation function is calculated, and the maximum peak position of the autocorrelation function within the delay range corresponding to the fundamental frequency of 60 Hz to 400 Hz is searched. The delay value corresponding to the maximum peak position is converted into a frequency value to obtain the candidate fundamental frequency value for that frame. The candidate fundamental frequency value sequence is then subjected to median smoothing with a median smoothing window length of 5 frames to eliminate outliers during the fundamental frequency extraction process. The smoothed value is used as the fundamental frequency feature value for that frame. The fundamental frequency feature vector of each enhanced sound source signal is composed of the fundamental frequency feature values ​​of all analyzed frames arranged in chronological order, and the dimension of the fundamental frequency feature vector is consistent with the number of analyzed frames of the enhanced sound source signal.

[0031] The Mel-frequency cepstral coefficient features and fundamental frequency features of the same enhanced sound source signal are concatenated. The concatenation method is as follows: for each analysis frame, the corresponding 13-dimensional Mel-frequency cepstral coefficient feature vector is concatenated with a fundamental frequency feature value from the same time point. The alignment of the fundamental frequency feature value is matched with the analysis frame index of the Mel-frequency cepstral coefficient features. The matched fundamental frequency feature value is appended as an extra dimension to the end of the Mel-frequency cepstral coefficient feature vector, forming a 14-dimensional joint feature vector. The joint feature vectors corresponding to all analysis frames of the entire enhanced sound source signal constitute a joint feature vector sequence.

[0032] Each joint feature vector in the joint feature vector sequence is input into a pre-trained Gaussian mixture model for classification inference. The Gaussian mixture model consists of two independent Gaussian mixture models: a speech category Gaussian mixture model and a non-speech category Gaussian mixture model, operating side-by-side. The speech category Gaussian mixture model contains... Gaussian components The value is set to 16, based on the fact that with 16 Gaussian components, the model can adequately fit the distribution of the speech feature space with limited computational cost. The non-speech category Gaussian mixture model includes... Gaussian components The value is 16, and the selection criteria are consistent with the speech category. Each Gaussian component consists of a mean vector, a covariance matrix, and a weighting coefficient. The covariance matrix is ​​in diagonal form to reduce the number of parameters and improve computational stability.

[0033] The training process of the Gaussian Mixture Model (GMM) is as follows: Training audio data containing both speech and non-speech segments is collected. A 14-dimensional joint feature vector is extracted from the training audio data using the same method described above, and each joint feature vector is labeled with either a speech label or a non-speech label. Two GMMs are trained, one for the set of joint feature vectors corresponding to speech labels and the other for the set of joint feature vectors corresponding to non-speech labels. The training algorithm employs the expectation-maximization (EM) algorithm. The initialization steps of the EEM algorithm for the speech category GMM are as follows: The speech label joint feature vector set is divided using the K-means clustering algorithm. There are *n* clusters, with the center of each cluster serving as the initial value for the mean vector of the corresponding Gaussian component. The covariance matrix of each Gaussian component is initialized as a unit diagonal matrix, and the weight coefficients of each Gaussian component are initialized as follows: In the expectation step, the posterior probability of each joint eigenvector belonging to each Gaussian component is calculated based on the current Gaussian component parameters. In the maximization step, the weight coefficients, mean vector, and covariance matrix of each Gaussian component are re-estimated based on the posterior probabilities. The expectation and maximization steps are repeated until the change in the log-likelihood function value of the model is less than [a certain value]. The training of the Gaussian mixture model for non-speech categories follows the same steps, except that the number of Gaussian components is [number missing]. The training data is a set of joint feature vectors of non-speech labels.

[0034] In classification reasoning, for a joint feature vector Calculate the posterior probability of the joint feature vector belonging to the Gaussian mixture model of the speech category. The formula for calculating the posterior probability is:

[0035] in, Represents a given joint eigenvector The posterior probability that the sound source signal belongs to the speech category is enhanced. This represents the total number of Gaussian components in the Gaussian mixture model for speech categories. In the Gaussian mixture model representing speech categories, the first... The weight coefficients of each Gaussian component are obtained through training and satisfy the following conditions: ; In the Gaussian mixture model representing speech categories, the first... The mean vector of Gaussian components, which has a dimension of 14, consistent with the dimension of the joint feature vector; In the Gaussian mixture model representing speech categories, the first... The covariance matrix of Gaussian components, which is a 14-row, 14-column diagonal matrix; Indicated by As the mean vector, with The probability density function of the multidimensional Gaussian distribution of the covariance matrix in The value at; This represents the total number of Gaussian components in the Gaussian mixture model for non-speech categories. In the non-speech category Gaussian mixture model, the first... The weight coefficients of each Gaussian component are obtained through training and satisfy the following conditions: ; In the non-speech category Gaussian mixture model, the first... The mean vector of Gaussian components, with a dimension of 14; In the non-speech category Gaussian mixture model, the first... The covariance matrix of Gaussian components, which is a 14-row, 14-column diagonal matrix; Indicated by As the mean vector, with The probability density function of the multidimensional Gaussian distribution of the covariance matrix in The value at that location.

[0036] When the posterior probability When the probability exceeds the preset speech determination threshold, the enhanced sound source signal from which the joint feature vector originates is marked as a speech source signal; otherwise, it is marked as a non-speech source signal. The speech determination threshold is set to 0.5, based on the following: when the posterior probability is greater than 0.5, it indicates that the weighted likelihood value of the enhanced sound source signal under the Gaussian mixture model of the speech category exceeds the weighted likelihood value under the Gaussian mixture model of the non-speech category, and therefore, in a probabilistic sense, it has a higher confidence in belonging to the speech category.

[0037] In practical implementation, a speech activity detection atlas is constructed based on the speech source signal. The speech source signal is divided into analysis frames of fixed duration, with a fixed duration of 20 milliseconds and a frame shift of 10 milliseconds. Short-time energy is calculated for each analysis frame by summing the squares of the amplitudes of all sampling points within the frame. Zero-crossing rate is calculated for each analysis frame by counting the number of sign changes between two adjacent sampling points within the frame and normalizing the result by dividing the result by the total number of sampling points in the frame minus one. An initial speech activity marker for each frame is constructed based on the ratio of short-time energy to zero-crossing rate: for the ... Each analysis frame calculates short-time energy. With zero crossing rate ratio ,in Given a very small positive number, avoid division by zero. Values Set a high threshold. and low threshold High threshold With low threshold The value is determined by analyzing the histogram distribution of the ratio of short-time energy to zero-crossing rate in silence segments and speech segments of the speech source signal, ensuring that the ratio is higher than [the specified value]. The frames mainly come from speech activity segments, with a ratio lower than [missing information]. The frames mainly come from silent segments. When At that time, the first The initial voice activity flag for a frame is set to 1, indicating voice activity; when At that time, the first The initial voice activity flag for a frame is set to 0, indicating non-voice activity; when At that time, the initial marker for the speech activity follows the first... The initial tag value for the speech activity of the frame, which is set to 0 if the frame is the first frame.

[0038] Median filtering is applied to the initial speech activity markers of all analysis frames. The median filtering window length is set to 5 frames. That is, for each analysis frame, five values ​​of the initial speech activity markers of that frame and the two frames before and after it are taken, and the median of these five values ​​is used as the smoothed speech activity marker. Median filtering eliminates isolated speech activity marker jumps with a duration of less than three consecutive frames. The smoothed speech activity marker sequence is expanded into a binary indicator function on the time axis. The binary indicator function takes a first value (1) at time points marked as speech activity and a second value (0) at time points marked as non-speech activity. This binary indicator function is used as the speech activity detection atlas.

[0039] In practice, the non-speech source signal is multiplied point-by-point with the speech activity detection map to obtain a noise signal segment containing only non-speech activity time periods. The point-by-point multiplication is implemented as follows: the speech activity detection map is inverted to generate an inverted speech activity detection map, where the value is 1 at non-speech activity time points and 0 at speech activity time points; then, the non-speech source signal and the inverted speech activity detection map are multiplied point-by-point along the time axis. The value of the non-speech source signal at a speech activity time point is multiplied by 0 and set to zero, while the value at a non-speech activity time point is multiplied by 1 and retained, thus obtaining a noise signal segment containing only non-speech activity time periods.

[0040] The opening operation, a mathematical morphology operation, is performed on the noise signal segment. The opening operation uses a flattened structuring element, a one-dimensional discrete sequence. The length of the flattened structuring element is set to 5 sampling points. The length of 5 is chosen because, at a typical audio sampling rate of 16 kHz, the time span corresponding to 5 sampling points is approximately 0.3125 milliseconds. This time span is shorter than the pulse duration of typical transient noises such as keyboard clicks and object collisions in intelligent conferencing scenarios, effectively extracting pulse-type components. The opening operation first performs erosion and then dilation. The erosion operation takes the minimum amplitude of all sampling points within the length of the flattened structuring element as the center, and uses this as the output value of that sampling point. The dilation operation takes the maximum amplitude of all sampling points within the length of the flattened structuring element as the center, and uses this as the output value of that sampling point. The output signal of the opening operation is the extracted transient noise pulse.

[0041] The noise signal segment is subtracted point-by-point from the transient noise pulse to obtain the remaining noise signal segment after the opening operation. A closing operation is then performed on the remaining noise signal segment, using the same flat structuring element, with a length of 5 sampling points. The closing operation first performs dilation and then erosion. The dilation operation is performed by taking the maximum amplitude of all sampling points within the length of the flat structuring element, centered on each sampling point of the remaining noise signal segment, as the output value of that sampling point. The erosion operation is performed by taking the minimum amplitude of all sampling points within the length of the flat structuring element, centered on each sampling point of the signal after the dilation operation, as the output value of that sampling point. The output signal of the closing operation is the extracted continuous and smooth noise background component, serving as the steady-state noise profile.

[0042] In practice, transient noise pulses are used to suppress pulses at corresponding time points in the speech source signal. The time position index of the transient noise pulse is mapped to the same time axis of the speech source signal to determine the target time region of the speech source signal contaminated by transient noise. The method for determining the target time region is as follows: in the transient noise pulse signal, the region of consecutive sampling points where the absolute value of the signal amplitude exceeds a preset pulse detection threshold is marked as a transient noise event; the start time index and end time index of the transient noise event are extracted and directly mapped to the same time index position in the speech source signal. The interval in the speech source signal corresponding to the start time index and end time index is a target time region. The pulse detection threshold is set to half of the maximum value of the absolute value of the transient noise pulse's own signal amplitude, based on the ability to detect the pulse position that protrudes from the base of the waveform.

[0043] For each target time region in the speech source signal, time-domain waveform segments within that target time region are extracted and processed using an adaptive threshold limiting algorithm. The adaptive threshold is dynamically calculated based on the mean and standard deviation of the waveform amplitudes of the non-noise regions before and after the target time region. The method for determining the preceding and following non-noise regions is as follows: In the speech source signal, after tracing back a guard interval from the start time index of the target time region, a segment with a length equal to the reference region length is selected as the preceding non-noise region; after extending backward a guard interval from the end time index of the target time region, a segment with a length equal to the reference region length is selected as the following non-noise region. The guard interval length is set to 10 milliseconds to avoid interference from the transition edges of transient noise pulses on the amplitude statistics of the reference region; the reference region length is set to 50 milliseconds to ensure sufficient sampling points can be obtained within 50 milliseconds to stably estimate the mean and standard deviation of the amplitudes, while maintaining local stationarity. The mean and standard deviation of the waveform amplitudes of all sampling points in the preceding and following non-noise regions are calculated as the mean and standard deviation of the preceding and following non-noise regions.

[0044] The adaptive threshold calculation formula used in the adaptive threshold limiting algorithm is as follows:

[0045] in, This represents the adaptive threshold, which is a positive real number. This represents the mean waveform amplitude of all sampling points within the non-noise region before and after the signal; This represents the standard deviation of the waveform amplitude at all sampling points within the non-noise region before and after the sampling point. This represents the threshold expansion factor, which is set to 3. The basis for this value is: assuming that the waveform amplitude in the non-noise region follows a normal distribution, the probability of data appearing outside three standard deviations is about 0.27%. Treating this part of abnormally high amplitude as transient noise can cover the amplitude shift caused by transient noise with a high probability, while avoiding excessive reduction of the normal speech waveform.

[0046] For each sampling point in the extracted time-domain waveform segment, if the amplitude of the sampling point is greater than... Then the amplitude of the sampling point is replaced with If the amplitude of the sampling point is less than Then the amplitude of the sampling point is replaced with If the amplitude of the sampling point is in If the range is within the specified range, the original amplitude value will remain unchanged.

[0047] After amplitude limiting processing is performed on all target time regions, the processed time-domain waveform segments are spliced ​​with the remaining parts of the speech source signal that are not marked as target time regions in the original time index order. During splicing, a direct time-domain connection method is used at the boundary of the target time region without windowing or overlapping processing to obtain a preliminary denoised speech signal.

[0048] In practice, a Short-Time Fourier Transform (SFT) is performed on the initial denoised speech signal. The SFT uses a Hamming window as the window function, with a window length of 1024 sampling points and a frame shift of 256 sampling points. The Discrete Fourier Transform (DFT) also uses a 1024-point window. After the SFT, the magnitude of each time-frequency point is extracted from the complex spectrum to form the amplitude spectrum of the initial denoised speech. The amplitude spectrum is a two-dimensional matrix, where row indices correspond to frequency indices and column indices correspond to time-frame indices. Matrix elements are non-negative real numbers. Simultaneously, the phase angle of each time-frequency point is extracted from the complex spectrum to form the phase spectrum of the initial denoised speech. The phase spectrum is a two-dimensional matrix with the same dimensions as the amplitude spectrum, where row indices correspond to frequency indices and column indices correspond to time-frame indices. Matrix element values ​​range from [value missing]. .

[0049] A short-time Fourier transform (SFT) is performed on the steady-state noise profile. The parameters for the SFT are exactly the same as those for the initial denoised speech signal: a Hamming window of 1024 sampling points, a frame shift of 256 sampling points, and a discrete Fourier transform (DFT) of 1024 points. The magnitude of each time-frequency point is extracted from the complex spectrum of the SFT of the steady-state noise profile to form the amplitude spectrum of the steady-state noise profile. The amplitude spectrum of the steady-state noise profile is a two-dimensional matrix, where the row indices correspond to the frequency indices, the column indices correspond to the time frame indices, and the matrix elements are non-negative real numbers.

[0050] In practice, the amplitude spectrum after noise reduction is obtained by subtracting the product of the amplitude spectrum of the steady-state noise profile and the over-subtraction factor from the amplitude spectrum of the initial denoised speech at each frequency point. The frequency-point-by-frequency operation is performed as follows: for the frequency index... and time frame index For a given time and frequency point, calculate the value of the denoised amplitude spectrum at that time and frequency point using the following formula:

[0051] in, Indicates the amplitude spectrum after noise reduction at the frequency index. and time frame index The value at that position is a non-negative real number; The amplitude spectrum of the initial denoised speech is represented by the frequency index. and time frame index The value at that position is a non-negative real number; The amplitude spectrum representing the steady-state noise profile at the frequency index and time frame index The value at that position is a non-negative real number; This indicates the over-subtraction factor, which is set to 1.5. The basis for this value is that 1.5 can suppress steady-state noise components with a moderate amplitude in typical conference noise environments, while avoiding distortion of the harmonic structure of the speech signal due to excessive attenuation. This represents the noise floor scaling factor, which is set to 0.05. The value is determined by retaining a noise floor equivalent to 5% of the steady-state noise profile amplitude spectrum at each time frequency point, so that the residual noise after noise reduction has a stable spectral characteristic and avoids introducing intermittent music noise.

[0052] In the above formula, This indicates taking the maximum of the two parameters. When... The calculation result is less than At that time, the value of the denoised amplitude spectrum at that time frequency point is set to , This is the noise floor threshold. When The calculation result is greater than or equal to At that time frequency, the amplitude spectrum after noise reduction remains at the value of [value missing]. After performing the above operations on all frequency indices and all time frame indices, the complete denoised amplitude spectrum is obtained.

[0053] In practice, the amplitude spectrum of the denoised speech and the phase spectrum of the initially denoised speech are reconstructed using complex numbers to obtain the reconstructed time-frequency representation. The complex reconstruction method is as follows: for the frequency index... and time frame index Given a specific time and frequency point, the amplitude spectrum after noise reduction is set to the value at that time and frequency point. As the magnitude of the reconstructed complex number, the values ​​of the phase spectrum of the initially denoised speech at the same frequency index and time frame index are taken. As the phase of the reconstructed complex number, the reconstructed complex value is: ,in The imaginary unit, is a natural constant. The above reconstruction operation is performed on all frequency indices and time frame indices to obtain the reconstructed time-frequency representation, which is a two-dimensional complex matrix where row indices correspond to frequency indices and column indices correspond to time frame indices.

[0054] The reconstructed time-frequency representation is subjected to an inverse short-time Fourier transform. The parameters of the inverse short-time Fourier transform are the same as those of the short-time Fourier transform. The discrete Fourier transform has 1024 points, the window function is a Hamming window with a window length of 1024 sampling points, and the frame shift is 256 sampling points. In the inverse short-time Fourier transform (ISFT) process, the overlapping addition method is used to eliminate inter-frame discontinuities. The specific operation of the overlapping addition method is as follows: For each frame, the reconstructed complex value is subjected to an inverse discrete Fourier transform to obtain a time-domain frame signal with a length of 1024 sampling points; the same Hamming window as in the IFT is applied to the time-domain frame signal for weighting; the weighted time-domain frame signal is superimposed frame by frame with a frame shift of 256 sampling points. During superposition, there is an overlapping area between adjacent frames, and the length of the overlapping area is 1024 minus 256 equals 768 sampling points. The signal value in the overlapping area is the sum of the sampling point values ​​of the two adjacent frames at that position; after all time frames are superimposed, a continuous time-domain signal is generated, which is the multi-source fused noise-reduced speech signal.

[0055] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A multi-source audio fusion and noise reduction method for intelligent conferencing, characterized in that, Includes the following steps: In a smart meeting scenario, acquire multi-channel raw audio streams synchronously collected by multiple microphone arrays, perform blind source separation on each array, and perform cross-array sound source matching based on the time difference of arrival to form cross-array sound source groups. Based on the phase difference and amplitude correlation of each initial sound source component within the cross-array sound source group, the fusion weight matrix corresponding to each cross-array sound source group is calculated, and the fusion weight matrix is ​​used to weight and superimpose all initial sound source components within the group to generate the enhanced sound source signal corresponding to each physical sound source. Speech source signals and non-speech source signals are identified from all enhanced sound source signals. A speech activity detection map is constructed based on the speech source signals. Morphological filtering is performed on the non-speech source signals according to the speech activity detection map to extract transient noise pulses and steady-state noise contours from the non-speech source signals. The transient noise pulse is used to suppress the corresponding time point in the speech source signal to obtain a preliminary denoised speech signal. The steady-state noise profile is then used to perform spectral subtraction on the preliminary denoised speech signal to obtain a multi-source fused denoised speech signal.

2. The multi-source audio fusion and noise reduction method for intelligent conferencing according to claim 1, characterized in that, The formation of cross-array sound source groups specifically includes: Acquire multi-channel raw audio streams synchronously collected by multiple microphone arrays in an intelligent meeting scenario, and perform blind source separation operation independently on each microphone array to obtain the initial sound source component set corresponding to each microphone array; Based on the spatial coordinates of each microphone array, a multi-source matching operation based on the time difference of arrival is performed on the initial sound source component set to associate the initial sound source components belonging to the same physical sound source in different microphone arrays, forming cross-array sound source groups. The step of independently performing blind source separation for each microphone array to obtain the initial set of sound source components corresponding to each microphone array is as follows: For each microphone array, the multi-channel raw audio stream acquired by the microphone array is constructed into an observation signal matrix, and the observation signal matrix is ​​preprocessed by centering and whitening to obtain a preprocessed signal matrix; Perform independent component analysis iterative operation on the preprocessed signal matrix. In each iteration, update the unmixing matrix based on the negative entropy maximization criterion until the unmixing matrix converges to obtain the final unmixing matrix. The final demixing matrix is ​​multiplied by the preprocessed signal matrix to obtain the separated signal matrix. Each row in the separated signal matrix represents an initial sound source component. All initial sound source components are combined to form the initial sound source component set corresponding to the microphone array.

3. The multi-source audio fusion and noise reduction method for intelligent conferencing according to claim 2, characterized in that, Perform a multi-source matching operation based on time difference of arrival on the initial sound source component set, and associate the initial sound source components belonging to the same physical sound source in different microphone arrays to form cross-array sound source groups, specifically: The arrival time feature of each initial sound source component is extracted from the initial sound source component set of each microphone array. The arrival time feature is determined by calculating the peak position of the cross-correlation function between each initial sound source component and each microphone in the microphone array. For any two initial sound source components in any two different microphone arrays, calculate the arrival time difference vector of the two initial sound source components, and match the arrival time difference vector with a preset theoretical arrival time difference lookup table. The theoretical arrival time difference lookup table is pre-calculated based on the spatial position coordinates of the microphone array and the candidate position grid of the sound source. When the Euclidean distance between the arrival time difference vector and the theoretical arrival time difference of a sound source location in the theoretical arrival time difference lookup table is less than a preset matching threshold, the two initial sound source components are marked as candidate components of the same physical sound source. After pairwise matching of all microphone arrays, the connected component algorithm in graph theory is used to aggregate all mutually matched initial sound source components into cross-array sound source groups.

4. The multi-source audio fusion and noise reduction method for intelligent conferencing according to claim 3, characterized in that, Based on the phase difference and amplitude correlation of each initial sound source component within the cross-array sound source group, the fusion weight matrix corresponding to each cross-array sound source group is calculated, specifically as follows: For each cross-array sound source group, the short-time Fourier transform representation of each initial sound source component within the group is extracted to obtain the time-frequency domain complex spectrum of each initial sound source component. The phase difference matrix and amplitude correlation coefficient matrix between each two initial sound source components are calculated based on the time-frequency domain complex spectrum. Each element in the phase difference matrix represents the phase angle difference between the two initial sound source components at the same time-frequency point, and each element in the amplitude correlation coefficient matrix represents the Pearson correlation coefficient of the amplitude sequence of the two initial sound source components in the entire time-frequency domain. The phase difference matrix and amplitude correlation coefficient matrix are input into the fusion weight calculator. The fusion weight calculator first performs Gaussian kernel mapping on the phase difference matrix to obtain the phase consistency weight, then multiplies the phase consistency weight with the amplitude correlation coefficient matrix element by element, and finally normalizes the multiplication result to obtain the fusion weight matrix.

5. A multi-source audio fusion and noise reduction method for intelligent conferencing according to claim 4, characterized in that, The fusion weight matrix is ​​used to weight and superimpose all initial sound source components within the group to generate an enhanced sound source signal corresponding to each physical sound source, specifically: The short-time Fourier transform representations of all initial sound source components within the cross-array sound source group are stacked into a three-dimensional tensor, and the three dimensions of the three-dimensional tensor correspond to the frequency index, time frame index, and sound source component index, respectively. The fusion weight matrix is ​​broadcast multiplied along the sound source component dimension of the three-dimensional tensor to obtain a weighted three-dimensional tensor. Then, a summation operation is performed along the sound source component dimension to obtain the fused time-frequency representation. An inverse short-time Fourier transform is performed on the fused time-frequency representation to obtain an enhanced sound source signal in the time domain, and the enhanced sound source signal is used as the output signal of the physical sound source corresponding to the cross-array sound source group.

6. A multi-source audio fusion and noise reduction method for intelligent conferencing according to claim 5, characterized in that, The process of identifying speech source signals and non-speech source signals from all enhanced sound source signals specifically involves: For each enhanced sound source signal, Mel frequency cepstral coefficient features and fundamental frequency features are extracted, and the Mel frequency cepstral coefficient features and fundamental frequency features are concatenated to form a joint feature vector; The joint feature vector is input into a pre-trained Gaussian mixture model for classification inference. The Gaussian mixture model includes Gaussian components for speech categories and Gaussian components for non-speech categories. The posterior probability of the joint feature vector belonging to the Gaussian component for speech categories is calculated. When the posterior probability is greater than the preset speech determination threshold, the enhanced sound source signal is marked as a speech sound source signal; otherwise, it is marked as a non-speech sound source signal.

7. A multi-source audio fusion and noise reduction method for intelligent conferencing according to claim 6, characterized in that, Based on the aforementioned speech source signals, a speech activity detection map is constructed, specifically as follows: The speech source signal is divided into analysis frames of fixed duration. Short-time energy and zero-crossing rate are calculated for each analysis frame. The initial marker of speech activity for each frame is constructed based on the ratio of short-time energy to zero-crossing rate. Median filtering is performed on the initial speech activity markers of all analysis frames to eliminate isolated speech activity marker jumps, resulting in a smoothed speech activity marker sequence. The smoothed speech activity marker sequence is expanded into a binary indicator function on the time axis. The binary indicator function takes a first value at the time point marked as speech activity and a second value at the time point marked as non-speech activity. The binary indicator function is used as a speech activity detection map.

8. A multi-source audio fusion and noise reduction method for intelligent conferencing according to claim 7, characterized in that, Based on the speech activity detection map, morphological filtering is performed on the non-speech source signal to extract transient noise pulses and steady-state noise contours from the non-speech source signal. Specifically: The non-speech source signal is multiplied point by point with the speech activity detection map to obtain a noise signal segment containing only the non-speech activity time period; The noise signal segment is subjected to an opening operation in mathematical morphology. The opening operation uses a flat structuring element to first erode the noise signal segment and then dilate it to extract the pulse-type components in the noise signal segment as transient noise pulses. A closing operation is performed on the remaining noise signal segment after the opening operation. The closing operation uses a flat structuring element to first perform an expansion operation and then an erosion operation to extract a continuous and smooth noise background component as a steady-state noise profile.

9. A multi-source audio fusion and noise reduction method for intelligent conferencing according to claim 8, characterized in that, The transient noise pulse is used to perform pulse suppression processing on the corresponding time points in the speech source signal to obtain a preliminary denoised speech signal, specifically as follows: The time position index of the transient noise pulse is mapped to the same time axis of the speech source signal to determine the target time region of the speech source signal that is contaminated by transient noise. For each target time region in the speech source signal, a time-domain waveform segment is extracted within the target time region. An adaptive threshold limiting algorithm is used to replace the sampling points in the time-domain waveform segment whose amplitude exceeds the adaptive threshold with the adaptive threshold. The adaptive threshold is dynamically calculated based on the mean and standard deviation of the waveform amplitude in the non-noise regions before and after the target time region. After amplitude limiting processing is performed on all target time regions, the processed time-domain waveform segments are spliced ​​with the remaining parts of the speech source signal that are not marked as time regions to obtain the preliminary denoised speech signal.

10. A multi-source audio fusion and noise reduction method for intelligent conferencing according to claim 9, characterized in that, The initial denoised speech signal is subjected to spectral subtraction using the steady-state noise profile to obtain a multi-source fused denoised speech signal, specifically as follows: Short-time Fourier transforms are performed on the initial denoised speech signal and the steady-state noise profile to obtain the amplitude spectrum and phase spectrum of the initial denoised speech signal and the amplitude spectrum of the steady-state noise profile. The amplitude spectrum after noise reduction is obtained by subtracting the product of the amplitude spectrum of the steady-state noise profile and the over-subtraction factor from the amplitude spectrum of the initial noise reduction speech at each frequency point, and setting all frequency points in the noise reduction amplitude spectrum that are less than the noise floor threshold as the noise floor threshold. The amplitude spectrum after noise reduction is reconstructed by complex number with the phase spectrum of the initial noise-reduced speech to obtain the reconstructed time-frequency representation. The inverse short-time Fourier transform is performed on the reconstructed time-frequency representation, and the overlapping addition method is used to eliminate inter-frame discontinuities to generate a multi-source fused noise-reduced speech signal.