Directional sound pickup method, system, device and storage medium

By combining MVDR beamforming, BFCC feature extraction and pre-trained speech/noise separation DNN, adaptively designing comb filters and using Bark frequency scales for speech enhancement, the problem of poor performance of traditional beamforming technology in non-static noise and reverberation environments is solved, achieving more efficient speech signal clarity and noise suppression.

CN119694333BActive Publication Date: 2025-08-19SHANGHAI HAOYI INFORMATION SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510215640.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-08-19
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

In the prior art, the beamforming technology based on microphone arrays has limited effect when dealing with non-static noise and reverb, making it difficult to effectively suppress background noise, affecting the clarity and intelligibility of the speech signal.

Method used

MVDR beamforming is used to combine BFCC feature extraction and pre-training speech/noise separation DNN, and adaptively design comb filters, and use Bark frequency scale for speech enhancement, fine control is performed through time-frequency mask and Bark gain coefficient, and finally inverse BFCC transformation is performed to improve the directional sound pickup effect.

Benefits of technology

It significantly improves the clarity and intelligibility of the voice signal, effectively suppresses noise interference in non-target directions, and improves the quality of the voice signal and subjective listening quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119694333B_ABST
    Figure CN119694333B_ABST
Patent Text Reader

Abstract

The present application provides a directional sound pickup method, system, device, and storage medium, relating to the field of sound pickup. The method comprises: performing MVDR beamforming processing on the initial speech signal input by the microphone array, and then extracting the BFCC features and fundamental frequency features of the speech signal. The BFCC features are processed using a pre-trained speech / noise separation DNN to obtain a time-frequency mask. Comb filter parameters are adaptively designed based on the fundamental frequency features, and the initial speech signal is comb filtered. The time-frequency mask is used as a weighting coefficient and multiplied with the comb-filtered speech signal and the initial speech signal, respectively, to estimate the target speech energy and noise energy. The frequency-dependent Bark gain coefficient is calculated on the Bark frequency scale and used to enhance the speech signal in the time-frequency domain. Finally, the enhanced speech signal is converted into a time-domain signal through an inverse BFCC transform to obtain the target enhanced speech pointing in the target direction. The above scheme improves the directional sound pickup effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of sound pickup, and in particular to a directional sound pickup method, system, device and storage medium. Background Art

[0002] Speech is one of the primary means of daily communication and information transfer. In real-world environments, speech signals are often accompanied by various background noises, such as ambient noise, reverberation, and the voices of other speakers. These noise interferences severely affect speech clarity and intelligibility, posing significant challenges to applications such as voice interaction and remote communications.

[0003] To improve the quality of speech signals and suppress background noise, various speech enhancement methods have been proposed. A common approach is beamforming technology based on microphone arrays. This technology spatially filters the multi-channel speech signals collected by the microphone array to create a spatial directivity toward the target sound source, thereby extracting the target speech and suppressing noise interference from non-target directions. However, traditional beamforming technology typically only considers the spatial characteristics of the speech signal, and its effectiveness in suppressing unsteady noise and reverberation is limited, resulting in poor sound pickup. Summary of the Invention

[0004] The present application provides a directional sound pickup method, system, device and storage medium, which can improve the directional sound pickup effect.

[0005] In a first aspect, the present application provides a directional sound pickup method, the method comprising:

[0006] Performing MVDR beamforming processing on the initial speech signal input by the microphone array to obtain the initial speech signal after beamforming;

[0007] Extract features from the initial speech signal to obtain BFCC features and fundamental frequency features;

[0008] Use the pre-trained speech / noise separation DNN to process the BFCC features and obtain the time-frequency mask;

[0009] Adaptively designing the parameters of the comb filter according to the fundamental frequency characteristics to obtain a target comb filter, and filtering the initial speech signal based on the comb filter to obtain a comb-filtered speech signal;

[0010] The time-frequency mask is used as a weighting coefficient and multiplied with the speech signal after comb filtering to obtain the energy of the weighted signal, and the energy of the weighted signal is used as the target speech energy estimation value;

[0011] The time-frequency mask is used as a weighting coefficient to multiply the initial speech signal to obtain the noise energy estimate;

[0012] Calculate the frequency-dependent Bark gain coefficient on the Bark frequency scale based on the target speech energy estimate and the noise energy estimate;

[0013] The speech signal after comb filtering is multiplied by the Bark gain coefficient to enhance the initial speech signal in the time-frequency domain to obtain the enhanced speech signal;

[0014] The enhanced speech signal is converted from the Bark frequency scale to a time domain signal using the inverse BFCC transform to obtain the target enhanced speech pointing in the target direction.

[0015] By adopting the above technical solution, the initial speech signal input by the microphone array is processed by MVDR beamforming to form a spatial filter pointing in the direction of the target speech, which can effectively suppress noise interference in non-target directions and lay a good foundation for subsequent feature extraction and speech enhancement.

[0016] The beamformed speech signal's BFCC features and fundamental frequency characteristics are then extracted, leveraging the speech characteristics in both the frequency and time domains. The BFCC features utilize the Bark frequency scale, which better aligns with human hearing and facilitates subsequent speech / noise separation and gain adjustment. The fundamental frequency features provide key parameters for the design of the adaptive comb filter, enabling targeted enhancement of the speech fundamental frequency and its harmonic components.

[0017] Next, the BFCC features are processed using a pre-trained speech / noise separation DNN to generate a time-frequency mask. This DNN fully exploits the distinguishing characteristics of speech and noise in the time-frequency domain, accurately estimating the time-frequency regions dominated by speech and noise. This time-frequency mask serves as an important basis for subsequent energy estimation and gain calculation, enabling more precise control over the strength and effectiveness of speech enhancement.

[0018] On this basis, an adaptive comb filter is designed to filter the speech signal, further enhancing the periodic components of speech and attenuating non-periodic noise. Simultaneously, the target speech energy and noise energy are estimated based on the time-frequency mask, and a Bark gain coefficient is calculated on the Bark frequency scale. This gain coefficient fully considers the energy distribution of speech and noise across different frequency bands, giving greater gain to speech-dominated time-frequency regions and more suppression to noise-dominated regions, thereby preserving the important components of speech while reducing noise.

[0019] Finally, the comb-filtered speech signal is multiplied by the Bark gain coefficient to enhance the speech. The enhanced speech signal is then converted to the time domain using an inverse BFCC transform, resulting in enhanced speech directed in the target direction. Compared to traditional speech enhancement methods, this method leverages the multidimensional characteristics of speech and incorporates the Bark frequency scale, which aligns with human hearing. Consequently, it offers significant advantages in improving speech clarity and subjective listening quality.

[0020] In a second aspect of the present application, a directional sound pickup method system is provided, comprising:

[0021] A data acquisition module is used to perform MVDR beamforming processing on the initial speech signal input by the microphone array to obtain the initial speech signal after beamforming;

[0022] The feature extraction module is used to extract features from the initial speech signal to obtain BFCC features and fundamental frequency features;

[0023] A special processing module is used to process the BFCC features using a pre-trained speech / noise separation DNN to obtain a time-frequency mask;

[0024] A filtering module is used to adaptively design the parameters of a comb filter according to the fundamental frequency characteristics to obtain a target comb filter, and filter the initial speech signal based on the comb filter to obtain a comb-filtered speech signal;

[0025] The target speech energy estimation module is used to multiply the time-frequency mask as a weighting coefficient with the speech signal after comb filtering to obtain the energy of the weighted signal, and use the energy of the weighted signal as the target speech energy estimation value;

[0026] A noise energy estimation module is used to multiply the time-frequency mask as a weighting coefficient with the initial speech signal to obtain a noise energy estimation value;

[0027] A gain coefficient calculation module is used to calculate the frequency-dependent Bark gain coefficient on the Bark frequency scale based on the target speech energy estimate and the noise energy estimate;

[0028] A first signal enhancement module is used to multiply the speech signal after comb filtering by the Bark gain coefficient to enhance the initial speech signal in the time-frequency domain to obtain an enhanced speech signal;

[0029] The second signal enhancement module is used to convert the enhanced speech signal from the Bark frequency scale into a time domain signal using an inverse BFCC transform to obtain a target enhanced speech pointing in the target direction.

[0030] In a third aspect of the present application, a computer storage medium is provided. The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the above method steps.

[0031] In the fourth aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs the above method. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 A schematic diagram of a flow chart of a directional sound pickup method provided in an embodiment of the present application;

[0033] Figure 2 An architectural diagram of a directional sound pickup system provided in an embodiment of the present application;

[0034] Figure 3 This is a schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION

[0035] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.

[0036] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.

[0037] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0038] In order to facilitate understanding of the method and system provided by the embodiments of the present application, before introducing the embodiments of the present application, the background of the embodiments of the present application is first introduced.

[0039] As one of the most natural and convenient means of communication for humans, voice plays an indispensable role in our daily lives and work. However, in real-world applications, voice signals are often contaminated by various background noises, such as ambient noise, reverberation, and interference from other speakers. These noises severely impact speech clarity and intelligibility, posing significant challenges to voice applications such as voice interaction and remote communications.

[0040] To improve the quality of speech signals and suppress background noise interference, various speech enhancement methods have been proposed. Among them, beamforming technology based on microphone arrays has attracted considerable attention. This technology uses a microphone array to collect multi-channel speech signals and spatially filters these signals to create a spatial directivity toward the target sound source, thereby extracting the target speech and suppressing noise from non-target directions.

[0041] While beamforming technology has improved speech signal quality to some extent, traditional beamforming methods still have limitations. These methods typically only consider the spatial characteristics of speech signals, leveraging the geometric layout of the microphone array and the orientation of the sound source to distinguish target speech from noise. However, noise in real-world environments often exhibits non-stationary characteristics and complex reverberation, making it difficult to achieve ideal noise suppression by relying solely on the spatial characteristics of speech.

[0042] Furthermore, traditional beamforming technology often performs poorly when dealing with non-stationary noise and reverberation. Non-stationary noise, such as sudden and intermittent noise, has statistical characteristics that vary over time, making it difficult to effectively eliminate using fixed spatial filters. Reverberation, on the other hand, causes multiple reflections and superposition of speech signals during propagation, resulting in temporal and spatial aliasing of the target speech and noise, making speech enhancement more difficult.

[0043] After the background introduction of the above content, those skilled in the art can understand the problems existing in the prior art. The technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.

[0044] On the basis of the above background technology, further, please refer to Figure 1 , Figure 1A flow chart of a directional sound pickup method provided in an embodiment of the present application is provided. The system can be implemented by a computer program or can be run as an independent tool application. In a preferred embodiment of the present invention, the method can be applied on a server in the embodiment of the present application, but can also be applied to electronic devices such as servers. A directional sound pickup method includes the following steps:

[0045] S101, performing MVDR beamforming processing on an initial speech signal input by a microphone array to obtain an initial speech signal after beamforming;

[0046] In a preferred embodiment of the present invention, firstly, MVDR (minimum variance distortionless response) beamforming processing is performed on the initial speech signal input by the microphone array to obtain the initial speech signal after beamforming.

[0047] MVDR beamforming is a commonly used spatial filtering technique that aims to minimize interference and noise from other directions while preserving the target signal. Specifically, it performs a weighted summation of the multi-channel speech signals collected by the microphone array to form a spatial filter oriented toward the target sound source. The filter achieves a gain of 1 in that direction and minimizes the gain in other directions, thereby extracting the target signal and suppressing noise.

[0048] In this step, the microphone array's steering vector is first determined based on the geometric arrangement of the microphone array and the azimuth of the target speech signal. This vector represents the propagation delay and attenuation differences of the target signal from the sound source to each microphone. Next, the optimal weight coefficients of the MVDR beamformer are calculated using the speech signal's spatial covariance matrix and the steering vector to minimize the variance of its output signal. Finally, the weight coefficients are weighted and summed with the signals from each microphone to obtain the initial speech signal after beamforming.

[0049] MVDR beamforming effectively extracts the target speech signal in the spatial domain and suppresses noise interference from non-target directions, laying a solid foundation for subsequent feature extraction and speech enhancement. Compared to traditional delay-sum beamformers, the MVDR beamformer adaptively adjusts the weights of each signal path, achieving higher noise suppression without introducing speech distortion. This results in a higher signal-to-noise ratio and clarity for the beamformed speech signal.

[0050] It's important to note that the spatial covariance matrix is a statistic that describes the spatial correlation between multiple speech signals. Its diagonal elements represent the power of each signal, while the off-diagonal elements represent the correlation coefficients between different signals. By performing eigendecomposition on the spatial covariance matrix, we can estimate the subspaces of the noise and target signals, and then optimize the weight coefficients of the MVDR beamformer to achieve adaptive noise suppression.

[0051] Based on the above embodiment, as an optional embodiment, performing MVDR beamforming processing on the initial speech signal input by the microphone array to obtain the beamformed initial speech signal includes:

[0052] S201, determining an array manifold vector based on geometric arrangement information of the microphone array and azimuth information of the initial speech signal;

[0053] In a preferred embodiment of the present invention, when performing MVDR beamforming processing on the initial speech signal input by the microphone array, the first step of the method of the present invention is to determine the array manifold vector based on the geometric arrangement information of the microphone array and the azimuth information of the initial speech signal.

[0054] A microphone array consists of multiple microphones arranged in a specific geometric configuration. Common configurations include linear, circular, and planar arrays. The geometric configuration of a microphone array includes parameters such as the number of microphones, their position coordinates, and their spacing. These parameters determine the array's spatial sampling characteristics and directional resolution. The azimuth angle of the initial speech signal represents the angle at which the speech signal arrives at the microphone array—that is, the angle between the speech signal and the normal to the microphone array. This azimuth angle can be estimated using sound source localization algorithms, such as those based on time delay estimation or subspace decomposition.

[0055] The array manifold vector is a vector that represents the response characteristics of a microphone array to speech signals from a specific direction. Its dimension is equal to the number of microphones. The array manifold vector is closely related to the geometric arrangement of the microphone array and the incident direction of the speech signal. Different geometric arrangements and incident directions will result in different array manifold vectors. The purpose of determining the array manifold vector is to optimize the weight vector in the subsequent MVDR beamforming process to maximize the gain of the beamformed speech signal in the target direction while minimizing the gain in other directions, thereby achieving spatial filtering directed in the target direction.

[0056] The specific implementation process is as follows: First, the steering vector of the microphone array is calculated based on the geometric arrangement information of the microphone array. The steering vector is a complex vector that represents the phase difference between the signals received by each microphone when a plane wave is incident on the microphone array from a specific direction. The calculation formula of the steering vector is:

[0057] a(θ) = [1, e^(-j2πfτ_1), ..., e^(-j2πfτ_M-1)]^T

[0058] Where θ is the azimuth angle of the speech signal, f is the frequency of the speech signal, τ_m is the time delay of the speech signal from the reference microphone to the mth microphone, and M is the number of microphones. The time delay τ_m can be calculated based on the position coordinates of the microphones and the azimuth angle of the speech signal.

[0059] The calculated steering vector is then normalized to obtain the unit steering vector d(θ). The purpose of normalization is to eliminate the amplitude differences of the steering vector and make it a unit vector, which facilitates subsequent weight calculation and optimization.

[0060] Finally, the unit steering vector d(θ) is used as the array manifold vector to characterize the response characteristics of the microphone array to the speech signal in the target direction.

[0061] By determining the array manifold vector based on the geometric arrangement of the microphone array and the azimuth angle of the initial speech signal, a spatial response model of the microphone array was established, laying the foundation for subsequent MVDR beamforming processing. The array manifold vector accurately describes the gain characteristics of the microphone array for speech signals in the target direction and is a key factor in achieving spatial filtering and directionality enhancement.

[0062] S202, calculating the weight coefficient of the MVDR beamformer according to the spatial covariance matrix of the initial speech signal and the array manifold vector to obtain the MVDR weight coefficient;

[0063] In a preferred embodiment of the present invention, after determining the array manifold vector, the next step of the method of the present invention is to calculate the weight coefficients of the MVDR beamformer based on the spatial covariance matrix of the initial speech signal and the array manifold vector to obtain the MVDR weight coefficients.

[0064] The spatial covariance matrix represents the correlation and energy distribution of the speech signal received by a microphone array across different microphones. It captures the spatial statistical characteristics of the speech signal and noise, reflecting their spatial distribution across the microphone array. The calculation of the spatial covariance matrix is typically based on short-term analysis of the speech signal, obtained by estimating the speech signal samples over a period of time.

[0065] The Minimum Variance Distortionless Response (MVDR) beamformer is an adaptive beamforming algorithm that aims to minimize the variance of the beamformed output signal while preserving the target speech signal's distortion, thereby suppressing noise and interference. The core of the MVDR beamformer is the calculation of weight coefficients, which determine the beamformer's gain and attenuation for signals from different directions.

[0066] The purpose of calculating the MVDR weight coefficients based on the spatial covariance matrix of the initial speech signal and the array manifold vector is to utilize the spatial distribution characteristics of the speech signal and noise to adaptively adjust the weights of the beamformer so that it enhances the speech signal in the target direction and suppresses the noise in the non-target direction, thereby achieving spatial filtering and speech enhancement.

[0067] The specific implementation process is as follows: First, the spatial covariance matrix is calculated based on the initial speech signal. The speech signal received by the microphone array is divided into several short time frames, and the sample covariance matrix is calculated for each frame. Then, the sample covariance matrices of all frames are averaged to obtain the spatial covariance matrix R_xx. The calculation formula of the spatial covariance matrix is:

[0068] R_xx = E[x(t)x^H(t)]

[0069] Where x(t) is the speech signal vector received by the microphone array, x^H(t) is the conjugate transpose of x(t), and E[] represents the mathematical expectation.

[0070] Then, the MVDR weight coefficient is calculated using the spatial covariance matrix R_xx and the array manifold vector d(θ). The calculation formula of the MVDR weight coefficient is:

[0071] w_mvdr = (R_xx^(-1)d(θ)) / (d^H(θ)R_xx^(-1)d(θ))

[0072] Where R_xx^(-1) is the inverse of the spatial covariance matrix, d(θ) is the array manifold vector in the target direction, and d^H(θ) is the conjugate transpose of d(θ).

[0073] The MVDR weight coefficient w_mvdr is a complex vector whose dimension is equal to the number of microphones. It represents the weighting coefficient of the beamformer on the signals of different microphones so that the output signal of the beamformer obtains the maximum gain in the target direction and the minimum gain in other directions.

[0074] Finally, the calculated MVDR weight coefficients are applied to the beamformer to perform spatial filtering and enhancement on the initial speech signal. The beamformed speech signal y(t) can be expressed as:

[0075] y(t) = w_mvdr^H * x(t)

[0076] Where w_mvdr^H is the conjugate transpose of the MVDR weight coefficient, and x(t) is the speech signal vector received by the microphone array.

[0077] Adaptive spatial filtering and speech enhancement are achieved by calculating the MVDR beamformer's weight coefficients based on the spatial covariance matrix of the initial speech signal and the array manifold vectors. Compared to fixed-weight beamformers, the MVDR beamformer dynamically adjusts the weight coefficients based on the spatial distribution characteristics of the speech signal and noise, resulting in better noise suppression and speech enhancement.

[0078] S203: Perform weighted summation on the initial speech signal using the MVDR weight coefficient to form a spatial filter pointing in the direction of the target speech, and perform MVDR beamforming processing on the initial speech signal based on the spatial filter to obtain the initial speech signal after beamforming.

[0079] In a preferred embodiment of the present invention, after calculating the MVDR weight coefficient, the next step of the method of the present invention is to use the MVDR weight coefficient to weighted sum the initial speech signal to form a spatial filter pointing in the direction of the target speech, and perform MVDR beamforming processing on the initial speech signal based on the spatial filter to obtain the initial speech signal after beamforming.

[0080] Spatial filtering is a signal processing technology based on microphone arrays. It enhances speech signals from target directions and suppresses noise from non-target directions by weighted combination of speech signals received by different microphones. The core of spatial filtering is the weight coefficient, which determines the contribution of different microphone signals to the combination.

[0081] MVDR beamforming is a commonly used spatial filtering technique. It uses the MVDR criterion to calculate the optimal weight coefficients so that the beamformed speech signal obtains maximum gain in the target direction and minimum gain in other directions, thereby achieving spatial filtering directed towards the target speech.

[0082] The purpose of weighted summation of the initial speech signal using the MVDR weight coefficients is to form an optimal spatial filter based on the MVDR criterion, enhancing the speech signal in the target direction while suppressing noise in non-target directions. Through weighted summation, the speech signals received by different microphones are assigned different weights, resulting in a higher signal-to-noise ratio and greater directionality for the beamformed speech signal.

[0083] The specific implementation process is as follows: First, the MVDR weight coefficient w_mvdr is weighted and summed with the speech signal x(t) received by the microphone array to obtain the beamformed speech signal y(t). The weighted summation process can be expressed as:

[0084] y(t) = w_mvdr^H * x(t)

[0085] Where w_mvdr^H is the conjugate transpose of the MVDR weight coefficient, x(t) is the speech signal vector received by the microphone array, and y(t) is the speech signal after beamforming.

[0086] The weighted summation process can be viewed as a spatial filter whose frequency response is determined by the MVDR weight coefficients. By selecting appropriate MVDR weight coefficients, the spatial filter can have high gain in the direction of the target speech and low gain in non-target directions, thereby achieving spatial filtering directed towards the target speech.

[0087] Then, the beamformed speech signal y(t) is used as the output of the MVDR beamforming process to obtain the beamformed initial speech signal.

[0088] The essence of MVDR beamforming is to optimize the initial speech signal using a spatial filter, minimizing the variance of the output signal while preserving the target direction's speech signal distortion, thereby suppressing noise and interference. By weighted summation of MVDR weight coefficients, adaptive spatial filtering is achieved, resulting in higher quality and clarity of the beamformed speech signal.

[0089] Finally, the initial speech signal after beamforming is used as the input of the subsequent speech enhancement algorithm for further processing and optimization.

[0090] S102, extracting features from the initial speech signal to obtain BFCC features and fundamental frequency features;

[0091] In a preferred embodiment of the present invention, after the MVDR beamforming process is performed on the initial speech signal, the next step of the method of the present invention is to perform feature extraction on the beamformed initial speech signal to obtain BFCC features and fundamental frequency features.

[0092] BFCC features, or Bark Frequency Cepstral Coefficients, are a method for extracting frequency-domain features of speech signals on the Bark frequency scale. Similar to traditional MFCCs (Mel-Frequency Cepstral Coefficients), BFCC features are derived through a series of transformations and compressions on the speech signal's spectrum. However, BFCC utilizes the Bark frequency scale, which better aligns with human hearing and is closer to human hearing in terms of frequency resolution and perceived pitch. The steps for extracting BFCC features are as follows: First, a short-time Fourier transform is performed on the initial beamformed speech signal to obtain a spectrum. Then, the spectrum is mapped to the Bark frequency scale to obtain a Bark spectrum. Next, the logarithm of the Bark spectrum is taken to obtain a logarithmic Bark spectrum. Finally, a discrete cosine transform is performed on the logarithmic Bark spectrum to obtain the BFCC feature coefficients. Extracting BFCC features fully utilizes the characteristics of speech signals on the Bark frequency scale, providing an effective frequency-domain feature representation for subsequent speech / noise separation and speech enhancement.

[0093] The fundamental frequency feature, also known as the pitch feature, reflects the fundamental periodicity of the speech signal. For voiced speech, the fundamental frequency corresponds to the frequency of vocal cord vibration and is the lowest harmonic frequency in the speech signal. Extracting the fundamental frequency feature provides critical time-domain feature information for subsequent adaptive filtering and speech enhancement. Common methods for estimating fundamental frequency include the autocorrelation function method, the average amplitude difference function method, and the cepstrum method. Here, the autocorrelation function method is used for fundamental frequency estimation. The specific steps are as follows: First, the initial speech signal after beamforming is framed; then, the autocorrelation function is calculated for each frame of speech signal; then, the delay time corresponding to the first maximum point in the autocorrelation function is found; this time is the fundamental frequency period of the speech frame; finally, the inverse of the fundamental frequency period is used as the estimated fundamental frequency value for the speech frame. By extracting the fundamental frequency feature, the periodic structure of the speech signal can be accurately characterized, providing important parameters for subsequent adaptive filter design.

[0094] Based on the above embodiment, as an optional embodiment, feature extraction is performed on the initial speech signal to obtain BFCC features, including:

[0095] Perform short-time Fourier transform on the initial speech signal to obtain the spectrum;

[0096] Map the spectrum to the Bark frequency scale to obtain the Bark spectrum;

[0097] Take the logarithm of the Bark spectrum to obtain the logarithmic Bark spectrum;

[0098] The logarithmic Bark spectrum is subjected to discrete cosine transform to obtain the BFCC feature.

[0099] In a preferred embodiment of the present invention, a short-time Fourier transform (STFT) is performed on the initial speech signal to obtain a spectrum. The STFT is a time-frequency analysis method that divides the speech signal into several short-time frames and performs a Fourier transform on each frame to obtain the spectrum of that frame. The STFT can be used to obtain the frequency components of the speech signal at different time points, revealing the time-varying characteristics of the speech signal.

[0100] The spectrum represents the energy distribution of a speech signal at different frequencies, but it does not directly reflect the human auditory perception characteristics. To better characterize the perceptual characteristics of speech signals, the spectrum needs to be mapped to a frequency scale that matches the human auditory characteristics.

[0101] Therefore, the next step is to map the spectrum to the Bark frequency scale to obtain the Bark spectrum. The Bark frequency scale is a nonlinear frequency scale based on the critical bandwidth of the human ear. It divides the frequency range into several critical bands, and the frequency components within each critical band have similar contributions to human perception. By mapping the spectrum to the Bark frequency scale, the perceptual characteristics of the speech signal can be more accurately characterized, highlighting the frequency components to which the human ear is sensitive.

[0102] The process of mapping to the Bark frequency scale can be achieved by the Bark frequency conversion formula, which converts the linear frequency into the Bark frequency. The Bark frequency conversion formula is:

[0103] B(f) = 13arctan(0.00076f) + 3.5arctan((f / 7500)^2)

[0104] Where f is the linear frequency and B(f) is the corresponding Bark frequency.

[0105] After obtaining the Bark spectrum, we take its logarithm to further highlight the dynamic changes in energy across frequency bands. This yields a logarithmic Bark spectrum. This operation compresses the spectrum's dynamic range, smoothing energy differences between frequency bands and aligning with the human ear's perception of loudness. Furthermore, this operation converts spectrum multiplication into addition, simplifying subsequent processing.

[0106] Finally, the logarithmic Bark spectrum is subjected to a discrete cosine transform (DCT) to obtain the BFCC features. The DCT is an orthogonal transform that compresses spectral information into a few low-order coefficients while maintaining good decorrelation and energy concentration. The DCT yields compact, independent BFCC feature coefficients that effectively characterize characteristics such as timbre and pitch of the speech signal.

[0107] The calculation formula of BFCC characteristic is:

[0108] BFCC(i) = sqrt(2 / N) * sum(log(Bark(j)) * cos(pi i (j-0.5) / N)), i=1,2,...,M

[0109] Where N is the number of Bark bands, M is the order of the BFCC feature, and Bark(j) is the energy value of the j-th Bark band.

[0110] S103, using a pre-trained speech / noise separation DNN to process the BFCC features to obtain a time-frequency mask;

[0111] In a preferred embodiment of the present invention, after obtaining the BFCC features of the initial speech signal, the next step of the method of the present invention is to process the BFCC features using a pre-trained speech / noise separation deep neural network (DNN) to obtain a time-frequency mask.

[0112] The speech / noise separation DNN is a speech enhancement model based on deep learning, whose purpose is to separate clear target speech and background noise from noise-contaminated speech signals. Through the powerful nonlinear mapping and representation learning capabilities of DNN, effective speech and noise differentiation information can be extracted from the characteristics of noisy speech, thereby achieving speech and noise separation. In the method of the present invention, a large number of pure speech and various noise samples are pre-trained on the speech / noise separation DNN, so that it learns to estimate the time-frequency regions dominated by speech and noise from the BFCC features of noisy speech, and obtain a time-frequency mask (also called an ideal binary mask).

[0113] Specifically, the time-frequency mask is a binary matrix of the same size as the speech signal's time-frequency spectrum. Elements with a value of 1 indicate that the time-frequency point at that point is dominated by speech, while elements with a value of 0 indicate that the time-frequency point is dominated by noise. The time-frequency mask provides an ideal criterion for speech / noise separation. Simply multiplying the time-frequency mask by the speech signal's time-frequency spectrum yields an ideal, clean speech time-frequency spectrum.

[0114] During this step, the extracted BFCC features are first input into a pre-trained speech / noise separation DNN. The DNN uses multiple layers of nonlinear transformations to map the BFCC features into a high-dimensional space, extracting the distinguishing features between speech and noise. Next, the DNN's output layer uses a sigmoid activation function to convert these distinguishing features into probability values between 0 and 1, indicating the likelihood that each time-frequency point belongs to speech or noise. Finally, the probability values in the output layer are binarized according to a certain threshold to obtain the final time-frequency mask.

[0115] By estimating the time-frequency mask using a pre-trained speech / noise separation DNN, the distribution of speech and noise can be precisely characterized in the time-frequency domain, providing important prior information for subsequent speech enhancement. The time-frequency mask not only indicates the time-frequency regions dominated by speech and noise, but also serves as a weighting factor to guide signal synthesis and filtering operations during the speech enhancement process, ensuring that the enhanced speech retains the greatest possible detail and naturalness while reducing noise.

[0116] Based on the above embodiment, as an optional embodiment, a pre-trained speech / noise separation DNN is used to process the BFCC features to obtain a time-frequency mask, including:

[0117] The BFCC features are input into the pre-trained speech / noise separation DNN, and the local features of the BFCC features are extracted through the convolutional neural network of the pre-trained speech / noise separation DNN;

[0118] The recursive neural network of the pre-trained speech / noise separation DNN is used to model the temporal context information of the local features of the BFCC features to obtain the temporal context information;

[0119] The local features of the BFCC feature and the temporal context information are fused through the fully connected layer of the pre-trained speech / noise separation DNN to output a soft mask;

[0120] The soft mask is binarized to obtain the time-frequency mask.

[0121] In a preferred embodiment of the present invention, after obtaining the BFCC features, the next step of the method is to process the BFCC features using a pretrained speech / noise separation DNN to obtain a time-frequency mask. Speech / noise separation is a core task in speech enhancement, aiming to isolate clear speech components from noise-contaminated speech signals. Traditional speech / noise separation methods are typically based on statistical models or signal processing techniques. However, in recent years, speech / noise separation methods based on deep neural networks (DNNs) have made significant progress.

[0122] The pre-trained speech / noise separation DNN is a deep learning-based speech enhancement model. It is trained on a large amount of speech and noise data to learn the distinguishing features of speech and noise, thereby effectively separating speech and noise. Compared with traditional methods, the pre-trained speech / noise separation DNN can automatically learn and extract high-level features of speech signals, and has stronger generalization and robustness.

[0123] The purpose of using a pretrained speech / noise separation DNN to process BFCC features is to leverage the DNN's powerful feature learning and mapping capabilities to extract information distinguishing speech and noise from the BFCC features, generating a time-frequency mask to guide the subsequent speech enhancement process. The time-frequency mask is a binary matrix of the same size as the speech signal's time-frequency representation, where 1 represents time-frequency points dominated by speech and 0 represents time-frequency points dominated by noise. This time-frequency mask allows for the selective retention and suppression of speech and noise in the time-frequency domain, thereby achieving speech enhancement.

[0124] The specific implementation process is as follows: First, the BFCC features are input into a pre-trained speech / noise separation DNN. The DNN's convolutional neural network (CNN) then extracts local features of the BFCC features. A convolutional neural network is a specialized neural network architecture that effectively extracts local features of the input data through local connections and weight sharing. In the speech / noise separation task, the convolutional neural network can capture the local patterns and correlations of the BFCC features in the time-frequency domain, extracting information that distinguishes speech from noise.

[0125] Specifically, the convolutional neural network performs a convolution operation on the BFCC features using sliding convolution kernels to generate a series of feature maps. Each convolution kernel can be regarded as a local feature extractor, which scans different regions of the BFCC features to extract local time-frequency patterns. The mathematical expression of the convolution operation is:

[0126] F(i,j) = sum(W(m,n) * BFCC(im,jn)) + b

[0127] Among them, F(i,j) is the feature map after convolution, W(m,n) is the weight coefficient of the convolution kernel, BFCC(im,jn) is the local area of the BFCC feature, and b is the bias term.

[0128] Through layer-by-layer extraction of multi-layer convolutional neural networks, local feature representations of BFCC features at different scales and abstraction levels can be obtained.

[0129] Then, a recurrent neural network (RNN) from a pre-trained speech / noise separation DNN is used to model the temporal contextual information of the local features of the BFCC features. Recurrent neural networks are a type of neural network structure suitable for processing sequential data. Through recurrent connections and the transfer of internal states, they can capture data dependencies and contextual information in the temporal dimension. In the speech / noise separation task, recurrent neural networks can model the continuity and contextual relevance of BFCC features in the temporal dimension, extracting temporal information that distinguishes speech from noise.

[0130] Specifically, the recurrent neural network gradually models the temporal context information of the BFCC feature by iteratively updating the internal state and output. The mathematical expression of the recurrent neural network is:

[0131] h(t) = f(W(h)*h(t-1) + W(x)*x(t) + b)

[0132] y(t) = g(W(y)*h(t) + c)

[0133] Where h(t) is the internal state of the recurrent neural network at time t, x(t) is the input feature at time t, y(t) is the output at time t, W(h), W(x), and W(y) are weight matrices, b and c are bias terms, and f and g are activation functions.

[0134] Through iterative calculation of recursive neural networks, the contextual information of BFCC features in the time dimension can be obtained, capturing the time domain evolution characteristics of speech and noise.

[0135] Next, the fully connected layer (FC) of the pre-trained speech / noise separation DNN fuses the local features of the BFCC features with the temporal context information to output a soft mask. A fully connected layer is a common neural network structure that maps input features to the output space through linear transformations and nonlinear activations. In the speech / noise separation task, the fully connected layer can integrate the local features extracted by the convolutional neural network and the temporal context information extracted by the recurrent neural network to generate a soft mask that represents the discriminability between speech and noise.

[0136] Specifically, the fully connected layer fuses local features and temporal context information through matrix multiplication and nonlinear transformation to generate a soft mask of the same size as the BFCC feature time-frequency representation. The mathematical expression of the fully connected layer is:

[0137] M(i,j) = sigmoid(W(fc)*[F(i,j);H(i)] + b)

[0138] Among them, M(i,j) is the time-frequency point value of the soft mask, W(fc) is the weight matrix of the fully connected layer, [F(i,j);H(i)] is the concatenation of local features and temporal context information, b is the bias term, and sigmoid is the activation function.

[0139] The soft mask value range is [0, 1], indicating the probability of each time-frequency bin belonging to speech or noise. A value close to 1 indicates that the time-frequency bin primarily contains speech components, while a value close to 0 indicates that the time-frequency bin primarily contains noise components.

[0140] Finally, the soft mask is binarized to obtain a time-frequency mask. Binarization converts the probability values in the soft mask to 0 or 1 to obtain a binary time-frequency mask. This process can be achieved by setting a threshold. Time-frequency points greater than the threshold are set to 1, indicating speech dominance; time-frequency points less than or equal to the threshold are set to 0, indicating noise dominance. The mathematical expression for binarization is:

[0141] BM(i,j) = (M(i,j)>threshold) 1 : 0

[0142] Among them, BM(i,j) is the time-frequency mask after binarization, and threshold is the binarization threshold.

[0143] Through binarization, the soft mask is converted into a binary matrix of the same size as the BFCC feature time-frequency representation, where 1 represents speech-dominated time-frequency points and 0 represents noise-dominated time-frequency points. The time-frequency mask provides clear guidance for the subsequent speech enhancement process, indicating which time-frequency points should retain speech components and which should suppress noise components.

[0144] S104, adaptively designing parameters of a comb filter according to the fundamental frequency characteristics to obtain a target comb filter, and filtering the initial speech signal based on the comb filter to obtain a comb-filtered speech signal;

[0145] In a preferred embodiment of the present invention, a comb filter is a linear filter that enhances the periodic components of a speech signal and suppresses non-periodic noise. Its frequency response exhibits a series of evenly spaced peaks, resembling the shape of a comb, hence its name. A key parameter of a comb filter is the spacing of its peak frequencies, which should match the fundamental frequency of the speech signal to achieve selective enhancement of the fundamental frequency and its harmonics.

[0146] During this step, the parameters of the comb filter are first adaptively designed based on the fundamental frequency characteristics extracted in the previous step. Specifically, the reciprocal of the fundamental frequency is used as the delay interval parameter of the comb filter. By adjusting the filter's feedback coefficient and feedforward coefficient, the peak frequency of its frequency response is aligned with the fundamental frequency and its harmonics. Because the fundamental frequency may vary between different time frames, it is necessary to dynamically adjust the filter parameters based on the estimated fundamental frequency of each frame to achieve adaptive filtering.

[0147] After obtaining the target comb filter, it is applied to the initial speech signal. Specifically, the initial speech signal is passed through the designed comb filter and subjected to a time-domain convolution operation. During the filtering process, frequency components corresponding to the fundamental frequency and its harmonics are selectively enhanced, while other frequency components are relatively suppressed. In this way, the comb filter effectively extracts periodic components from the speech signal and reduces the influence of non-periodic noise, making the filtered speech signal clearer and more stable.

[0148] The comb-filtered speech signal not only retains the basic structure and prosodic characteristics of speech but also achieves a higher harmonic-to-noise ratio (HNR), meaning that the energy of the speech's harmonic components is more prominent than that of the noise components. This periodic enhancement helps improve speech clarity and naturalness, laying a good foundation for subsequent speech synthesis and enhancement processing.

[0149] It should be noted that the design and implementation of the comb filter is a key technical step. To achieve efficient, real-time adaptive filtering, a comb filter based on an IIR (Infinite Impulse Response) structure is employed. This structure offers low computational complexity and excellent frequency selectivity. Furthermore, a smooth parameter adjustment mechanism is designed to avoid drastic changes in the filter response when sudden changes occur in the pitch frequency estimate, thereby improving the stability and robustness of the filtering process.

[0150] Based on the above embodiment, as an optional embodiment, adaptively designing the parameters of the comb filter according to the fundamental frequency characteristics to obtain the target comb filter includes:

[0151] The order of the comb filter is determined according to the fundamental frequency characteristics, and the order is inversely proportional to the fundamental frequency;

[0152] Determine the passband center frequency of the comb filter according to the fundamental frequency characteristics;

[0153] Determine the bandwidth of the comb filter according to the fundamental frequency characteristics;

[0154] According to the order, passband center frequency and bandwidth, the filter coefficients of the comb filter are determined to obtain the target comb filter.

[0155] In a preferred embodiment of the present invention, after obtaining the fundamental frequency characteristics, the next step of the method of the present invention is to adaptively design the parameters of the comb filter based on the fundamental frequency characteristics to obtain a target comb filter. The comb filter is a commonly used speech enhancement and noise suppression filter that achieves speech enhancement and purification by selectively filtering speech and noise in the frequency domain. Traditional comb filters generally use fixed parameter settings, such as order, passband center frequency, and bandwidth. However, the method of the present invention proposes an adaptive comb filter design method that dynamically adjusts the filter parameters based on the fundamental frequency characteristics of the speech signal to better match the frequency structure of the speech and improve the filtering effect.

[0156] Pitch frequency is a key characteristic of speech signals, reflecting the vibration characteristics of the speaker's vocal tract and the pitch of the speech. Pitch frequency varies between speakers and for different speech contents. Therefore, designing a comb filter based on this frequency characteristic allows the filter to better adapt to the frequency structure of different speech sounds, achieving more precise speech enhancement.

[0157] The specific implementation process is as follows: First, the order of the comb filter is determined based on the fundamental frequency characteristics. The order of the comb filter determines the filter's periodicity and selectivity in the frequency domain. A higher order results in denser peaks in the filter's frequency response curve, and a stronger ability to distinguish between speech and noise. However, an excessively high order also increases computational complexity and reduces filtering effectiveness. Therefore, the order of the comb filter needs to be appropriately selected based on the fundamental frequency characteristics.

[0158] The method of the present invention proposes that the order of the comb filter is inversely proportional to the fundamental frequency. The higher the fundamental frequency, the higher the pitch of the speech and the denser the frequency structure, in which case a lower order should be selected; conversely, the lower the fundamental frequency, the lower the pitch of the speech and the sparser the frequency structure, in which case a higher order should be selected. The calculation formula for the comb filter order is:

[0159] N = round(fs / (k * F0))

[0160] Among them, N is the order of the comb filter, fs is the sampling frequency of the speech signal, F0 is the fundamental frequency, and k is a proportional coefficient used to adjust the size of the order.

[0161] Through the above formula, the order of the comb filter can be adaptively determined according to the fundamental frequency characteristics of the speech signal, so that it matches the frequency structure of the speech and improves the accuracy of filtering.

[0162] The comb filter's passband center frequency is then determined based on the fundamental frequency characteristics. The passband center frequency is the center position of the comb filter in the frequency domain and determines the region where the filter enhances speech frequency components. To effectively enhance speech, the comb filter's passband center frequency should align with the primary frequency components of speech, which are typically concentrated near the fundamental frequency and its multiples.

[0163] The method of the present invention proposes to use the frequency multiplication of the fundamental frequency as the passband center frequency of the comb filter. The frequency multiplication of the fundamental frequency can be expressed as:

[0164] Fc(m) = m * F0

[0165] Where Fc(m) is the center frequency of the mth passband, F0 is the fundamental frequency, and m is the frequency multiplication coefficient, which is a positive integer.

[0166] By using the multiple of the fundamental frequency as the passband center frequency, the comb filter can achieve gain on the main frequency components of the speech and effectively enhance the speech signal.

[0167] Next, the comb filter bandwidth is determined based on the fundamental frequency characteristics. Bandwidth is the width of the comb filter in the frequency domain, which determines the filter's frequency range for speech and noise. Narrower bandwidths improve the filter's ability to distinguish between speech and noise, but may also cause speech distortion. Wider bandwidths weaken the filter's ability to distinguish between speech and noise, but also preserve more speech details. Therefore, the comb filter bandwidth must be appropriately selected based on the fundamental frequency characteristics.

[0168] The method of the present invention proposes to use a certain proportion of the fundamental frequency as the bandwidth of the comb filter. The calculation formula of the bandwidth is:

[0169] BW = α * F0

[0170] Among them, BW is the bandwidth of the comb filter, F0 is the fundamental frequency, and α is a proportional coefficient used to adjust the size of the bandwidth.

[0171] By using a certain proportion of the fundamental frequency as the bandwidth, the comb filter can adaptively adjust the frequency selection range according to the pitch characteristics of the speech, achieving a balance between speech enhancement and detail preservation.

[0172] Finally, the filter coefficients of the comb filter are determined based on the order, passband center frequency, and bandwidth to obtain the target comb filter. The filter coefficients of the comb filter determine the specific response characteristics of the filter in the time domain and frequency domain, and are the key parameters for implementing comb filtering.

[0173] The method of the present invention proposes to design the filter coefficients of the comb filter using the window function method. First, a comb function in the frequency domain is generated according to the order of the comb filter and the passband center frequency. Its mathematical expression is:

[0174] H(k) = 1, if |k - mN| ≤ BW / 2= 0, otherwise

[0175] Where H(k) is the value of the comb function in the frequency domain, k is the index of the frequency domain sample point, m is the index of the passband, N is the order of the comb filter, and BW is the bandwidth of the passband.

[0176] A window function is then applied to the comb function to smooth the filter's frequency response and reduce ringing artifacts. Common window functions include the Hamming window, the Hanning window, and the Blackman window. The choice of window function depends on the specific application scenario and performance requirements.

[0177] Finally, the inverse Fourier transform is performed on the windowed comb function to obtain the filter coefficients of the comb filter in the time domain. The mathematical expression of the filter coefficients is:

[0178] h(n) = IDFT[H(k) · W(k)]

[0179] Wherein, h(n) is the filter coefficient of the comb filter, H(k) is the windowed comb function, W(k) is the window function, and IDFT stands for inverse discrete Fourier transform.

[0180] By designing the filter coefficients of the comb filter using the window function method, a parameter-adaptive target comb filter can be obtained. It can dynamically adjust the filter order, passband center frequency and bandwidth according to the fundamental frequency characteristics of the speech signal, thereby achieving more accurate and effective speech enhancement.

[0181] S105, multiplying the time-frequency mask as a weighting coefficient with the comb-filtered speech signal to obtain the energy of the weighted signal, and using the energy of the weighted signal as the target speech energy estimate;

[0182] In a preferred embodiment of the present invention, after the initial speech signal is subjected to adaptive comb filtering, the next step of the method of the present invention is to multiply the time-frequency mask as a weighting coefficient with the comb-filtered speech signal to obtain the energy of the weighted signal, and use the energy of the weighted signal as the target speech energy estimate.

[0183] The time-frequency mask, estimated by the speech / noise separation DNN in the previous step, represents the dominant region of the speech signal in the time-frequency domain. Using the time-frequency mask as a weighting factor selectively enhances the signal energy in speech-dominated regions while suppressing the energy in noise-dominated regions, thereby achieving accurate estimation of speech signal energy.

[0184] During this step, the comb-filtered speech signal is first subjected to a short-time Fourier transform (STFT) to obtain its time-frequency representation. The time-frequency mask is then multiplied with the speech signal's time-frequency spectrum to obtain a weighted time-frequency spectrum. During this dot-wise multiplication, the time-frequency points corresponding to elements with a value of 1 in the time-frequency mask remain unchanged, while the time-frequency points corresponding to elements with a value of 0 are suppressed to 0. This preserves the energy of speech-dominated time-frequency points in the weighted time-frequency spectrum, while effectively suppressing the energy of noise-dominated points. Next, the energy of the weighted time-frequency spectrum is summed to obtain the total energy of the weighted signal. Specifically, the energy of all frequency points within each frame of the weighted time-frequency spectrum is summed to obtain the energy value for that frame. The energy values of all frames are accumulated to obtain the total energy of the entire weighted signal. This total energy reflects the primary energy distribution of the speech signal in the time-frequency domain, eliminating most of the noise energy interference, and can serve as a relatively accurate estimate of the target speech energy.

[0185] By multiplying the comb-filtered speech signal with the time-frequency mask as a weighting factor and calculating the energy of the weighted signal, an estimate of the target speech energy is obtained. This estimate comprehensively considers the speech signal's time-domain periodicity and frequency-domain energy distribution, fully utilizing the fundamental frequency characteristics and time-frequency mask information extracted in the previous steps, resulting in high reliability and accuracy.

[0186] Based on the above embodiment, as an optional embodiment, the time-frequency mask is used as a weighting coefficient and multiplied with the speech signal after comb filtering to obtain the energy of the weighted signal, including:

[0187] The elements representing the target speech in the time-frequency mask are set to 1, and the elements representing the noise are set to 0 to obtain the speech time-frequency mask;

[0188] The speech time-frequency mask is used as a weighting coefficient and multiplied with the speech signal after comb filtering to obtain a weighted signal;

[0189] Calculate the average energy of the weighted signal in the time dimension and use the average energy as the target speech energy estimate.

[0190] In a preferred embodiment of the present invention, after obtaining the target comb filter, the next step of the method is to multiply the comb-filtered speech signal using the time-frequency mask as a weighting factor to obtain the energy of the weighted signal. The purpose of this step is to use the time-frequency mask to weight the comb-filtered speech signal, highlighting the target speech components, suppressing the noise components, and estimating the energy of the target speech.

[0191] The time-frequency mask is a matrix of the same size as the time-frequency representation of the speech signal. It represents the energy distribution of the speech signal on the time-frequency plane. Elements in the time-frequency mask typically range from [0, 1], where elements close to 1 indicate that the time-frequency point primarily contains the target speech component, and elements close to 0 indicate that the time-frequency point primarily contains noise. The time-frequency mask can be used to distinguish the target speech from noise in the speech signal, providing guidance for subsequent speech enhancement processing.

[0192] The specific implementation process is as follows: First, the elements representing the target speech in the time-frequency mask are set to 1, and the elements representing noise are set to 0, to obtain the speech time-frequency mask. This step binarizes the original time-frequency mask, converting it into a matrix containing only 0s and 1s. Elements with a value of 1 indicate that the corresponding time-frequency point belongs to the target speech, while elements with a value of 0 indicate that the corresponding time-frequency point belongs to noise. The speech time-frequency mask can be generated by setting an energy threshold. Time-frequency points above the threshold are considered speech-dominated, while those below the threshold are considered noise-dominated. The threshold should be adjusted based on the specific signal-to-noise ratio and speech characteristics to achieve optimal speech / noise separation. The speech time-frequency mask is then used as a weighting coefficient and multiplied with the comb-filtered speech signal to obtain a weighted signal. The comb-filtered speech signal is processed by the adaptive comb filter, which enhances the target speech and suppresses noise in the time-frequency domain. Multiplying the speech time-frequency mask with the comb-filtered speech signal is equivalent to performing a second weighting on the speech signal, further highlighting the target speech component and weakening the remaining noise component. The weighting process can be described by the following mathematical expression:

[0193] S_weighted(t,f) = S_comb(t,f) · M_speech(t,f)

[0194] Among them, S_weighted(t,f) represents the weighted speech signal, S_comb(t,f) represents the speech signal after comb filtering, M_speech(t,f) represents the speech time-frequency mask, and t and f represent the indexes of the time and frequency dimensions respectively.

[0195] Through weighted operations, a new time-frequency representation is obtained, in which the target speech component is amplified and the noise component is suppressed, and the clarity and purity of the speech are further improved.

[0196] Finally, the average energy of the weighted signal over time is calculated and used as the target speech energy estimate. Energy is an important indicator of signal strength, reflecting the amplitude and variability of the signal. By calculating the average energy of the weighted signal over time, we can estimate the energy level of the target speech in the entire speech segment, providing a reference for subsequent signal-to-noise ratio estimation and speech quality assessment. The formula for calculating average energy is:

[0197] E_speech = (1 / T) * sum(S_weighted(t,f)^2, t=1,2,...,T)

[0198] Among them, E_speech represents the average energy of the target speech, T represents the total duration of the speech segment, and sum represents the summation of the time dimension.

[0199] Using the average energy of the weighted signal as the target speech energy estimate can reflect the proportion and intensity of the target speech component in the speech signal, provide an energy reference for the speech enhancement algorithm, and assist in adjusting the gain coefficient and suppression factor to achieve better speech enhancement effects.

[0200] S106, multiplying the time-frequency mask as a weighting coefficient with the initial speech signal to obtain a noise energy estimation value;

[0201] In a preferred embodiment of the present invention, after obtaining the target speech energy estimate, the next step of the method of the present invention is to multiply the time-frequency mask as a weighting coefficient with the initial speech signal to obtain a noise energy estimate.

[0202] The purpose of this step is to estimate the energy of the background noise in the speech signal, providing the necessary information for the subsequent calculation of the signal-to-noise ratio and speech gain. Unlike the previous step, which used the time-frequency mask to estimate the target speech energy, this step uses the complement of the time-frequency mask (one minus the time-frequency mask) to weight the initial speech signal to obtain an estimate of the noise energy.

[0203] The specific implementation process is as follows: First, the initial speech signal is subjected to a short-time Fourier transform (STFT) to obtain its time-frequency representation. Then, the complement of the time-frequency mask is calculated, that is, the time-frequency mask is subtracted from the all-one matrix to obtain the indicator matrix of the noise-dominated area. Next, the indicator matrix of the noise-dominated area is point-multiplied with the time-frequency spectrum of the initial speech signal to obtain the weighted noise time-frequency spectrum. During the point-multiplication process, the time-frequency points corresponding to the elements with a value of 1 in the noise-dominated area indicator matrix remain unchanged, while the time-frequency points corresponding to the elements with a value of 0 are suppressed to 0. In this way, in the weighted noise time-frequency spectrum, the energy of the noise-dominated time-frequency points is retained, while the energy of the speech-dominated time-frequency points is effectively suppressed.

[0204] Finally, the energy of the weighted noise time-spectrum is summed to obtain the total energy estimate of the noise signal. Specifically, the energy of all frequency points within each frame of the weighted noise time-spectrum is added to obtain the noise energy estimate for that frame. The noise energy estimates of all frames are accumulated to obtain the total energy estimate of the entire noise signal.

[0205] In this way, using the complement of the time-frequency mask as a weighting factor effectively extracts the energy of background noise from the initial speech signal, suppressing interference in speech-dominated areas. Because the time-frequency mask, estimated by the speech / noise separation DNN in the previous step, accurately distinguishes speech from noise, weighting based on the complement of the time-frequency mask yields a reliable estimate of noise energy.

[0206] Obtaining a noise energy estimate is crucial for speech enhancement algorithms. With this estimate, we can further calculate the energy ratio of the speech signal to the noise signal, known as the signal-to-noise ratio (SNR). This SNR is a key metric for measuring speech quality and noise contamination. Its calculation formula is: SNR = 10 * log10(speech energy / noise energy). In subsequent steps, the target speech energy estimate and the noise energy estimate are used to calculate the SNR. This is used as a basis for designing a speech gain function to adaptively enhance the speech signal, ultimately achieving high-quality noise suppression and speech extraction.

[0207] S107, calculating a frequency-dependent Bark gain coefficient on a Bark frequency scale according to the target speech energy estimate and the noise energy estimate;

[0208] In a preferred embodiment of the present invention, after obtaining the target speech energy estimate and the noise energy estimate, the next step of the method of the present invention is to calculate the frequency-dependent Bark gain coefficient on the Bark frequency scale based on the two energy estimates.

[0209] The Bark frequency scale is a frequency scale based on the human hearing characteristics. It divides frequencies into several critical bands, each of which widens as the frequency increases, more closely matching the frequency resolution characteristics of the human ear. In the field of speech enhancement, using the Bark frequency scale for gain calculation can better match the human auditory perception, achieving a more natural and comfortable enhancement effect.

[0210] The frequency-dependent Bark gain coefficient is a frequency-dependent gain function defined on the Bark frequency scale. Its purpose is to adaptively adjust the gain coefficient of each frequency band based on the energy distribution of speech and noise signals in different Bark frequency bands, thereby achieving frequency-selective speech enhancement.

[0211] During the implementation of this step, the energy of the speech signal and noise signal in each Bark frequency band is first calculated based on the target speech energy estimate and noise energy estimate. Specifically, the time-frequency spectrum energy of the speech signal and noise signal is accumulated according to the division of the Bark frequency band to obtain the speech energy and noise energy in each Bark frequency band. Then, based on the speech energy and noise energy in each Bark frequency band, the signal-to-noise ratio (SNR) of the frequency band is calculated. The calculation formula for the Bark frequency band signal-to-noise ratio is: Bark_SNR(i) = 10 * log10(Bark_Speech_Energy(i) / Bark_Noise_Energy(i)), where i represents the i-th Bark frequency band.

[0212] Next, the Bark band signal-to-noise ratio is used to calculate the frequency-dependent Bark gain coefficient. Common gain calculation formulas include Wiener filter, MMSE estimator, log spectrum amplitude estimator, etc. Here, an improved Wiener filter formula is adopted, introducing additional gain adjustment factors and noise suppression factors to balance the trade-off between speech distortion and noise suppression. The specific Bark gain coefficient calculation formula is as follows:

[0213] Bark_Gain(i) = max(min((Bark_SNR(i) / (Bark_SNR(i) + 1))^α * β, 1),0)

[0214] Among them, α is the gain adjustment factor, β is the noise suppression factor, and the max and min functions are used to ensure that the gain coefficient is between 0 and 1.

[0215] Finally, the calculated Bark gain coefficients are mapped back to the time-frequency domain to obtain a frequency-dependent, time-varying gain function. For each time-frequency point, the corresponding gain coefficient is equal to the gain coefficient of the Bark frequency band in which it resides. This results in a frequency-dependent Bark gain function defined in the time-frequency domain that can be used for subsequent speech enhancement processing.

[0216] Based on the above embodiment, as an optional embodiment, the frequency-dependent Bark gain coefficient is calculated on the Bark frequency scale according to the target speech energy estimate and the noise energy estimate, including:

[0217] Smoothing the noise energy estimation value and the target speech energy estimation value on the Bark frequency scale to obtain a smoothed target speech energy estimation value and a smoothed noise energy estimation value;

[0218] Calculate the ratio of the smoothed target speech energy estimate to the smoothed noise energy estimate to obtain the signal-to-noise ratio on the Bark frequency scale;

[0219] The signal-to-noise ratio is mapped to the Bark gain coefficient on the Bark frequency scale through a preset exponential function.

[0220] In a preferred embodiment of the present invention, after obtaining the target speech energy estimate and the noise energy estimate, the method of the present invention next calculates a frequency-dependent Bark gain coefficient on the Bark frequency scale based on these two energy estimates. The Bark gain coefficient is a frequency-dependent gain adjustment factor that reflects the relative magnitude of speech and noise energy in different frequency regions and is used to adaptively enhance speech signals in the frequency domain.

[0221] The Bark frequency scale is a nonlinear frequency scale based on the human hearing characteristics. It divides the frequency range into several critical bands, each of which has similar frequency components in its perception. Processing on the Bark frequency scale better accounts for the human hearing characteristics, achieving speech enhancement effects that align with human perception.

[0222] The specific implementation process is as follows: First, the noise energy estimate and the target speech energy estimate are smoothed on the Bark frequency scale to obtain the smoothed target speech energy estimate and the smoothed noise energy estimate. The purpose of smoothing is to reduce short-term fluctuations and estimation errors in the energy estimate, thereby improving the stability and reliability of the energy estimate. Smoothing can be achieved by applying a smoothing window function on the Bark frequency scale. Commonly used smoothing window functions include triangular windows, Hanning windows, and Gaussian windows. The choice of smoothing window function should be based on a trade-off between the specific frequency resolution and smoothness requirements.

[0223] The smoothing process can be described by the following mathematical expression:

[0224] E_speech_smooth(b) = sum(E_speech(f) * W(b,f), f=1,2,...,F)

[0225] E_noise_smooth(b) = sum(E_noise(f) * W(b,f), f=1,2,...,F)

[0226] Where E_speech_smooth(b) and E_noise_smooth(b) represent the smoothed target speech energy estimate and the smoothed noise energy estimate, respectively. E_speech(f) and E_noise(f) represent the original target speech energy estimate and the noise energy estimate, respectively. W(b,f) represents the smoothing window function, b and f represent the indices of the Bark frequency band and the linear frequency, respectively. F represents the total number of frequency points.

[0227] Through smoothing, a more stable and reliable energy estimate is obtained on the Bark frequency scale, laying the foundation for subsequent signal-to-noise ratio calculation and gain estimation. Then, the ratio of the smoothed target speech energy estimate to the smoothed noise energy estimate is calculated to obtain the signal-to-noise ratio on the Bark frequency scale. The signal-to-noise ratio is an important indicator for measuring the quality of speech signals. It reflects the relative strength of the speech component and the noise component in the speech signal. By calculating the signal-to-noise ratio on the Bark frequency scale, the relative strength of speech and noise in different frequency regions can be evaluated, providing a basis for adaptive gain control. The formula for calculating the signal-to-noise ratio on the Bark frequency scale is:

[0228] SNR_bark(b) = E_speech_smooth(b) / E_noise_smooth(b)

[0229] Wherein, SNR_bark(b) represents the signal-to-noise ratio of the b-th Bark frequency band.

[0230] The degree of dominance of speech and noise in different Bark frequency bands can be determined based on the size of the signal-to-noise ratio. When the signal-to-noise ratio is high, it indicates that the speech component is dominant and the gain of the frequency band should be amplified; when the signal-to-noise ratio is low, it indicates that the noise component is dominant and the gain of the frequency band should be suppressed.

[0231] Finally, the signal-to-noise ratio is mapped to the Bark gain coefficient on the Bark frequency scale using a preset exponential function. The exponential function is a commonly used nonlinear mapping function that can compress the variation range of the signal-to-noise ratio into a limited gain coefficient range, achieving smooth gain adjustment. The general form of the exponential function is:

[0232] G_bark(b) = a * exp(b * SNR_bark(b))

[0233] Wherein, G_bark(b) represents the Bark gain coefficient of the b-th Bark frequency band, a and b are parameters for controlling the dynamic range of the gain coefficient, and can be set according to specific gain adjustment requirements.

[0234] Through exponential function mapping, a set of frequency-dependent Bark gain coefficients is derived. These reflect the relative magnitude of speech and noise energy within different Bark frequency bands and are used to adaptively enhance speech signals in the frequency domain. Bark gain coefficients typically range from [0, 1]. Values close to 1 indicate that the speech component in the corresponding frequency band is dominant, requiring a significant gain amplification; values close to 0 indicate that the noise component in the corresponding frequency band is dominant, requiring a significant gain suppression.

[0235] S108, multiplying the comb-filtered speech signal by the Bark gain coefficient to enhance the initial speech signal in the time-frequency domain to obtain an enhanced speech signal;

[0236] In a preferred embodiment of the present invention, after calculating the frequency-dependent Bark gain coefficient, the next step of the method of the present invention is to multiply the comb-filtered speech signal by the Bark gain coefficient, enhance the initial speech signal in the time-frequency domain, and obtain an enhanced speech signal.

[0237] The purpose of this step is to use an adaptive frequency-selective gain function to perform time-varying gain adjustment on the comb-filtered speech signal, improving speech clarity and naturalness while suppressing background noise. By multiplying the comb-filtered speech signal with the Bark gain coefficient, refined speech enhancement processing can be achieved in the time-frequency domain.

[0238] The specific implementation process is as follows: First, the comb-filtered speech signal undergoes a short-time Fourier transform (STFT) to obtain its time-frequency representation. The comb-filtered speech signal here is obtained by processing the initial speech signal with the adaptive comb filter designed based on the fundamental frequency characteristics in the previous step. It contains the enhanced speech fundamental frequency and its harmonic components, as well as the suppressed noise components. Next, the Bark gain coefficients are mapped to the dimensions of the time-frequency spectrum, resulting in a gain matrix of the same size as the time-frequency spectrum. Specifically, for each time-frequency point in the time-frequency spectrum, the Bark frequency band to which it belongs is determined based on its frequency value, and the gain coefficient for the corresponding Bark frequency band is assigned to that time-frequency point, forming a gain matrix. Next, the time-frequency spectrum of the comb-filtered speech signal is dot-multiplied with the gain matrix to obtain the enhanced speech time-frequency spectrum. During this dot-multiplication, speech-dominated time-frequency points in the time-frequency spectrum are amplified by the Bark gain coefficients, while noise-dominated time-frequency points are suppressed by the Bark gain coefficients. This time-varying, frequency-selective gain adjustment flexibly balances speech enhancement and noise suppression at different time and frequency locations.

[0239] Finally, the enhanced speech time-spectrum is subjected to an inverse short-time Fourier transform (ISTFT) to obtain the enhanced speech signal. The ISTFT process involves performing phase recovery and overlap-add operations on the time-spectrum to reconstruct the speech waveform in the time domain.

[0240] By multiplying the comb-filtered speech signal by the Bark gain coefficient, adaptive speech enhancement is achieved in the time-frequency domain. Compared with traditional time-domain filtering methods, this time-frequency domain enhancement method can more flexibly and finely adjust the gain coefficient at different time and frequency locations, fully utilizing the time-varying and frequency-varying characteristics of the speech signal to achieve a more optimized enhancement effect.

[0241] S109 , using an inverse BFCC transform to convert the enhanced speech signal from the Bark frequency scale into a time domain signal, to obtain a target enhanced speech pointing in the target direction.

[0242] In a preferred embodiment of the present invention, after the initial speech signal is enhanced in the time-frequency domain, the last step of the method of the present invention is to use an inverse BFCC transform to convert the enhanced speech signal from the Bark frequency scale to a time domain signal to obtain a target enhanced speech pointing in the target direction.

[0243] BFCC (Bark Frequency Cepstral Coefficients) is a speech feature representation method similar to MFCC (Mel Frequency Cepstral Coefficients), but uses the Bark frequency scale instead of the Mel frequency scale during calculation. Because the Bark frequency scale better reflects human hearing, BFCC features are widely used in speech enhancement and speech recognition. In the previous step, the speech signal was enhanced at the Bark frequency scale, resulting in an enhanced speech signal in the Bark frequency domain. To convert the enhanced speech signal back to the time domain, an inverse BFCC transform is required.

[0244] The specific implementation process is as follows: First, logarithmic energy calculation is performed on the enhanced speech signal in the Bark frequency domain to obtain the logarithmic energy spectrum of the Bark frequency band. This step converts the energy values of the Bark frequency band to a logarithmic scale, making it more consistent with the human ear's loudness perception characteristics. Then, a discrete cosine transform (DCT) is performed on the logarithmic energy spectrum of the Bark frequency band to obtain the BFCC feature coefficients. The DCT transform compresses the band energy spectrum into a small number of cepstral coefficients while maintaining good decorrelation and energy concentration. The BFCC feature coefficients represent the envelope characteristics of the speech signal on the Bark frequency scale. Next, an inverse discrete cosine transform (IDCT) is performed on the BFCC feature coefficients to convert them back to the logarithmic energy spectrum of the Bark frequency band. The IDCT transform is the inverse process of the DCT transform and can restore the cepstral coefficients to the band energy spectrum. Then, an exponential operation is performed on the logarithmic energy spectrum of the Bark frequency band to convert the energy values on the logarithmic scale to a linear scale. This step restores the band energy spectrum to its original energy values for subsequent band synthesis.

[0245] Finally, a Bark band synthesis filter is used to convert the energy values of the Bark bands into a time-domain signal, resulting in the target enhanced speech signal. The Bark band synthesis filter is a set of bandpass filters corresponding to the Bark band divisions, with a frequency response that matches the critical band characteristics of human hearing. By multiplying the energy value of each Bark band by the corresponding synthesis filter and summing the filtered results across all bands, the enhanced speech signal in the time domain is obtained.

[0246] See also Figure 2 , Figure 2 This is an architecture diagram of a directional sound pickup system provided in an embodiment of the present application. The directional sound pickup system may include:

[0247] Data acquisition module 1, used to perform MVDR beamforming processing on the initial speech signal input by the microphone array to obtain the initial speech signal after beamforming;

[0248] Feature extraction module 2 is used to extract features from the initial speech signal to obtain BFCC features and fundamental frequency features;

[0249] Special processing module 3 is used to process the BFCC features using the pre-trained speech / noise separation DNN to obtain the time-frequency mask;

[0250] Filtering module 4, used to adaptively design the parameters of the comb filter according to the fundamental frequency characteristics to obtain a target comb filter, and filter the initial speech signal based on the comb filter to obtain a comb-filtered speech signal;

[0251] The target speech energy estimation module 5 is used to multiply the time-frequency mask as a weighting coefficient with the speech signal after comb filtering to obtain the energy of the weighted signal, and use the energy of the weighted signal as the target speech energy estimation value;

[0252] Noise energy estimation module 6, used to multiply the time-frequency mask as a weighting coefficient with the initial speech signal to obtain a noise energy estimation value;

[0253] A gain coefficient calculation module 7 is used to calculate the frequency-dependent Bark gain coefficient on the Bark frequency scale according to the target speech energy estimate and the noise energy estimate;

[0254] A first signal enhancement module 8 is configured to multiply the comb-filtered speech signal by a Bark gain coefficient to enhance the initial speech signal in the time-frequency domain to obtain an enhanced speech signal;

[0255] The second signal enhancement module 9 is configured to convert the enhanced speech signal from the Bark frequency scale into a time domain signal by using an inverse BFCC transform, so as to obtain a target enhanced speech pointing in a target direction.

[0256] It should be noted that the above embodiments provide systems that implement their functions using only the division of the above functional modules as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0257] Please refer to Figure 3 The present application also discloses an electronic device. Figure 3 The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302 or end-to-end wireless communication.

[0258] The communication bus 302 is used to implement the connection and communication between these components.

[0259] The user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.

[0260] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).

[0261] The processor 301 may include one or more processing cores. Using various interfaces and circuits, the processor 301 connects to various components within the server. It executes instructions, programs, code sets, or instruction sets stored in the memory 305, as well as accesses data stored in the memory 305, to perform various server functions and process data. Optionally, the processor 301 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 301 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 301 but implemented as a separate chip.

[0262] Among them, the memory 305 may include a random access memory (Random Access Memory, RAM) and may also include a read-only memory (Read~Only Memory). Optionally, the memory 305 includes a non-transitory computer-readable medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 305 may also optionally be at least one storage system located away from the aforementioned processor 301. Reference Figure 3 , the memory 305 as a computer storage medium may include an operating system, a network communication module, a user interface module and an application program of a directional sound pickup method.

[0263] exist Figure 3 In the electronic device 300 shown, the user interface 303 is mainly used to provide an input interface for the user and obtain the data input by the user; and the processor 301 can be used to call the application program storing the directional sound pickup method in the memory 305. When executed by one or more processors 301, the electronic device 300 executes one or more methods in the above-mentioned embodiments. It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that this application is not limited to the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application. In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0264] In the several embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative, such as the division of modules, which is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, the indirect coupling or communication connection of the system or module can be electrical or other forms.

[0265] Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of these modules may be selected to achieve the purpose of this embodiment based on actual needs.

[0266] The present application also provides a computer storage medium that can store multiple instructions, which are suitable for being loaded and executed by a processor as described above. Figure 1 The directional sound pickup method of the embodiment shown, the specific execution process can be found in Figure 1 The detailed description of the illustrated embodiment will not be repeated here.

[0267] In addition, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.

[0268] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of this application. The aforementioned memory includes various media that can store program code, such as USB flash drives, mobile hard drives, magnetic disks, or optical disks.

[0269] The above are merely exemplary embodiments of the present disclosure and are not intended to limit the scope of the present disclosure. In other words, any equivalent variations and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the disclosure and the practical implications thereof.

[0270] This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not described herein. The description and examples are to be considered as exemplary only, and the scope and spirit of the present disclosure are to be defined by the claims.

Claims

1. A directional sound pickup method, characterized in that: The method comprises: Performing MVDR beamforming processing on the initial speech signal input by the microphone array to obtain the initial speech signal after beamforming; Extracting features of the initial speech signal to obtain BFCC features and fundamental frequency features; Processing the BFCC features using a pre-trained speech / noise separation DNN to obtain a time-frequency mask; Adaptively designing parameters of a comb filter according to the fundamental frequency characteristics to obtain a target comb filter, and filtering the initial speech signal based on the comb filter to obtain a comb-filtered speech signal; Multiplying the time-frequency mask as a weighting coefficient with the comb-filtered speech signal to obtain energy of a weighted signal, and using the energy of the weighted signal as an estimated value of target speech energy; Multiplying the initial speech signal by the time-frequency mask as a weighting coefficient to obtain a noise energy estimation value; Calculating a frequency-dependent Bark gain coefficient on a Bark frequency scale according to the target speech energy estimate and the noise energy estimate; Multiplying the comb-filtered speech signal by a Bark gain coefficient to enhance the initial speech signal in the time-frequency domain to obtain an enhanced speech signal; Converting the enhanced speech signal from the Bark frequency scale to a time domain signal using an inverse BFCC transform to obtain a target enhanced speech directed in a target direction; The performing of MVDR beamforming processing on the initial speech signal input by the microphone array to obtain the initial speech signal after beamforming includes: Determining an array manifold vector according to geometric arrangement information of the microphone array and azimuth information of the initial speech signal; Calculating a weight coefficient of an MVDR beamformer according to the spatial covariance matrix of the initial speech signal and the array manifold vector to obtain an MVDR weight coefficient; The initial speech signal is weighted and summed using the MVDR weight coefficient to form a spatial filter pointing to the target speech direction, and the initial speech signal is subjected to MVDR beamforming processing based on the spatial filter to obtain the initial speech signal after beamforming.

2. The method according to claim 1, characterized in that Performing feature extraction on the initial speech signal to obtain BFCC features, including: Performing short-time Fourier transform on the initial speech signal to obtain a frequency spectrum; Mapping the spectrum to the Bark frequency scale to obtain a Bark spectrum; Taking the logarithm of the Bark spectrum to obtain a logarithmic Bark spectrum; Performing discrete cosine transform on the logarithmic Bark spectrum to obtain the BFCC feature.

3. The method according to claim 1, characterized in that The BFCC feature is processed using a pre-trained speech / noise separation DNN to obtain a time-frequency mask, including: Inputting the BFCC feature into the pre-trained speech / noise separation DNN, and extracting local features of the BFCC feature through the convolutional neural network of the pre-trained speech / noise separation DNN; Modeling the time domain context information of the local features of the BFCC features using the recursive neural network of the pre-trained speech / noise separation DNN to obtain the time domain context information; fusing the local features of the BFCC features and the temporal context information through the fully connected layer of the pre-trained speech / noise separation DNN to output a soft mask; Binarization is performed on the soft mask to obtain the time-frequency mask.

4. The method according to claim 1, wherein The method of adaptively designing the parameters of the comb filter according to the fundamental frequency characteristics to obtain a target comb filter includes: Determining the order of the comb filter according to the fundamental frequency characteristic, wherein the order is inversely proportional to the fundamental frequency; Determining the passband center frequency of the comb filter according to the fundamental frequency characteristics; Determining the bandwidth of the comb filter according to the fundamental frequency characteristics; The filter coefficients of the comb filter are determined according to the order, the passband center frequency, and the bandwidth to obtain the target comb filter.

5. The method according to claim 1, wherein The step of multiplying the comb-filtered speech signal by the time-frequency mask as a weighting coefficient to obtain energy of a weighted signal includes: The value of the element representing the target speech in the time-frequency mask is set to 1, and the value of the element representing the noise is set to 0, to obtain a speech time-frequency mask; The speech time-frequency mask is used as a weighting coefficient and multiplied by the speech signal after the comb filter to obtain the weighted signal; The average energy of the weighted signal in the time dimension is calculated, and the average energy is used as the target speech energy estimation value.

6. The method according to claim 1, characterized in that described Calculating a frequency-dependent Bark gain coefficient on a Bark frequency scale according to the target speech energy estimate and the noise energy estimate, including: Smoothing the noise energy estimate and the target speech energy estimate on a Bark frequency scale to obtain a smoothed target speech energy estimate and a smoothed noise energy estimate; Calculating the ratio of the smoothed target speech energy estimate to the smoothed noise energy estimate to obtain a signal-to-noise ratio on a Bark frequency scale; The signal-to-noise ratio is mapped to a Bark gain coefficient on a Bark frequency scale through a preset exponential function.

7. A directional sound pickup system for implementing the directional sound pickup method according to claim 1, characterized in that: The system comprises: A data acquisition module is used to perform MVDR beamforming processing on the initial speech signal input by the microphone array to obtain the initial speech signal after beamforming; A feature extraction module is used to extract features from the initial speech signal to obtain BFCC features and fundamental frequency features; A special processing module, configured to process the BFCC features using a pre-trained speech / noise separation DNN to obtain a time-frequency mask; a filtering module, configured to adaptively design parameters of a comb filter according to the fundamental frequency characteristics to obtain a target comb filter, and filter the initial speech signal based on the comb filter to obtain a comb-filtered speech signal; a target speech energy estimation module, configured to multiply the time-frequency mask as a weighting coefficient with the comb-filtered speech signal to obtain the energy of the weighted signal, and use the energy of the weighted signal as the target speech energy estimation value; a noise energy estimation module, configured to multiply the time-frequency mask as a weighting coefficient by the initial speech signal to obtain a noise energy estimation value; A gain coefficient calculation module, configured to calculate a frequency-dependent Bark gain coefficient on a Bark frequency scale based on the target speech energy estimate and the noise energy estimate; A first signal enhancement module is configured to multiply the comb-filtered speech signal by a Bark gain coefficient to enhance the initial speech signal in the time-frequency domain to obtain an enhanced speech signal; The second signal enhancement module is configured to convert the enhanced speech signal from the Bark frequency scale into a time domain signal using an inverse BFCC transform to obtain a target enhanced speech pointing in a target direction.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executed by a method according to any one of claims 1 to 6.

9. An electronic device, characterized in that: The electronic device comprises a processor, a memory and a transceiver, wherein the memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice separation method, voice recognition method and related equipment

    CN110459237A

  • Noise suppression method and device, medium and electronic equipment

    CN113571078A