Multi-source direction of arrival estimation method, apparatus and non-transitory storage medium
Patent Information
- Application Number
- CN202510749627.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
[0004]本申请实施例提供了一种多声源到达方向估计方法、装置及非易失性存储介质,以至少解决由于相关技术中无法对信号的噪声成分和混响成分进行有高效去除导致的多声源到达方向估计结果不准确的技术问题
[0018]In this embodiment, an initial dereverberation signal is obtained by eliminating late reverberation signals in the observed signal, wherein the number of sound sources in the observed signal is one or more; the initial dereverberation signal is non-uniformly divided into multiple first sub-bandwidths, wherein each first sub-bandwidth contains multiple frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; the focusing covariance matrix corresponding to each first sub-bandwidth is determined, and the focusing frequency spatial spectrum of each first sub-bandwidth is obtained based on the focusing covariance matrix; the initial estimation result of the sound source arrival direction of the observed signal is determined based on the focusing frequency spatial spectrum of the first sub-bandwidth; based on the initial estimation result, the observed signal is iteratively processed using a beamformer and a dereverberation filter to obtain the target estimation result of the sound source arrival direction, wherein, in the iterative processing... In each iteration, based on the estimated direction of arrival of the sound sources obtained from the previous iteration, a dereverberation filter is used to eliminate the reverberation signal in the observed signal, and a beamformer is used to eliminate the noise signal in the observed signal. Based on the dereverberated and denoised signal obtained after eliminating the reverberation and noise signals, the estimated direction of arrival of the sound sources for the current iteration is determined. By using filters and beamformers to iteratively eliminate noise and reverberation signals in the observed signal, the goal of efficiently removing noise and reverberation components from the signal is achieved. This improves the accuracy of the multi-source direction of arrival estimation results and solves the technical problem of inaccurate multi-source direction of arrival estimation results caused by the inability to efficiently remove noise and reverberation components from the signal in related technologies.
Smart Images

Figure CN120630101B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio signal processing, and more specifically, to a method, apparatus, and non-volatile storage medium for estimating the direction of arrival of multiple sound sources. Background Technology
[0002] In related technologies, when estimating the direction of arrival (DOA) of a sound source wave, it is impossible to efficiently remove noise and reverberation components from the signal, resulting in inaccurate estimation results.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This application provides a method, apparatus, and non-volatile storage medium for estimating the direction of arrival of multiple sound sources, in order to at least solve the technical problem of inaccurate estimation results of the direction of arrival of multiple sound sources due to the inability to efficiently remove the noise and reverberation components of the signal in related technologies.
[0005] According to one aspect of the embodiments of this application, a method for estimating the direction of arrival (DOA) of multiple sound sources is provided, comprising: eliminating late reverberation signals in an observed signal to obtain an initial dereverberation signal, wherein the number of sound sources in the observed signal is one or more; non-uniformly dividing the initial dereverberation signal into multiple first sub-bandwidths, wherein each first sub-bandwidth contains multiple frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; determining the focusing covariance matrix corresponding to each first sub-bandwidth, and obtaining the sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the focusing covariance matrix; and focusing the frequency spectrum of the sub-bandwidth corresponding to the first sub-bandwidth based on the sub-bandwidth focusing frequency spectrum. The initial estimate of the sound source arrival direction of the observed signal is determined by the rate space spectrum. Based on the initial estimate, the observed signal is iteratively processed using a beamformer and a dereverberation filter to obtain the target estimate of the sound source arrival direction. In each iteration of the iterative processing, based on the sound source arrival direction estimate obtained in the previous iteration, a dereverberation filter is used to eliminate the reverberation signal in the observed signal, and a beamformer is used to eliminate the noise signal in the observed signal. Based on the dereverberation and noise signal obtained after eliminating the reverberation and noise signals, the estimated sound source arrival direction corresponding to the current iteration is determined.
[0006] Optionally, determining the focusing covariance matrix corresponding to each first sub-bandwidth includes: determining the covariance matrix corresponding to each first sub-bandwidth and performing a phase transformation on the covariance matrix; for each first sub-bandwidth, determining the focusing covariance matrix corresponding to each first sub-bandwidth based on the phase-transformed covariance matrix; obtaining the sub-bandwidth focusing frequency space spectrum corresponding to each first sub-bandwidth based on the focusing covariance matrix includes: performing singular value decomposition on the focusing covariance matrix to obtain the noise subspace corresponding to each first sub-bandwidth; and determining the sub-bandwidth focusing frequency space spectrum corresponding to each first sub-bandwidth based on the noise subspace.
[0007] Optionally, the observation signal is iteratively processed using a beamformer and a dereverberation filter to obtain the target estimation result of the sound source arrival direction, including: in each iteration, a beamformer is used to determine the time-varying variance based on the sound source arrival direction estimation result obtained in the previous iteration, wherein the time-varying variance is used to reflect the energy change of the observation signal within a preset frequency range and time period, and the beamformer is used to determine the time-varying variance based on the initial estimation result in the first iteration; a dereverberation filter is used to eliminate the reverberation signal in the observation signal based on the time-varying variance to obtain a dereverberation signal; a beamformer is used to eliminate the noise signal in the dereverberation signal to obtain a dereverberation-denoising signal, wherein the dereverberation-denoising signal includes multiple second sub-bandwidths, each second sub-bandwidth in the dereverberation-denoising signal contains multiple frequency band signals, and there is frequency overlap between adjacent second sub-bandwidths; the focusing covariance matrix corresponding to each second sub-bandwidth is determined, and the sub-bandwidth focusing frequency space spectrum corresponding to each second sub-bandwidth is obtained based on the focusing covariance matrix; the sound source arrival direction estimation result of this iteration is determined based on the sub-bandwidth focusing frequency space spectrum corresponding to the second sub-bandwidth.
[0008] Optionally, the dereverberation filter includes a multi-channel linear prediction filter; using the dereverberation filter to eliminate the reverberation signal in the observed signal based on the time-varying variance to obtain the dereverberation signal includes: sequentially using each channel of the multi-channel linear prediction filter as a reference channel, and using the filter to eliminate the reverberation signal in the observed signal based on the reference channel and the time-varying variance to obtain a dereverberation signal matrix, wherein the dereverberation signal matrix includes the dereverberation sub-signals obtained after eliminating the reverberation signal with each channel as a reference channel; and using the dereverberation signal matrix as the dereverberation signal.
[0009] Optionally, the beamformer includes a multi-target minimum variance distortionless response beamforming filter; using the beamformer to eliminate noise signals in the dereverberation signal to obtain a dereverberation and denoising signal includes: sequentially using each channel of the multi-target minimum variance distortionless response beamforming filter as a reference channel, and using the multi-target minimum variance distortionless response beamforming filter to eliminate noise signals in the dereverberation signal based on the reference channels to obtain a dereverberation and denoising signal.
[0010] Optionally, using a beamformer to eliminate noise signals in the dereverberation signal to obtain a dereverberation and noise-reduced signal includes: when there are multiple sound sources, determining the steering vector of each sound source; and using a beamformer to eliminate noise signals in the dereverberation signal based on the steering vector of each sound source to obtain a dereverberation and noise-reduced signal.
[0011] Optionally, non-uniformly dividing the initial dereverberation signal into multiple first sub-bandwidths includes: discarding the portion of the initial dereverberation signal with frequencies lower than a first preset frequency; for a first signal portion of the initial dereverberation signal with a frequency not less than the first preset frequency and less than a second preset frequency, dividing the first signal portion into multiple first sub-bandwidths according to a first preset frequency spacing, wherein the sub-bandwidths corresponding to two adjacent frequencies in the first signal portion are superimposed with a first preset number of frequency bands; for a second signal portion of the initial dereverberation signal with a frequency not less than the second preset frequency and less than a third preset frequency, dividing the second signal portion into multiple first sub-bandwidths according to a second preset frequency spacing, wherein the sub-bandwidths corresponding to two adjacent frequencies in the second signal portion are superimposed with a first preset number of frequency bands; Two first sub-bandwidths are superimposed with a second preset number of frequency bands; for the third signal portion of the initial dereverberation signal with a frequency not less than a third preset frequency and less than a fourth preset frequency, the third signal portion is divided into multiple first sub-bandwidths according to the third preset frequency spacing, wherein two first sub-bandwidths with adjacent frequencies in the third signal portion are superimposed with a third preset number of frequency bands; for the fourth signal portion of the initial dereverberation signal with a frequency not less than a fourth preset frequency and less than a fifth preset frequency, the fourth signal portion is divided into multiple first sub-bandwidths according to the fourth preset frequency spacing, wherein two first sub-bandwidths with adjacent frequencies in the fourth signal portion are superimposed with a fourth preset number of frequency bands.
[0012] Optionally, the initial estimation result for determining the sound source arrival direction of the observed signal based on the sub-bandwidth focusing frequency spatial spectrum includes: summing the focusing frequency spatial spectra of each sub-bandwidth to obtain the significant peak value of the overall focusing frequency spatial spectrum of each sub-bandwidth; and determining the initial estimation result based on the significant peak value.
[0013] Optionally, the initial estimation results for determining the sound source arrival direction of the observed signal based on the sub-bandwidth focusing frequency spatial spectrum include: determining the significant peaks of each sub-bandwidth focusing frequency spatial spectrum, and the estimation results of the sound source arrival direction corresponding to the significant peaks; determining the significant peak-estimation result set, and determining the distribution information of the significant peak-estimation result set; and determining the initial estimation results based on the distribution information.
[0014] According to another aspect of the embodiments of this application, a multi-source direction-of-arrival estimation device is also provided, comprising: a first processing module, configured to eliminate late reverberation signals in an observed signal to obtain an initial dereverberation signal, wherein the number of sound sources in the observed signal is one or more; a second processing module, configured to non-uniformly divide the initial dereverberation signal into multiple first sub-bandwidths, wherein each first sub-bandwidth contains multiple frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; a third processing module, configured to determine the focusing covariance matrix corresponding to each first sub-bandwidth, and obtain the sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the focusing covariance matrix; and a fourth processing module, configured to... The first sub-bandwidth focuses the frequency spatial spectrum to determine the initial estimation result of the sound source arrival direction of the observed signal; the fifth processing module is used to iteratively process the observed signal using a beamformer and a dereverberation filter based on the initial estimation result to obtain the target estimation result of the sound source arrival direction. In each iteration of the iterative processing, based on the sound source arrival direction estimation result obtained in the previous iteration, a dereverberation filter is used to eliminate the reverberation signal in the observed signal, and a beamformer is used to eliminate the noise signal in the observed signal. Based on the dereverberation and noise signal obtained after eliminating the reverberation and noise signals, the estimation result of the sound source arrival direction corresponding to the current iteration is determined.
[0015] According to an embodiment of this application, a non-volatile storage medium is also provided, wherein a program is stored in the non-volatile storage medium, and the program controls the device where the non-volatile storage medium is located to execute a multi-source arrival direction estimation method when it runs.
[0016] According to an embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes a multi-source direction of arrival estimation method during runtime.
[0017] According to an embodiment of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements a multi-source direction-of-arrival estimation method.
[0018] In this embodiment, an initial dereverberation signal is obtained by eliminating late reverberation signals in the observed signal, wherein the number of sound sources in the observed signal is one or more; the initial dereverberation signal is non-uniformly divided into multiple first sub-bandwidths, wherein each first sub-bandwidth contains multiple frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; the focusing covariance matrix corresponding to each first sub-bandwidth is determined, and the focusing frequency spatial spectrum of each first sub-bandwidth is obtained based on the focusing covariance matrix; the initial estimation result of the sound source arrival direction of the observed signal is determined based on the focusing frequency spatial spectrum of the first sub-bandwidth; based on the initial estimation result, the observed signal is iteratively processed using a beamformer and a dereverberation filter to obtain the target estimation result of the sound source arrival direction, wherein, in the iterative processing... In each iteration, based on the estimated direction of arrival of the sound sources obtained from the previous iteration, a dereverberation filter is used to eliminate the reverberation signal in the observed signal, and a beamformer is used to eliminate the noise signal in the observed signal. Based on the dereverberated and denoised signal obtained after eliminating the reverberation and noise signals, the estimated direction of arrival of the sound sources for the current iteration is determined. By using filters and beamformers to iteratively eliminate noise and reverberation signals in the observed signal, the goal of efficiently removing noise and reverberation components from the signal is achieved. This improves the accuracy of the multi-source direction of arrival estimation results and solves the technical problem of inaccurate multi-source direction of arrival estimation results caused by the inability to efficiently remove noise and reverberation components from the signal in related technologies. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0020] Figure 1 This is a schematic diagram of the structure of a computer terminal (mobile terminal) according to an embodiment of this application;
[0021] Figure 2 This is a flowchart illustrating a multi-source direction-of-arrival estimation method according to an embodiment of this application;
[0022] Figure 3 This is a flowchart illustrating a multi-source direction-of-arrival estimation process according to an embodiment of this application;
[0023] Figure 4 This is a schematic diagram of a multi-source arrival direction estimation device provided according to an embodiment of this application. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained below:
[0027] Direction of Arrival (DOA): refers to the azimuth information of the signal source incident on the array. It can effectively characterize the spatial positional relationship between the signal source and the array, and plays a very important role in signal processing fields such as speech, antenna, sonar, ultrasound, and radar.
[0028] Frequency focusing is a processing method that uses a focusing matrix in the frequency domain to transfer signals from different frequency bands to the same frequency band. It has wide applications in signal processing fields such as radar, sonar, antennas, and speech.
[0029] Convolutional beamforming is a type of enhancement filter that uses historical frame information and current frame information to jointly achieve dereverberation and denoising tasks. It can use multi-channel linear prediction to achieve dereverberation, use beamforming to achieve denoising, and use time-varying variance to combine the two for alternating optimization.
[0030] The purpose of multi-source localization is to estimate the location information of multiple sound sources in acoustic space, providing spatial information for subsequent tasks such as sound source enhancement, sound source separation, and target sound source extraction. The objectives of multi-source localization can be mainly divided into three categories: sound source location estimation, time difference of arrival (TDOA) estimation, and direction of arrival (DOA) estimation. Among these, sound source location estimation is often difficult to use in practical tasks due to its high algorithm complexity or the difficulty in accurately estimating the azimuth distance. TDOA estimation accuracy is limited by the element distance of the microphone array, which has led to wider research and application of DOA estimation.
[0031] In ideal far-field environments with no noise or reverberation, numerous methods have been proven effective in estimating the DOA (Directivity to Orientation) information of sound sources. Examples include the Steered response power (SRP) method based on beamforming, the Multi-signal classification (MUSIC) method based on signal subspace, and the SRP-based phase transform (SRP-PHAT) method. However, noise is unavoidable in real-world environments, and reverberation is present in typical indoor environments, significantly increasing the difficulty of DOA estimation. Furthermore, the random spatial distribution of sound sources in multi-source scenarios necessitates algorithms with high resolution to cover situations where sound sources have similar azimuth angles. Therefore, multi-source localization in complex scenarios is one of the most challenging problems in acoustics. To address this issue, several neural network-based methods have been proposed, such as the MUSIC algorithm based on deep neural networks (DNNs) and the SRP-PHAT algorithm based on conventional neural networks (CNNs). However, due to differences in array topology, environmental noise, and reverberation, there is a lack of general sound source localization datasets with precise annotations, resulting in poor generalization performance of many neural network-based methods in practical tasks.
[0032] To address the aforementioned issues, this application provides a solution that combines frequency-focusing-based sound source localization with reverberation- and noise-reducing convolutional beamforming. This proposed method for estimating the DOA (Direct Occurrence Area) of multiple sound sources using a joint frequency-focusing and convolutional beamforming approach offers at least the following technical advantages: First, based on the distribution of frequency components in the speech signal, the speech bandwidth is non-uniformly divided into multiple sub-bandwidths, and these sub-bandwidths are overlapped to improve feature stability and frequency dimension resolution. Second, a phase transformation method is applied to the signal covariance matrix to obtain the frequency-focusing spatial spectrum of the phase transformation, mitigating the impact of amplitude feature instability and improving the noise and reverberation resistance of the frequency-focusing spatial spectrum. Third, addressing the issue of drastic performance degradation in low signal-to-noise ratio and strong reverberation environments, a multi-source convolutional beamformer is modeled using multi-channel linear prediction (MCLP) and minimum variance distortion-less response (MVDR) beamforming to suppress reverberation and noise effects. This is then jointly optimized with frequency-focusing-based DOA through an iterative alternation process. The details are described below.
[0033] According to an embodiment of this application, a method embodiment for estimating the direction of arrival of multiple sound sources is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0034] The methods and embodiments provided in this application can be executed on mobile terminals, computer terminals, or similar computing devices. Figure 1 A hardware block diagram of a computer terminal (or mobile device) for implementing a multi-source direction-of-arrival estimation method is shown. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0035] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0036] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the multi-source direction of arrival estimation method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned multi-source direction of arrival estimation method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0037] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0038] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0039] Under the aforementioned operating environment, embodiments of this application provide a method for estimating the direction of arrival of multiple sound sources, such as... Figure 2 As shown, the method includes the following steps:
[0040] Step S202: Eliminate the late reverberation signal in the observed signal to obtain the initial dereverberation signal, wherein the number of sound sources in the sound source wave is one or more;
[0041] In some embodiments of this application, the observation signal can be an observation signal obtained using a microphone array. Optionally, in a far-field environment, it can be obtained from... M The frequency domain model of the observed signal acquired by the microphone array with one element can be represented as follows:
[0042] (1)
[0043] Where X( k , l ) M×1 For the observed signal, X dir ( k , l ) M×1 For the direct sound component, X early ( k , l ) M ×1 For early reverberation components, X late ( k , l ) M×1 For late-stage reverberation components, N( k , l ) M×1 Noise components; k and l These are indices representing frequency and frame number, respectively. Let X represent the complex space. In the above formula, X... dir ( k , l This can be represented as follows:
[0044] (2)
[0045] Where, a( θ q , k , l ) is the direct sound steering vector. For the first The angle of incidence of a sound source, S q ( k , l ) is the first One sound source signal, Q This represents the number of sound sources. The purpose of sound source DOA estimation is to estimate the incident angle of each sound source. Give me an estimate.
[0046] In some embodiments of this application, a multi-channel linear prediction filter (MCLP) can be used to eliminate late reverberation, thereby mitigating the impact of the reverberation component on DOA estimation. Assuming the dereverberated signal follows a zero-mean complex Gaussian distribution, the MCLP filter G(…) can be optimized based on the maximum likelihood objective function of the weighted prediction error (WPE) algorithm. k , l The solution is as follows:
[0047] (3)
[0048] Among them, G( k , l ) MLg×M , L g The length (also called the order) of the linear prediction filter. =[X T ( k , l – b ),…, X T ( k , l – b – L g +1)] T MLg×1 A vector stacked from historical frames. b For linear prediction time extension, σ ( k , l The variance is time-varying. The superscript H indicates the conjugate transpose. τ This represents the index of the historical frames used for linear predictive filtering.
[0049] Based on the above MCLP filter G( k , l The initial dereverberation signal Z( ) can be obtained after dereverberation. k , l ) M×1 The solution is as follows:
[0050] (4)
[0051] In the above formula, G H This represents the conjugate transpose of G.
[0052] Step S204: The initial dereverberation signal is non-uniformly divided into multiple first sub-bandwidths, wherein each first sub-bandwidth contains multiple frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths;
[0053] In the technical solution provided in step S204, the step of non-uniformly dividing the initial dereverberation signal into multiple first sub-bandwidths includes: discarding the portion of the initial dereverberation signal with a frequency lower than a first preset frequency; for the first signal portion of the initial dereverberation signal with a frequency not less than the first preset frequency and less than the second preset frequency, dividing the first signal portion into multiple first sub-bandwidths according to a first preset frequency spacing, wherein the sub-bandwidths corresponding to two adjacent first sub-bandwidths in the first signal portion are superimposed with a first preset number of frequency bands; for the second signal portion of the initial dereverberation signal with a frequency not less than the second preset frequency and less than the third preset frequency, dividing the second signal portion into multiple first sub-bandwidths according to a second preset frequency spacing, wherein the second signal portion... In the first signal portion, the sub-bandwidths corresponding to two adjacent first sub-bandwidths are superimposed with a second preset number of frequency bands; for the third signal portion where the frequency of the initial dereverberation signal is not less than the third preset frequency and less than the fourth preset frequency, the third signal portion is divided into multiple first sub-bandwidths according to the third preset frequency spacing, wherein the sub-bandwidths corresponding to two adjacent first sub-bandwidths in the third signal portion are superimposed with a third preset number of frequency bands; for the fourth signal portion where the frequency of the initial dereverberation signal is not less than the fourth preset frequency and less than the fifth preset frequency, the fourth signal portion is divided into multiple first sub-bandwidths according to the fourth preset frequency spacing, wherein the sub-bandwidths corresponding to two adjacent first sub-bandwidths in the fourth signal portion are superimposed with a fourth preset number of frequency bands.
[0054] Optionally, given the characteristics of concentrated low-frequency energy and gradual attenuation of high-frequency energy in speech signals, a non-uniform sub-bandwidth partitioning strategy is adopted to more effectively utilize the frequency band information of the speech signal. Specifically, the first formant of the speech signal is generally between 300 Hz and 1000 Hz, the second formant is generally between 800 Hz and 2500 Hz, the third formant is generally between 2000 Hz and 3500 Hz, and higher-order formants are above 3500 Hz. Furthermore, the speech fundamental tone is generally between 85 Hz and 300 Hz. Considering the above factors, the non-uniform flat band partitioning strategy can be described in the following steps:
[0055] (1) Discard the noise component below 100 Hz;
[0056] (2) Between 100 Hz and 900 Hz, the signal is divided into sub-bandwidths at intervals of 100 Hz, and two frequency bands are superimposed between the bandwidths;
[0057] (3) Between 900 Hz and 2300 Hz, the signal is divided into sub-bandwidths at intervals of 200 Hz, and three frequency bands are superimposed between the bandwidths;
[0058] (4) Between 2300 Hz and 3500 Hz, the signal is divided into sub-bandwidths at intervals of 400 Hz, and four frequency bands are superimposed between the bandwidths;
[0059] (5) Between 3500 Hz and 8000 Hz, the signal is divided into sub-bandwidths at intervals of 800 Hz, and 6 frequency bands are superimposed between the bandwidths;
[0060] Assuming the audio signal sampling rate is 16 kHz, the short-time fourier transform (STFT) frame length and window length are both 512 samples, and the frame shift is 50%, then a total of 47 sub-bandwidths are divided according to the above method.
[0061] Step S206: Determine the focusing covariance matrix corresponding to each first sub-bandwidth, and obtain the sub-bandwidth focusing frequency space spectrum corresponding to each first sub-bandwidth based on the focusing covariance matrix.
[0062] In the technical solution provided in step S206, such as Figure 3 As shown, the steps for determining the focusing covariance matrix corresponding to each first sub-bandwidth include: determining the covariance matrix corresponding to each first sub-bandwidth and performing phase transformation processing on the covariance matrix; for each first sub-bandwidth, determining the focusing covariance matrix corresponding to each first sub-bandwidth based on the covariance matrix after phase transformation processing; and obtaining the sub-bandwidth focusing frequency space spectrum corresponding to each first sub-bandwidth based on the focusing covariance matrix, which includes: performing singular value decomposition on the focusing covariance matrix to obtain the noise subspace corresponding to each first sub-bandwidth; and determining the sub-bandwidth focusing frequency space spectrum corresponding to each first sub-bandwidth based on the noise subspace.
[0063] Additionally from Figure 3 As can be seen from this, determining the initial estimation result requires performing the following steps: calculating the covariance matrix, phase transformation of the covariance matrix, determining the focusing covariance matrix within the sub-bandwidth, determining the focusing covariance matrix within the sub-bandwidth, calculating the noise subspace, determining the focusing frequency space spectrum within the sub-bandwidth, and determining the initial estimation result.
[0064] Optionally, when calculating the covariance matrix, assume Z i,j ( k , l ) M×1 For the first i The first of the sub-bandwidths j For a signal with a frequency band, its covariance matrix can be calculated as follows:
[0065] (5)
[0066] Among them, R i,j ( k , l ) M×M Let covariance matrix be the variance matrix. τ For the index of snapshots, Γ This is the number of snapshots, indicated by the superscript. H Represents the conjugate transpose of a matrix. Γ -1 is the index of the maximum number of snapshots.
[0067] Optionally, considering the instability of amplitude information caused by various factors such as noise, signal absorption, and attenuation in the far-field environment, a phase transformation method is used to avoid the influence of amplitude factors. The phase transformation of equation (5) is shown in the following equation:
[0068] (6)
[0069] In the above formula, |.| represents the operation of taking the film value.
[0070] Optionally, when calculating the focusing covariance matrix within the sub-bandwidth, the first... i The first of the sub-bandwidths j The focusing matrix of the signal in each frequency band is C i,j ( f k , l If the focusing covariance matrix of this frequency band is:
[0071] (7)
[0072] in, f k To focus the reference frequency, the focusing matrix C i,j ( f k , l The following method can be used to solve this problem:
[0073] (8)
[0074] Among them, U s ( f k , l U is the signal subspace of the covariance matrix of the focused reference frequency band. s (k , l ) represents the signal subspace of the current frequency band covariance matrix. The signal subspace can be obtained by the eigenvectors / singular vectors corresponding to significantly large eigenvalues / singular values. C in equation (8) i,j ( f k , l The solution can be obtained through the bilateral transformation method.
[0075] Optionally, after obtaining the focusing matrix, the focusing covariance matrix shown in formula (7) can be further obtained. Then, the focusing covariance matrices of multiple frequency bands within the sub-bandwidth can be averaged to obtain the sub-bandwidth focusing covariance matrix, as shown in the following formula:
[0076] (9)
[0077] in, J i For the first i The total number of frequency bands within a sub-bandwidth.
[0078] Optionally, when calculating the noise subspace, singular value decomposition can be performed on the focusing covariance matrix of the corresponding frequency band to obtain the noise subspace of the corresponding frequency band, as shown in the following equation:
[0079] (10)
[0080] (11)
[0081] in, P For a significantly large number of eigenvalues, V i ( f k , l ), Σ i ( f k , l ) and U i ( f k , l ) are the left singular vector, the singular value diagonal matrix, and the right singular matrix, respectively. For the noise subspace, To arrange the first from the large direction P +1 singular value vectors corresponding to singular values.
[0082] Optionally, when calculating the sub-bandwidth focused frequency spatial spectrum, the first can be obtained based on equation (11). i The focusing frequency spectrum of each sub-bandwidth is:
[0083] (12)
[0084] Where, a( , f k , l )for Directional guidance vector, [0°, 360°] represents the guide vector search angle. I By repeating this process with each sub-bandwidth, a broadband frequency-focused spatial spectrum can be obtained. The broadband frequency-focused spatial spectrum can be viewed as... I The matrix is obtained by splicing together the spatial spectra of the narrowband frequencies corresponding to each sub-bandwidth.
[0085] Step S208: Determine the initial estimation result of the sound source arrival direction of the observed signal based on the spatial spectrum of the sub-bandwidth focusing frequency corresponding to the first sub-bandwidth;
[0086] There are generally two methods for estimating DOA based on broadband frequency-focused spatial spectrum. The first method involves using the broadband frequency-focused spatial spectrum... I The first method involves summing the focused spatial spectra of each narrowband frequency, and then estimating the DOA information of the sound source from the significant peaks of the summed narrowband spatial spectra. The second method involves constructing a peak set based on the significant peaks of each narrowband frequency focused spatial spectrum, and then using a statistical approach to estimate the final DOA information from the peak set. When using the second method, such as... Figure 3 As shown, Gaussian distribution or von Neumann distribution can be selected. The results of fitting the histogram statistics to the Mises distribution are used, and the estimated DOA information is obtained based on the peak value of the fitted curve θ=[ θ 1, θ 2 ,..., θ Q ].
[0087] In some embodiments of this application, the initial estimation result of determining the sound source arrival direction of the observed signal based on the sub-bandwidth focusing frequency spatial spectrum using method one includes: summing the focusing frequency spatial spectra of each sub-bandwidth to obtain the significant peak value of the overall focusing frequency spatial spectrum of each sub-bandwidth; and determining the initial estimation result based on the significant peak value.
[0088] The advantage of Method 1 lies in its computational simplicity, but the summation can easily lead to a decrease in the resolution of the spatial spectrum peaks, resulting in significant errors in the DOA estimation results, and even the inability to effectively detect some weaker sound sources. Method 2, on the other hand, is more computationally complex, but it can filter out the influence of many clutter noises and does not suffer from the problem of decreased spatial spectrum peak resolution due to summation.
[0089] In some embodiments of this application, the initial estimation result of the sound source arrival direction of the observed signal using Method 2 based on the sub-bandwidth focusing frequency spatial spectrum includes: determining the significant peaks of each sub-bandwidth focusing frequency spatial spectrum, and the estimation results of the sound source arrival direction corresponding to the significant peaks; determining the significant peak-estimation result set, and determining the distribution information of the significant peak-estimation result set; and determining the initial estimation result based on the distribution information. Specific implementation details include the following steps:
[0090] 1. Based on the significant peak values of the focused spatial spectrum of each narrowband frequency, obtain the corresponding DOA values, and construct a peak DOA set from these DOA values;
[0091] 2. Based on the peak DOA set, select an appropriate quantization precision and construct a statistical histogram, such as... Figure 3 As shown;
[0092] 3. Adopt Feng Mises distribution, fit histogram, and obtain DOA distribution curve;
[0093] 4. Based on the DOA distribution curve, obtain the DOA interval corresponding to the significant peak value, and take the midpoint of the interval as the estimated DOA information.
[0094] Step S210: Based on the initial estimation result, the observed signal is iteratively processed using a beamformer and a dereverberation filter to obtain the target estimation result of the sound source arrival direction. In each iteration of the iterative processing, based on the sound source arrival direction estimation result obtained in the previous iteration, a dereverberation filter is used to eliminate the reverberation signal in the observed signal, and a beamformer is used to eliminate the noise signal in the observed signal. Based on the dereverberation and noise signal obtained after eliminating the reverberation and noise signals, the sound source arrival direction estimation result corresponding to the current iteration is determined.
[0095] In some embodiments of this application, such as Figure 3As shown, the steps for iteratively processing the observed signal using a beamformer and a dereverberation filter to obtain the target estimation result of the sound source arrival direction include: in each iteration, using a beamformer, determining the time-varying variance based on the sound source arrival direction estimation result obtained in the previous iteration, wherein the time-varying variance is used to reflect the energy change of the observed signal within a preset frequency range and time period; in the first iteration, using a beamformer to determine the time-varying variance based on the initial estimation result; using a dereverberation filter to eliminate the reverberation signal in the observed signal based on the time-varying variance to obtain a dereverberated signal; using a beamformer to eliminate the noise signal in the dereverberated signal to obtain a dereverberated and denoised signal, wherein the dereverberated and denoised signal includes multiple second sub-bandwidths, each second sub-bandwidth in the dereverberated and denoised signal contains multiple frequency band signals, and there is frequency overlap between adjacent second sub-bandwidths; determining the focusing covariance matrix corresponding to each second sub-bandwidth, and obtaining the sub-bandwidth focusing frequency space spectrum corresponding to each second sub-bandwidth based on the focusing covariance matrix; and determining the sound source arrival direction estimation result of this iteration based on the sub-bandwidth focusing frequency space spectrum corresponding to the second sub-bandwidth. in Figure 3 The blue dashed line represents the transmission flow of the initial estimation result, while the red solid line represents the data transmission direction of the iterative process.
[0096] It should be noted that although the combined formulas (4) and (12) reduce the influence of reverberation on the sound source location, this process does not consider the influence of noise. In order to further eliminate the noise component in the observed signal, a combined beamformer can be used to suppress the noise component. The noise reduction process requires keeping the multi-source signal undistorted, which makes general beamformers unable to perform this task. Therefore, in this embodiment, a multi-target minimum variance distortionless response (MT-MVDR) beamformer for multi-source noise reduction is modeled to ensure that the direct-arrival (DOA) signal of the multi-source is undistorted. For ease of description, a dual-source scenario is used as an example in the following process. The MT-MVDR beamformer can then be modeled as:
[0097] (13)
[0098] Among them, W( k , l ) is the filter weight vector for beamforming, a( θ 1, k , l ) and a( θ 2, k , l ) represents the steering vectors of sound source 1 and sound source 2. θ 1 and θ 2 represents the DOA information for sound source 1 and sound source 2, respectively. (Superscript) HThis indicates the conjugate transpose. R zz ( k , l Let be the covariance matrix of the dereverberated signal, which can be represented as follows:
[0099] (14)
[0100] Based on the Lagrange multiplier method, the solution to equation (13) is:
[0101] (15)
[0102] in,
[0103] (16)
[0104] (17)
[0105] (18)
[0106] (19)
[0107] Based on the denoised signal from equation (4) and the beamforming denoising method from equation (15), the denoised and reverberated signal can be obtained:
[0108] (20)
[0109] In some embodiments of this application, the dereverberation filter includes a multi-channel linear prediction filter; using the dereverberation filter to eliminate the reverberation signal in the observed signal based on the time-varying variance to obtain a dereverberation signal includes: sequentially using each channel of the multi-channel linear prediction filter as a reference channel, and using the filter to eliminate the reverberation signal in the observed signal based on the reference channel and the time-varying variance to obtain a dereverberation signal matrix, wherein the dereverberation signal matrix includes a dereverberation sub-signal obtained after eliminating the reverberation signal with each channel as a reference channel; and using the dereverberation signal matrix as the dereverberation signal.
[0110] In some embodiments of this application, the beamformer includes a multi-target minimum variance distortionless response beamforming filter; using the beamformer to eliminate noise signals in the dereverberation signal to obtain a dereverberation and denoising signal includes: sequentially using each channel of the multi-target minimum variance distortionless response beamforming filter as a reference channel, and using the multi-target minimum variance distortionless response beamforming filter to eliminate noise signals in the dereverberation signal according to the reference channels to obtain a dereverberation and denoising signal.
[0111] In some embodiments of this application, by adjusting a( θ 1, k ,l ) and a( θ 2, k , l By selecting different reference channels, multi-channel denoised and denoised signals can be obtained, which can then be used for subsequent processing. Figure 3 As shown, a more robust DOA estimation is achieved by repeatedly executing the phase transform-based frequency-focusing DOA estimation method. However, although the cascaded WPE dereverberation filter (MCLP filter) and MT-MVDR beamformer denoising system can suppress the effects of reverberation and noise, the performance of the cascaded system is often suboptimal. Therefore, a convolutional beamformer is constructed based on the WPE dereverberation filter and MT-MVDR beamformer to achieve better performance.
[0112] As an alternative implementation, in convolutional beamformers, the WPE dereverberation filter and the MT-MVDR beamformer are often correlated through time-varying variance to achieve an alternating optimization process. First, the time-varying variance needs to be initialized. Let the... l Using one channel as a reference channel, a cascaded WPE dérenoising system and a multi-target MVDR beamforming denoising system are used to estimate the initial time-varying variance in order to enable faster algorithm convergence. σ 1,0 ( k , l Based on the result of equation (20), it can be expressed as follows:
[0113] (twenty one)
[0114] in, σ 1,0 ( k , l The 1 in ) represents the first l There are 1 channel, where 0 represents the initial time-varying variance, Z1( k , l ) indicates that the first l Each channel serves as a reference channel for the dereverberation signal. By selecting different reference channels, one can obtain... σ 0( k , l )=[ σ 1,0 ( k , l ), σ 2,0 ( k , l ),..., σ M,0 ( k , l )).
[0115] Optional, such as Figure 3 As shown, based on the time-varying variance of formula (21), the WPE dereverberation filter and MT-MVDR beamformer shown in formulas (3) and (15) can be jointly iterated. The specific iteration process is as follows:
[0116] (a) Clause WPE dereverb in the next iteration
[0117] No. In the next iteration, the first m A dereverberation filter with one reference channel can be represented as follows:
[0118] (twenty two)
[0119] Among them, G m, ( k , l )and σ m, ( k , l ) are respectively the first In the next iteration, the first m Each channel serves as a reference channel for the dereverberation filter and time-varying variance.
[0120] Based on equation (22), the de-reverberation signal can be calculated as follows:
[0121] (twenty three)
[0122] By selecting different reference channels, the first... The dereverberation signal matrix Z in the next iteration ( k , l )=[Z 1, ( k , l ), Z 2, ( k , l ),...,Z M, ( k , l )).
[0123] (b) The Multi-target beamforming denoising in the next iteration
[0124] No. In the next iteration, the first mA denoised beamformer with one reference channel can be represented as follows:
[0125] (twenty four)
[0126] (25)
[0127] Based on equations (23) and (24), the denoised signal can be calculated as follows:
[0128] (26)
[0129] The corresponding time-varying variance can be updated as follows:
[0130] (27)
[0131] By selecting different reference channels, the first... The time-varying variance matrix of the next iteration ( k , l )=[ σ 1, ( k , l ), σ 2, ( k , l ),..., σ M, ( k , l It's important to note that each channel will be selected as the reference channel. Furthermore, the reference channel for the denoising process and the reference channel for the dereverberation process are the same.
[0132] The multi-channel denoised signal can then be represented as follows based on formula (26):
[0133] (28)
[0134] Then you can ( k , l Replace the initial dereverberation signal Z( k , l ), re-execute Figure 3 The frequency focusing DOA estimation process based on phase transformation is used to obtain the estimation results of this iteration.
[0135] As an optional implementation, the step of using a beamformer to eliminate noise signals in the dereverberation signal to obtain a dereverberation and noise-reduced signal includes: when there are multiple sound sources, determining the steering vector of each sound source; and using a beamformer to eliminate noise signals in the dereverberation signal based on the steering vector of each sound source to obtain a dereverberation and noise-reduced signal.
[0136] It can be seen that, as Figure 3 As shown, each iteration includes three core steps: frequency focusing DOA estimation based on non-uniform frequency band division, demeveraging based on MCLP, and denoising based on multi-target MVDR beamforming. The steps of each iteration are as follows:
[0137] (1) First, the initial DOA information θ The 0 is fed into the multi-target MVDR beamforming denoising module to obtain the updated time-varying variance. ( k , l );
[0138] (2) Secondly, ( k , l Update the MCLP filter, optimize the convolutional beamforming enhancement system combining MCLP filtering and multi-target MVDR beamforming, and obtain the optimized enhanced signal. ( k , l );
[0139] (3) Then, ( k , l The DOA information is then fed into the frequency-focusing DOA estimation module based on non-uniform frequency band division to re-estimate the DOA information. ;
[0140] (4) Next, The signal is fed into a multi-target MVDR beamforming denoising module for a new round of iteration.
[0141] When the estimated DOA converges or reaches the maximum number of iterations, the final estimated multi-source DOA information can be obtained.
[0142] An initial dereverberation signal is obtained by eliminating late reverberation in the observed signal, where the number of sound sources is one or more. The initial dereverberation signal is divided into multiple sub-bandwidth signals, each corresponding to a sub-bandwidth, with overlapping frequencies between different sub-bandwidths. The covariance matrix corresponding to the multiple sub-bandwidth signals is determined, and the focusing frequency spatial spectrum of the sub-bandwidth signals is determined based on the covariance matrix. An initial estimate of the sound source arrival direction of the observed signal is determined based on the sub-bandwidth focusing frequency spatial spectrum. A beamformer and filter are used to iteratively process the initial estimate and the observed signal to obtain the target estimate of the sound source arrival direction. In each iteration of the iterative processing, the previous round's calculation is used as a reference. The estimated direction of arrival (DOA) of the sound sources obtained through the iterative process is used to eliminate reverberation signals in the observed signals using filters and noise signals using beamformers. The estimated DOA of the sound sources in this iteration is determined based on the denoised and denoised signals obtained after eliminating reverberation and noise signals. By using filters and beamformers to iteratively eliminate noise and reverberation signals in the observed signals, the goal of efficiently removing noise and reverberation components from the signals is achieved. This improves the accuracy of the DOA estimation results for multiple sound sources and solves the technical problem of inaccurate DOA estimation results caused by the inability to efficiently remove noise and reverberation components from signals in related technologies.
[0143] This application provides a multi-source direction-of-arrival estimation device. Figure 4 This is a schematic diagram of the device. From Figure 4As can be seen from the diagram, the device includes: a first processing module 40, used to eliminate late reverberation signals in the observed signal to obtain an initial dereverberation signal, wherein the number of sound sources in the observed signal is one or more; a second processing module 42, used to non-uniformly divide the initial dereverberation signal into multiple first sub-bandwidths, wherein each first sub-bandwidth contains multiple frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; a third processing module 44, used to determine the focusing covariance matrix corresponding to each first sub-bandwidth, and obtain the sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the focusing covariance matrix; and a fourth processing module 46, used to determine the focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the sub-bandwidth focusing covariance matrix. The bandwidth focusing frequency spatial spectrum determines the initial estimation result of the sound source arrival direction of the observed signal; the fifth processing module 48 is used to iteratively process the observed signal using a beamformer and a dereverberation filter based on the initial estimation result to obtain the target estimation result of the sound source arrival direction. In each iteration of the iterative processing, based on the sound source arrival direction estimation result obtained in the previous iteration, a dereverberation filter is used to eliminate the reverberation signal in the observed signal, and a beamformer is used to eliminate the noise signal in the observed signal. Based on the dereverberation and noise signal obtained after eliminating the reverberation and noise signals, the estimation result of the sound source arrival direction corresponding to the current iteration is determined.
[0144] In some embodiments of this application, the second processing module 42 non-uniformly divides the initial dereverberation signal into multiple first sub-bandwidths, including: discarding the portion of the initial dereverberation signal with a frequency lower than a first preset frequency; for a first signal portion of the initial dereverberation signal with a frequency not less than the first preset frequency and less than a second preset frequency, dividing the first signal portion into multiple first sub-bandwidths according to a first preset frequency spacing, wherein the sub-bandwidths corresponding to two adjacent first sub-bandwidths in the first signal portion are superimposed with a first preset number of frequency bands; for a second signal portion of the initial dereverberation signal with a frequency not less than the second preset frequency and less than a third preset frequency, dividing the second signal portion into multiple first sub-bandwidths according to a second preset frequency spacing, wherein the second signal portion... In the first signal portion, the sub-bandwidths corresponding to two adjacent first sub-bandwidths are superimposed with a second preset number of frequency bands; for the third signal portion where the frequency of the initial dereverberation signal is not less than the third preset frequency and less than the fourth preset frequency, the third signal portion is divided into multiple first sub-bandwidths according to the third preset frequency spacing, wherein the sub-bandwidths corresponding to two adjacent first sub-bandwidths in the third signal portion are superimposed with a third preset number of frequency bands; for the fourth signal portion where the frequency of the initial dereverberation signal is not less than the fourth preset frequency and less than the fifth preset frequency, the fourth signal portion is divided into multiple first sub-bandwidths according to the fourth preset frequency spacing, wherein the sub-bandwidths corresponding to two adjacent first sub-bandwidths in the fourth signal portion are superimposed with a fourth preset number of frequency bands.
[0145] In some embodiments of this application, the step of the third processing module 44 in determining the focusing covariance matrix corresponding to each first sub-bandwidth includes: determining the covariance matrix corresponding to each first sub-bandwidth and performing phase transformation processing on the covariance matrix; for each first sub-bandwidth, determining the focusing covariance matrix corresponding to each first sub-bandwidth based on the covariance matrix after phase transformation processing. The step of obtaining the sub-bandwidth focusing frequency space spectrum corresponding to each first sub-bandwidth based on the focusing covariance matrix includes: performing singular value decomposition on the focusing covariance matrix to obtain the noise subspace corresponding to each first sub-bandwidth; and determining the sub-bandwidth focusing frequency space spectrum corresponding to each first sub-bandwidth based on the noise subspace.
[0146] In some embodiments of this application, the step of the fourth processing module 46 in determining the initial estimation result of the sound source arrival direction of the observed signal based on the sub-bandwidth focusing frequency spatial spectrum includes: summing the sub-bandwidth focusing frequency spatial spectra to obtain the significant peak value of the overall sub-bandwidth focusing frequency spatial spectrum; and determining the initial estimation result based on the significant peak value.
[0147] In some embodiments of this application, the step of the fourth processing module 46 in determining the initial estimation result of the sound source arrival direction of the observed signal based on the sub-bandwidth focusing frequency spatial spectrum includes: determining the significant peaks of each sub-bandwidth focusing frequency spatial spectrum and the estimation result of the sound source arrival direction corresponding to the significant peaks; determining the significant peak-estimation result set and determining the distribution information of the significant peak-estimation result set; and determining the initial estimation result based on the distribution information.
[0148] In some embodiments of this application, the fifth processing module 48 uses a beamformer and a dereverberation filter to iteratively process the observed signal to obtain a target estimation result for the direction of arrival of the sound source. This includes: in each iteration, using a beamformer, determining the time-varying variance based on the estimation result of the direction of arrival of the sound source obtained in the previous iteration. The time-varying variance reflects the energy change of the observed signal within a preset frequency range and time period. In the first iteration, the beamformer determines the time-varying variance based on the initial estimation result. The dereverberation filter is then used to eliminate reverberation in the observed signal based on the time-varying variance. The signal is obtained by first obtaining a denoised signal; then, a beamformer is used to eliminate noise signals in the denoised signal to obtain a denoised and denoised signal. The denoised and denoised signal includes multiple second sub-bandwidths, each of which contains multiple frequency band signals, and there is frequency overlap between adjacent second sub-bandwidths. The focusing covariance matrix corresponding to each second sub-bandwidth is determined, and the focusing frequency space spectrum of each second sub-bandwidth is obtained based on the focusing covariance matrix. The estimation result of the sound source arrival direction for this iteration is determined based on the focusing frequency space spectrum of the second sub-bandwidth.
[0149] In some embodiments of this application, the dereverberation filter includes a multi-channel linear prediction filter; the fifth processing module 48 uses the dereverberation filter to eliminate the reverberation signal in the observed signal based on the time-varying variance to obtain a dereverberation signal, which includes: sequentially using each channel of the multi-channel linear prediction filter as a reference channel, and using the filter to eliminate the reverberation signal in the observed signal based on the reference channel and the time-varying variance to obtain a dereverberation signal matrix, wherein the dereverberation signal matrix includes a dereverberation sub-signal obtained after eliminating the reverberation signal with each channel as a reference channel; and using the dereverberation signal matrix as the dereverberation signal.
[0150] In some embodiments of this application, the beamformer includes a multi-target minimum variance distortionless response beamforming filter; the fifth processing module 48 uses the beamformer to eliminate noise signals in the dereverberation signal to obtain a dereverberation and denoising signal, which includes: sequentially using each channel in the multi-target minimum variance distortionless response beamforming filter as a reference channel, and using the multi-target minimum variance distortionless response beamforming filter to eliminate noise signals in the dereverberation signal according to the reference channels to obtain a dereverberation and denoising signal.
[0151] In some embodiments of this application, the step of the fifth processing module 48 using a beamformer to eliminate noise signals in the dereverberation signal to obtain a dereverberation and noise-reduced signal includes: when there are multiple sound sources, determining the steering vector of each sound source; and using a beamformer to eliminate noise signals in the dereverberation signal based on the steering vector of each sound source to obtain a dereverberation and noise-reduced signal.
[0152] It should be noted that each module in the above-mentioned multi-source direction of arrival estimation device can be a program module (e.g., a set of program instructions to implement a certain function) or a hardware module. For the latter, it can be manifested in the following forms, but is not limited to them: each of the above modules is manifested as a processor, or the functions of each of the above modules are implemented by a processor.
[0153] According to an embodiment of this application, a non-volatile storage medium is provided, which stores a program. During program execution, the device containing the non-volatile storage medium performs the following sound source arrival estimation method: eliminating late reverberation signals in the observed signal to obtain an initial dereverberation signal, wherein the number of sound sources in the observed signal is one or more; non-uniformly dividing the initial dereverberation signal into multiple first sub-bandwidths, wherein each first sub-bandwidth contains multiple frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; determining the focusing covariance matrix corresponding to each first sub-bandwidth, and obtaining the sub-bandwidth focusing frequency corresponding to each first sub-bandwidth based on the focusing covariance matrix. The frequency spatial spectrum is used to determine the initial estimation result of the sound source arrival direction of the observed signal based on the sub-bandwidth focusing frequency spatial spectrum corresponding to the first sub-bandwidth. Based on the initial estimation result, the observed signal is iteratively processed using a beamformer and a dereverberation filter to obtain the target estimation result of the sound source arrival direction. In each iteration of the iterative processing, based on the sound source arrival direction estimation result obtained in the previous iteration, a dereverberation filter is used to eliminate the reverberation signal in the observed signal, and a beamformer is used to eliminate the noise signal in the observed signal. Based on the dereverberation and noise signal obtained after eliminating the reverberation and noise signals, the estimation result of the sound source arrival direction corresponding to the current iteration is determined.
[0154] According to an embodiment of this application, an electronic device is provided, including a memory and a processor. The processor is used to run a program stored in the memory, wherein the program executes the following multi-source direction-of-arrival estimation method: eliminating late reverberation signals in the observed signal to obtain an initial dereverberation signal, wherein the number of sound sources in the observed signal is one or more; non-uniformly dividing the initial dereverberation signal into multiple first sub-bandwidths, wherein each first sub-bandwidth contains multiple frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; determining the focusing covariance matrix corresponding to each first sub-bandwidth, and obtaining the sub-bandwidth focusing frequency space corresponding to each first sub-bandwidth based on the focusing covariance matrix. The initial estimation result of the sound source arrival direction of the observed signal is determined based on the spatial spectrum of the focused frequency of the first sub-bandwidth. Based on the initial estimation result, the observed signal is iteratively processed using a beamformer and a dereverberation filter to obtain the target estimation result of the sound source arrival direction. In each iteration of the iterative processing, based on the sound source arrival direction estimation result obtained in the previous iteration, a dereverberation filter is used to eliminate the reverberation signal in the observed signal, and a beamformer is used to eliminate the noise signal in the observed signal. Based on the dereverberation and noise signal obtained after eliminating the reverberation and noise signals, the estimation result of the sound source arrival direction corresponding to the current iteration is determined.
[0155] According to an embodiment of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the following multi-source direction-of-arrival estimation method: eliminating late reverberation signals in an observed signal to obtain an initial dereverberation signal, wherein the number of sound sources in the observed signal is one or more; non-uniformly dividing the initial dereverberation signal into multiple first sub-bandwidths, wherein each first sub-bandwidth contains multiple frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; determining the focusing covariance matrix corresponding to each first sub-bandwidth, and obtaining the sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the focusing covariance matrix; and based on the first... The initial estimation result of the sound source arrival direction of the observed signal is determined by the sub-bandwidth focusing frequency spatial spectrum corresponding to the sub-bandwidth. Based on the initial estimation result, the observed signal is iteratively processed using a beamformer and a dereverberation filter to obtain the target estimation result of the sound source arrival direction. In each iteration of the iterative processing, based on the sound source arrival direction estimation result obtained in the previous iteration, a dereverberation filter is used to eliminate the reverberation signal in the observed signal, and a beamformer is used to eliminate the noise signal in the observed signal. Based on the dereverberation and noise signal obtained after eliminating the reverberation and noise signals, the estimation result of the sound source arrival direction corresponding to the current iteration is determined.
[0156] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0157] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0158] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0159] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0160] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0161] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for estimating the direction of arrival of multiple sound sources, characterized in that, include: The late reverberation signal in the observed signal is removed to obtain the initial dereverberation signal, wherein the number of sound sources in the observed signal is one or more; The initial dereverberation signal is non-uniformly divided into multiple first sub-bandwidths, wherein each first sub-bandwidth contains multiple frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths. Determine the focusing covariance matrix corresponding to each of the first sub-bandwidths, and obtain the sub-bandwidth focusing frequency space spectrum corresponding to each of the first sub-bandwidths based on the focusing covariance matrix; The initial estimation result of the sound source arrival direction of the observed signal is determined based on the sub-bandwidth focusing frequency spatial spectrum corresponding to the first sub-bandwidth; Based on the initial estimation result, the observed signal is iteratively processed using a beamformer and a dereverberation filter to obtain the target estimation result of the sound source arrival direction. In each iteration of the iterative processing, based on the sound source arrival direction estimation result obtained in the previous iteration, the dereverberation filter is used to eliminate the reverberation signal in the observed signal, and the beamformer is used to eliminate the noise signal in the observed signal. Based on the dereverberation and noise signal obtained after eliminating the reverberation and noise signals, the estimation result of the sound source arrival direction corresponding to the current iteration is determined. The observed signal is iteratively processed using a beamformer and a dereverberation filter to obtain the target estimation results for the direction of arrival of the sound source, including: In each iteration, the beamformer is used to determine the time-varying variance based on the estimation result of the sound source arrival direction obtained in the previous iteration. The time-varying variance is used to reflect the energy change of the observed signal within a preset frequency range and time period. In the first iteration, the beamformer is used to determine the time-varying variance based on the initial estimation result. The reverberation filter is used to eliminate the reverberation signal in the observed signal based on the time-varying variance to obtain a reverberation-free signal. The beamformer is used to eliminate noise signals in the dereverberation signal to obtain the dereverberation-denoising signal, wherein the dereverberation-denoising signal includes a plurality of second sub-bandwidths, each of the second sub-bandwidths in the dereverberation-denoising signal contains a plurality of frequency band signals, and there is frequency overlap between adjacent second sub-bandwidths; Determine the focusing covariance matrix corresponding to each of the second sub-bandwidths, and obtain the sub-bandwidth focusing frequency spatial spectrum corresponding to each of the second sub-bandwidths based on the focusing covariance matrix; The estimation result of the sound source arrival direction in this iteration is determined based on the sub-bandwidth focusing frequency spatial spectrum corresponding to the second sub-bandwidth; The dereverberation filter includes a multi-channel linear prediction filter; using the dereverberation filter to eliminate the reverberation signal in the observed signal based on the time-varying variance, the resulting dereverberation signal includes: Each channel in the multi-channel linear prediction filter is used as a reference channel in sequence, and the reverberation signal in the observed signal is eliminated by the de-reverberation filter based on the reference channel and the time-varying variance to obtain a de-reverberation signal matrix. The de-reverberation signal matrix includes the de-reverberation sub-signals obtained after eliminating the reverberation signal with each channel as a reference channel. The dereverberation signal matrix is used as the dereverberation signal.
2. The multi-source arrival direction estimation method according to claim 1, characterized in that, Determining the focusing covariance matrix corresponding to each of the first sub-bandwidths includes: Determine the covariance matrix corresponding to each of the first sub-bandwidths, and perform phase transformation processing on the covariance matrix; For each of the first sub-bandwidths, the focusing covariance matrix corresponding to each of the first sub-bandwidths is determined based on the covariance matrix after phase transformation processing. Based on the focusing covariance matrix, obtaining the sub-bandwidth focusing frequency spatial spectrum corresponding to each of the first sub-bandwidths includes: Perform singular value decomposition on the focusing covariance matrix to obtain the noise subspace corresponding to each of the first sub-bandwidths; Based on the noise subspace, the subbandwidth focusing frequency space spectrum corresponding to each of the first subbandwidths is determined.
3. The multi-source arrival direction estimation method according to claim 1, characterized in that, The beamformer includes a multi-target minimum variance distortionless response beamforming filter; Using the beamformer to eliminate noise in the dederobatic signal to obtain the dederobatic and denoised signal includes: Each channel in the multi-objective minimum variance distortionless response beamforming filter is used as a reference channel in sequence, and the noise signal in the dereverberation signal is eliminated based on the reference channel using the multi-objective minimum variance distortionless response beamforming filter to obtain the dereverberation and noise-reduced signal.
4. The multi-source arrival direction estimation method according to claim 1, characterized in that, Using the beamformer to eliminate noise in the dederobatic signal to obtain the dederobatic and denoised signal includes: When there are multiple sound sources, determine the steering vector of each sound source; The beamformer is used to eliminate noise signals in the dereverberation signal based on the steering vector of each sound source, thereby obtaining the dereverberation and noise-reducing signal.
5. The multi-source arrival direction estimation method according to claim 1, characterized in that, Dividing the initial dereverberation signal non-uniformly into multiple first sub-bandwidths includes: Discard the portion of the initial dereverberation signal whose frequency is lower than the first preset frequency; For the first signal portion of the initial dereverberation signal whose frequency is not less than the first preset frequency and less than the second preset frequency, the first signal portion is divided into multiple first sub-bandwidths according to the first preset frequency spacing, wherein a first preset number of frequency bands are superimposed between the sub-bandwidths corresponding to two adjacent first sub-bandwidths in the first signal portion. For the second signal portion of the initial de-reverberation signal whose frequency is not less than the second preset frequency and less than the third preset frequency, the second signal portion is divided into multiple first sub-bandwidths according to the second preset frequency spacing, wherein the sub-bandwidths corresponding to two adjacent first sub-bandwidths in the second signal portion are superimposed with a second preset number of frequency bands. For the third signal portion of the initial dereverberation signal whose frequency is not less than the third preset frequency and less than the fourth preset frequency, the third signal portion is divided into multiple first sub-bandwidths according to the third preset frequency spacing, wherein the sub-bandwidths corresponding to two adjacent first sub-bandwidths in the third signal portion are superimposed with a third preset number of frequency bands. For the fourth signal portion of the initial dereverberation signal whose frequency is not less than a fourth preset frequency and less than a fifth preset frequency, the fourth signal portion is divided into multiple first sub-bandwidths according to the fourth preset frequency spacing. Among them, the sub-bandwidths corresponding to two adjacent first sub-bandwidths in the fourth signal portion are superimposed with a fourth preset number of frequency bands.
6. The method for estimating the direction of arrival of multiple sound sources according to claim 1, characterized in that, The initial estimation results for determining the sound source arrival direction of the observed signal based on the sub-bandwidth focusing frequency spatial spectrum include: Summing the spatial spectrum of each sub-bandwidth focusing frequency spectrum yields the overall significant peak value of each sub-bandwidth focusing frequency spectrum. The initial estimation result is determined based on the significant peak value.
7. The method for estimating the direction of arrival of multiple sound sources according to claim 1, characterized in that, The initial estimation results for determining the sound source arrival direction of the observed signal based on the sub-bandwidth focusing frequency spatial spectrum include: The significant peak values of the focusing frequency spatial spectrum of each of the sub-bandwidths are determined, along with the estimated source arrival directions corresponding to the significant peak values. Determine the set of significant peak estimates and determine the distribution information of the set of significant peak estimates. The initial estimation result is determined based on the distribution information.
8. A multi-source direction-of-arrival estimation device, characterized in that, include: The first processing module is used to eliminate late reverberation signals in the observed signal to obtain an initial dereverberation signal, wherein the number of sound sources in the observed signal is one or more; The second processing module is used to non-uniformly divide the initial dereverberation signal into multiple first sub-bandwidths, wherein each first sub-bandwidth contains multiple frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths. The third processing module is used to determine the focusing covariance matrix corresponding to each of the first sub-bandwidths, and to obtain the sub-bandwidth focusing frequency space spectrum corresponding to each of the first sub-bandwidths based on the focusing covariance matrix. The fourth processing module is used to determine the initial estimation result of the sound source arrival direction of the observed signal based on the sub-bandwidth focusing frequency spatial spectrum corresponding to the first sub-bandwidth; The fifth processing module is used to iteratively process the observed signal using a beamformer and a dereverberation filter based on the initial estimation result to obtain the target estimation result of the sound source arrival direction. In each iteration of the iterative processing, based on the sound source arrival direction estimation result obtained in the previous iteration, the dereverberation filter is used to eliminate the reverberation signal in the observed signal, and the beamformer is used to eliminate the noise signal in the observed signal. Based on the dereverberation and noise signal obtained after eliminating the reverberation and noise signals, the estimation result of the sound source arrival direction corresponding to the current iteration is determined. The dereverberation filter includes a multi-channel linear prediction filter. The fifth processing module is also used for: In each iteration, the beamformer is used to determine the time-varying variance based on the estimation result of the sound source arrival direction obtained in the previous iteration. The time-varying variance is used to reflect the energy change of the observed signal within a preset frequency range and time period. In the first iteration, the beamformer is used to determine the time-varying variance based on the initial estimation result. The reverberation filter is used to eliminate the reverberation signal in the observed signal based on the time-varying variance to obtain a reverberation-free signal. The beamformer is used to eliminate noise signals in the dereverberation signal to obtain the dereverberation-denoising signal, wherein the dereverberation-denoising signal includes a plurality of second sub-bandwidths, each of the second sub-bandwidths in the dereverberation-denoising signal contains a plurality of frequency band signals, and there is frequency overlap between adjacent second sub-bandwidths; Determine the focusing covariance matrix corresponding to each of the second sub-bandwidths, and obtain the sub-bandwidth focusing frequency spatial spectrum corresponding to each of the second sub-bandwidths based on the focusing covariance matrix; The estimation result of the sound source arrival direction in this iteration is determined based on the sub-bandwidth focusing frequency spatial spectrum corresponding to the second sub-bandwidth; The fifth processing module is also used for: Each channel in the multi-channel linear prediction filter is used as a reference channel in sequence, and the reverberation signal in the observed signal is eliminated by the de-reverberation filter based on the reference channel and the time-varying variance to obtain a de-reverberation signal matrix. The de-reverberation signal matrix includes the de-reverberation sub-signals obtained after eliminating the reverberation signal with each channel as a reference channel. The dereverberation signal matrix is used as the dereverberation signal.
9. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores a program, wherein when the program is executed, it controls the device containing the non-volatile storage medium to perform the multi-source arrival direction estimation method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, include: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, executes the multi-source direction-of-arrival estimation method according to any one of claims 1 to 7.
11. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the multi-source direction-of-arrival estimation method according to any one of claims 1 to 7.