Multi-sound-source arrival direction estimation method and device and nonvolatile storage medium

By eliminating the late reverberation signal in the observation signal and using the focused covariance matrix and beamformer for iterative processing, the problem of removing noise and reverberation components in the signal is solved, and the accuracy of the arrival direction estimation of multiple sound sources is improved.

CN120630101AActive Publication Date: 2025-09-12CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510749627.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The existing technology cannot effectively remove the noise and reverberation components in the signal, resulting in inaccurate estimation results of the direction of arrival of multiple sound sources.

Method used

By eliminating the late reverberation signal in the observation signal, non-uniformly dividing it into multiple sub-bandwidths, and iteratively processing it using the focused covariance matrix and beamformer, combined with the dereverberation filter and beamformer, the noise and reverberation signals are eliminated round by round to obtain the target estimation result of the sound source arrival direction.

Benefits of technology

The noise and reverberation components in the signal are effectively removed, and the accuracy of the estimation of the direction of arrival of multiple sound sources is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120630101A_ABST
    Figure CN120630101A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-sound-source arrival direction estimation method and device and a nonvolatile storage medium. The method comprises the following steps: eliminating a late reverberation signal in an observation signal to obtain an initial reverberation-removed signal; non-uniformly dividing the initial de-reverberation signal into a plurality of first sub-bandwidths; determining a focusing covariance matrix corresponding to each first sub-bandwidth, and obtaining a sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth according to the focusing covariance matrix; determining an initial estimation result of the sound source arrival direction of the observation signal according to the sub-bandwidth focusing frequency spatial spectrum corresponding to the first sub-bandwidth; and performing iterative processing on the observation signal by using a beam former and a dereverberation filter according to the initial estimation result to obtain a target estimation result of the direction of arrival of the sound source. The technical problem that the multi-sound-source arrival direction estimation result is inaccurate due to the fact that the noise component and the reverberation component of the signal cannot be efficiently removed in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio signal processing, and more specifically, to a method and device for estimating the direction of arrival of multiple sound sources, and a non-volatile storage medium. Background Art

[0002] In related technologies, when estimating the direction of arrival (DOA) of a sound source wave, it is impossible to effectively remove the noise and reverberation components in the signal, resulting in inaccurate estimation results.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] Embodiments of the present application provide a method, apparatus, and non-volatile storage medium for estimating the directions of arrival of multiple sound sources, to at least address the technical problem of inaccurate estimation results of the directions of arrival of multiple sound sources due to the inability to efficiently remove the noise and reverberation components of the signal in the related art.

[0005] According to one aspect of an embodiment of the present application, a method for estimating the direction of arrival of multiple sound sources is provided, comprising: eliminating a late reverberation signal in an observation signal to obtain an initial dereverberation signal, wherein the number of sound sources of the observation signal is one or more; non-uniformly dividing the initial dereverberation signal into a plurality of first sub-bandwidths, wherein each first sub-bandwidth contains a plurality of frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; determining a focusing covariance matrix corresponding to each first sub-bandwidth, and obtaining a sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the focusing covariance matrix; and obtaining a sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the sub-bandwidth focusing frequency spatial spectrum corresponding to the first sub-bandwidth. An initial estimation result of the direction of arrival of the sound source of the observation signal is determined by using a rate space spectrum; based on the initial estimation result, the observation signal is iteratively processed using a beamformer and a dereverberation filter to obtain a target estimation result of the direction of arrival of the sound source, wherein, in each round of iterative processing, based on the estimation result of the direction of arrival of the sound source obtained in the previous round of iterative processing, the reverberation signal in the observation signal is eliminated by using a dereverberation filter, and the noise signal in the observation signal is eliminated by using a beamformer, and the estimation result of the direction of arrival of the sound source corresponding to the current round of iteration is determined based on the dereverberation and denoising signal obtained after eliminating the reverberation signal and the noise signal.

[0006] Optionally, determining the focusing covariance matrix corresponding to each first sub-bandwidth includes: determining the covariance matrix corresponding to each first sub-bandwidth, and performing phase transformation processing on the covariance matrix; for each first sub-bandwidth, determining the focusing covariance matrix corresponding to each first sub-bandwidth based on the covariance matrix after the phase transformation processing; obtaining the sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the focusing covariance matrix includes: performing singular value decomposition on the focusing covariance matrix to obtain the noise subspace corresponding to each first sub-bandwidth; and determining the sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the noise subspace.

[0007] Optionally, iteratively processing the observation signal using a beamformer and a dereverberation filter to obtain a target estimation result of the direction of arrival of the sound source includes: in each round of iteration, using the beamformer to determine a time-varying variance based on the estimation result of the direction of arrival of the sound source obtained in the previous round of iteration, wherein the time-varying variance is used to reflect the energy change of the observation signal within a preset frequency range and time period, and in the first iteration, using the beamformer to determine the time-varying variance based on the initial estimation result; using the dereverberation filter to eliminate the reverberation signal in the observation signal based on the time-varying variance to obtain a dereverberation signal; using the beamformer to eliminate the noise signal in the dereverberation signal to obtain a dereverberation and denoised signal, wherein the dereverberation and denoised signal includes multiple second sub-bandwidths, each second sub-bandwidth in the dereverberation and denoised signal contains multiple frequency band signals, and there is frequency overlap between adjacent second sub-bandwidths; determining a focused covariance matrix corresponding to each second sub-bandwidth, and obtaining a sub-bandwidth focused frequency spatial spectrum corresponding to each second sub-bandwidth based on the focused covariance matrix; and determining the estimation result of the direction of arrival of the sound source for the current iteration based on the sub-bandwidth focused frequency spatial spectrum corresponding to the second sub-bandwidth.

[0008] Optionally, the dereverberation filter includes a multi-channel linear prediction filter; using the dereverberation filter to eliminate the reverberation signal in the observation signal based on the time-varying variance to obtain the dereverberation signal includes: sequentially using each channel in the multi-channel linear prediction filter as a reference channel, and using the filter to eliminate the reverberation signal in the observation signal based on the reference channel and the time-varying variance to obtain a dereverberation signal matrix, wherein the dereverberation signal family includes dereverberation sub-signals obtained after eliminating the reverberation signal with each channel as the reference channel; and using the dereverberation signal matrix as the dereverberation signal.

[0009] Optionally, the beamformer includes a multi-objective minimum variance distortionless response beamforming filter; using the beamformer to eliminate the noise signal in the dereverberation signal to obtain the dereverberation and denoised signal includes: sequentially using each channel in the multi-objective minimum variance distortionless response beamforming filter as a reference channel, and using the multi-objective minimum variance distortionless response beamforming filter to eliminate the noise signal in the dereverberation signal according to the reference channel to obtain the dereverberation and denoised signal.

[0010] Optionally, using a beamformer to eliminate noise signals in the dereverberation signal to obtain the dereverberation and denoised signal includes: when there are multiple sound sources, determining a steering vector for each sound source; and using a beamformer to eliminate noise signals in the dereverberation signal based on the steering vectors for each sound source to obtain the dereverberation and denoised signal.

[0011] Optionally, the non-uniform division of the initial dereverberation signal into a plurality of first sub-bandwidths includes: discarding a portion of the initial dereverberation signal with a frequency lower than a first preset frequency; for a first signal portion of the initial dereverberation signal with a frequency not less than the first preset frequency and less than a second preset frequency, dividing the first signal portion into a plurality of first sub-bandwidths according to a first preset frequency interval, wherein a first preset number of frequency bands are overlapped between sub-bandwidths corresponding to two first sub-bandwidths with adjacent frequencies in the first signal portion; for a second signal portion of the initial dereverberation signal with a frequency not less than the second preset frequency and less than a third preset frequency, dividing the second signal portion into a plurality of first sub-bandwidths according to a second preset frequency interval, wherein two first sub-bandwidths with adjacent frequencies in the second signal portion are overlapped with each other. A second preset number of frequency bands are overlapped between the sub-widths corresponding to the two first sub-bandwidths; for a third signal portion of the initial dereverberation signal having a frequency not less than a third preset frequency and less than a fourth preset frequency, the third signal portion is divided into a plurality of first sub-bandwidths according to a third preset frequency spacing, wherein a third preset number of frequency bands are overlapped between the sub-widths corresponding to two first sub-bandwidths having adjacent frequencies in the third signal portion; for a fourth signal portion of the initial dereverberation signal having a frequency not less than a fourth preset frequency and less than a fifth preset frequency, the fourth signal portion is divided into a plurality of first sub-bandwidths according to a fourth preset frequency spacing, wherein a fourth preset number of frequency bands are overlapped between the sub-widths corresponding to two first sub-bandwidths having adjacent frequencies in the fourth signal portion.

[0012] Optionally, determining an initial estimation result of the sound source arrival direction of the observation signal based on the sub-bandwidth focused frequency spatial spectrum includes: summing the sub-bandwidth focused frequency spatial spectra to obtain a significant peak of the overall sub-bandwidth focused frequency spatial spectrum; and determining the initial estimation result based on the significant peak.

[0013] Optionally, determining the initial estimation result of the sound source arrival direction of the observation signal based on the sub-bandwidth focused frequency spatial spectrum includes: determining the significant peaks of each sub-bandwidth focused frequency spatial spectrum, and the estimation result of the sound source arrival direction corresponding to the significant peak; determining a significant peak-estimation result set, and determining the distribution information of the significant peak-estimation result set; and determining the initial estimation result based on the distribution information.

[0014] According to another aspect of an embodiment of the present application, a multi-source arrival direction estimation device is also provided, comprising: a first processing module for eliminating late reverberation signals in an observation signal to obtain an initial dereverberation signal, wherein the number of sound sources of the observation signal is one or more; a second processing module for unevenly dividing the initial dereverberation signal into a plurality of first sub-bandwidths, wherein each first sub-bandwidth contains a plurality of frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; a third processing module for determining a focusing covariance matrix corresponding to each first sub-bandwidth, and obtaining a sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the focusing covariance matrix; a fourth processing module for determining a focusing covariance matrix corresponding to each first sub-bandwidth based on the focusing covariance matrix; and a fourth processing module for determining a focusing covariance matrix corresponding to each first sub-bandwidth based on the focusing covariance matrix. The sub-bandwidth focused frequency spatial spectrum corresponding to the first sub-bandwidth determines an initial estimation result of the sound source arrival direction of the observation signal; a fifth processing module is used to iteratively process the observation signal using a beamformer and a dereverberation filter based on the initial estimation result to obtain a target estimation result of the sound source arrival direction, wherein, in each round of iterative processing, based on the estimation result of the sound source arrival direction obtained in the previous round of iterative processing, a dereverberation filter is used to eliminate the reverberation signal in the observation signal, and a beamformer is used to eliminate the noise signal in the observation signal, and the estimation result of the sound source arrival direction corresponding to the current round of iteration is determined based on the dereverberation and denoising signal obtained after eliminating the reverberation signal and the noise signal.

[0015] According to an embodiment of the present application, a non-volatile storage medium is further provided, in which a program is stored. When the program is running, the device where the non-volatile storage medium is located is controlled to execute a method for estimating the direction of arrival of multiple sound sources.

[0016] According to an embodiment of the present application, an electronic device is further provided, including a memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the method for estimating directions of arrival of multiple sound sources is executed when the program is run.

[0017] According to an embodiment of the present application, a computer program product is further provided, including a computer program, which implements a method for estimating directions of arrival of multiple sound sources when executed by a processor.

[0018] In an embodiment of the present application, a late reverberation signal in an observation signal is eliminated to obtain an initial dereverberation signal, wherein the number of sound sources of the observation signal is one or more; the initial dereverberation signal is unevenly divided into a plurality of first sub-bandwidths, wherein each first sub-bandwidth contains a plurality of frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; a focusing covariance matrix corresponding to each first sub-bandwidth is determined, and a sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth is obtained based on the focusing covariance matrix; an initial estimation result of the sound source arrival direction of the observation signal is determined based on the sub-bandwidth focusing frequency spatial spectrum corresponding to the first sub-bandwidth; based on the initial estimation result, the observation signal is iteratively processed using a beamformer and a dereverberation filter to obtain a target estimation result of the sound source arrival direction, wherein, in the iterative processing In each round of iteration, based on the estimated result of the sound source arrival direction obtained in the previous round of iteration, a dereverberation filter is used to eliminate the reverberation signal in the observation signal, and a beamformer is used to eliminate the noise signal in the observation signal. Based on the dereverberation and denoised signal obtained after eliminating the reverberation signal and the noise signal, the estimated result of the sound source arrival direction corresponding to the current iteration is determined. By iteratively eliminating the noise signal and the reverberation signal in the observation signal using the filter and the beamformer, the purpose of efficiently removing the noise component and the reverberation component in the signal is achieved, thereby achieving the technical effect of improving the accuracy of the estimation result of the arrival direction of multiple sound sources, and further solving the technical problem of inaccurate estimation results of the arrival direction of multiple sound sources caused by the inability to efficiently remove the noise component and the reverberation component of the signal in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0020] Figure 1 1 is a schematic diagram of the structure of a computer terminal (mobile terminal) provided according to an embodiment of the present application;

[0021] Figure 2 1 is a flow chart of a method for estimating the direction of arrival of multiple sound sources provided in an embodiment of the present application;

[0022] Figure 3 This is a flowchart of a multi-sound source arrival direction estimation process provided in an embodiment of the present application;

[0023] Figure 4 3 is a schematic structural diagram of a device for estimating the direction of arrival of multiple sound sources provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0026] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:

[0027] Direction of Arrival (DOA) refers to the azimuth angle of the signal source incident on the array. It can effectively characterize the spatial position relationship between the signal source and the array and plays a very important role in signal processing fields such as voice, antennas, sonar, ultrasound, and radar.

[0028] Frequency Focusing: This is a processing method that uses a focusing matrix in the frequency domain to relocate signals from different frequency bands to the same frequency band. It is widely used in signal processing fields such as radar, sonar, antennas, and voice.

[0029] Convolutional Beamforming: This is a type of enhancement filter that uses historical and current frame information to jointly implement dereverberation and denoising tasks. It can achieve dereverberation tasks using multi-channel linear prediction and noise reduction using beamforming, and combines the two for alternating optimization through time-varying variance.

[0030] The goal of multi-source localization is to estimate the azimuth information of multiple sound sources in the acoustic space, providing spatial information for subsequent tasks such as sound source enhancement, sound source separation, and target sound source extraction. The objectives of multi-source localization can be mainly divided into three categories: sound source position estimation, time difference of arrival (TDOA) estimation, and direction of arrival (DOA) estimation. Among them, sound source position estimation is often difficult to use in practical tasks due to the high algorithm complexity or the difficulty in accurately estimating the azimuth distance. The accuracy of TDOA estimation is limited by the array element distance of the microphone array, which has led to the DOA estimation task being more widely studied and applied.

[0031] In an ideal, noiseless, and reverberation-free far-field environment, numerous methods have been shown to effectively estimate the Direction of Arrival (DOA) of sound sources. These include the Steered Response Power (SRP) method based on beamforming, the Multi-Signal Classification (MUSIC) method based on signal subspaces, and the SRP-based Phase Transform (SRP-PHAT) method. However, noise is unavoidable in real-world environments, and reverberation is also present in conventional indoor environments, greatly increasing the difficulty of DOA estimation. Furthermore, the random spatial distribution of sound sources in multi-source scenarios requires algorithms to have high resolution to cover situations where the sound sources have similar azimuth angles. Therefore, the problem of multi-source localization in complex scenarios is one of the most challenging problems in acoustics. To address this issue, a number of neural network-based methods have been proposed, such as the MUSIC algorithm based on deep neural networks (DNNs) and the SRP-PHAT algorithm based on conventional neural networks (CNNs). However, due to differences in array topology, environmental noise and reverberation, there is a lack of general sound source localization datasets with precise annotations, resulting in poor generalization performance of many neural network-based methods in practical tasks.

[0032] In order to solve the above problems, the embodiments of the present application provide a relevant solution. By combining the sound source localization method based on frequency focusing with the convolution beamforming for denoising and reverberation, a multi-source DOA estimation method combining frequency focusing and convolution beamforming is proposed, which has at least the following technical effects: First, according to the distribution of the frequency components of the speech signal, the speech bandwidth is unevenly divided into multiple sub-bandwidths, and each sub-bandwidth is overlapped to a certain extent, thereby improving the feature stability while improving the resolution of the frequency dimension; Second, the phase transformation method is used for the signal covariance matrix to obtain the frequency-focused spatial spectrum of the phase transformation, avoiding the influence of the unstable amplitude feature and improving the anti-noise and anti-reverberation ability of the frequency-focused spatial spectrum; Third, to address the problem of algorithm performance degradation in low signal-to-noise ratio and strong reverberation environment, a multi-source convolution beamformer is modeled based on multi-channel linear prediction (MCLP) and multi-objective minimum variance distortion-less response (MVDR) beamforming to suppress the influence of reverberation and noise, and to achieve joint optimization with the DOA based on frequency focusing in an alternating iterative manner. The following is a detailed description.

[0033] According to an embodiment of the present application, a method embodiment of a method for estimating the direction of arrival of multiple sound sources is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0034] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for estimating the direction of arrival of multiple sound sources. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more (illustrated as 102a, 102b, ..., 102n) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0035] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0036] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the multi-sound source arrival direction estimation method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned multi-sound source arrival direction estimation method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0037] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0038] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0039] In the above operating environment, the embodiment of the present application provides a method for estimating the direction of arrival of multiple sound sources, such as Figure 2 As shown, the method includes the following steps:

[0040] Step S202, eliminating the late reverberation signal in the observation signal to obtain an initial dereverberation signal, wherein the number of sound sources of the sound source wave is one or more;

[0041] In some embodiments of the present application, the above-mentioned observation signal may be an observation signal obtained by a microphone array. Optionally, the frequency domain model of the observation signal collected by a microphone array with M elements in a far-field environment may be expressed as follows:

[0042] X(k,l)=X dir (k,l)+X early (k,l)+X late (k,l)+N(k,l) (1)

[0043] in, To observe the signal, is the direct sound component, is the early reverberation component, is the late reverberation component, is the noise component; k and l represent the index of frequency and frame number respectively. represents the complex space. X in the above formula dir (k,l) can be expressed as follows:

[0044]

[0045] Among them, a(θ q ,k,l) ​​is the direct sound guidance vector, θ q is the incident angle of the qth sound source, S q (k, l) is the qth sound source signal, Q is the number of sound sources. The purpose of sound source DOA estimation is to calculate the incident angle θ of each sound source. q Give an estimate.

[0046] In some embodiments of the present application, a multi-channel linear prediction filter (MCLP) can be used to eliminate late reverberation, thereby reducing the impact of the reverberation on DOA estimation. Assuming that the dereverberation signal obeys a zero-mean complex Gaussian distribution, the MCLP filter G(k,l) can be solved based on the maximum likelihood objective function of the weighted prediction error (WPE) algorithm as follows:

[0047]

[0048] in, Lg is the length of the linear prediction filter (also called the order), is the vector of historical frame stacking, b is the linear prediction delay, σ(k,l) is the time-varying variance. The superscript H represents the conjugate transpose, and τ represents the historical frame index used for linear prediction filtering.

[0049] Based on the above MCLP filter G(k,l), the initial dereverberation signal after dereverberation can be The solution is as follows:

[0050]

[0051] In the above formula, G H represents the conjugate transpose of G.

[0052] Step S204: non-uniformly dividing the initial dereverberation signal into a plurality of first sub-bandwidths, wherein each first sub-bandwidth contains a plurality of frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths;

[0053] In the technical solution provided in step S204, the step of unevenly dividing the initial dereverberation signal into a plurality of first sub-bandwidths includes: discarding a portion of the initial dereverberation signal having a frequency lower than a first preset frequency; for a first signal portion of the initial dereverberation signal having a frequency not less than the first preset frequency and less than a second preset frequency, dividing the first signal portion into a plurality of first sub-bandwidths according to a first preset frequency interval, wherein a first preset number of frequency bands are overlapped between sub-bandwidths corresponding to two first sub-bandwidths having adjacent frequencies in the first signal portion; for a second signal portion of the initial dereverberation signal having a frequency not less than the second preset frequency and less than a third preset frequency, dividing the second signal portion into a plurality of first sub-bandwidths according to a second preset frequency interval, wherein the second signal portion a second preset number of frequency bands are overlapped between sub-generation widths corresponding to two first sub-bandwidths with adjacent frequencies in the portion; for a third signal portion whose frequency of the initial dereverberation signal is not less than a third preset frequency and less than a fourth preset frequency, the third signal portion is divided into a plurality of first sub-bandwidths according to a third preset frequency spacing, wherein a third preset number of frequency bands are overlapped between sub-generation widths corresponding to two first sub-bandwidths with adjacent frequencies in the third signal portion; for a fourth signal portion whose frequency of the initial dereverberation signal is not less than a fourth preset frequency and less than a fifth preset frequency, the fourth signal portion is divided into a plurality of first sub-bandwidths according to a fourth preset frequency spacing, wherein a fourth preset number of frequency bands are overlapped between sub-generation widths corresponding to two first sub-bandwidths with adjacent frequencies in the fourth signal portion.

[0054] Optionally, given the characteristics of speech signals, where low-frequency energy is concentrated and high-frequency energy gradually decays, a non-uniform sub-bandwidth division strategy is adopted to more effectively utilize the frequency band information of the speech signal. Specifically, the first resonance peak of a speech signal is generally between 300Hz and 1000Hz, the second resonance peak is generally between 800Hz and 2500Hz, the third resonance peak is generally between 2000Hz and 3500Hz, and higher-order resonance peaks are above 3500Hz. In addition, the fundamental pitch of speech is generally between 85Hz and 300Hz. Taking the above factors into consideration, the non-uniform flatband division strategy can be expressed in steps as follows:

[0055] (1) Discard the noise below 100 Hz;

[0056] (2) Between 100 Hz and 900 Hz, the signal is divided into sub-bandwidths at intervals of 100 Hz, with two frequency bands superimposed between the bandwidths;

[0057] (3) Between 900 Hz and 2300 Hz, the signal is divided into sub-bandwidths at intervals of 200 Hz, with three frequency bands superimposed between the bandwidths;

[0058] (4) Between 2300 Hz and 3500 Hz, the signal is divided into sub-bandwidths at intervals of 400 Hz, with four frequency bands superimposed between the bandwidths;

[0059] (5) Between 3500 Hz and 8000 Hz, the signal is divided into sub-bandwidths at intervals of 800 Hz, with six frequency bands overlapped between the bandwidths;

[0060] Assuming that the speech signal sampling rate is 16 kHz, the short-time Fourier transform (STFT) frame length and window length are both 512 samples, and the frame shift is 50%, a total of 47 sub-bandwidths are divided according to the above method.

[0061] Step S206, determining a focusing covariance matrix corresponding to each first sub-bandwidth, and obtaining a sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth according to the focusing covariance matrix;

[0062] In the technical solution provided in step S206, if Figure 3As shown, the step of determining the focusing covariance matrix corresponding to each first sub-bandwidth includes: determining the covariance matrix corresponding to each first sub-bandwidth, and performing phase transformation processing on the covariance matrix; for each first sub-bandwidth, determining the focusing covariance matrix corresponding to each first sub-bandwidth based on the covariance matrix after the phase transformation processing; the step of obtaining the sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the focusing covariance matrix includes: performing singular value decomposition on the focusing covariance matrix to obtain the noise subspace corresponding to each first sub-bandwidth; and determining the sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the noise subspace.

[0063] In addition, from Figure 3 It can be seen that when determining the initial estimation result, it is necessary to perform steps such as calculating the covariance matrix, covariance matrix phase transformation, determining the sub-bandwidth focusing covariance matrix, determining the sub-bandwidth focusing covariance matrix, calculating the noise subspace, determining the sub-bandwidth focusing frequency space spectrum, and determining the initial estimation result.

[0064] Optionally, when computing the covariance matrix, assume is the signal of the jth frequency band in the i-th sub-bandwidth, then its covariance matrix can be calculated as follows:

[0065]

[0066] in, is the covariance matrix, τ is the snapshot index, Γ is the snapshot number, and the superscript H represents the conjugate transpose of the matrix, and Γ-1 is the maximum snapshot number index.

[0067] Optionally, considering the instability of amplitude information caused by various factors such as noise, signal absorption and attenuation in the far-field environment, a phase transformation method is used to avoid the influence of amplitude factors. The phase transformation of formula (5) is shown as follows:

[0068]

[0069] In the above formula, |.| represents the membrane value operation.

[0070] Optionally, when calculating the focusing covariance matrix within the sub-bandwidth, the focusing matrix of the j-th frequency band signal in the i-th sub-bandwidth can be set to C i,j (f k ,l), then the focusing covariance matrix of the frequency band can be expressed as:

[0071]

[0072] Among them, f k To focus the reference frequency, the focusing matrix C i,j (f k,l) can be solved by the following method:

[0073]

[0074] Among them, U s (f k ,l) is the signal subspace of the covariance matrix of the focused reference band, U s (k, l) is the signal subspace of the current band covariance matrix. The signal subspace can be obtained by the eigenvector / singular value vector corresponding to the significantly large eigenvalue / singular value. C in formula (8) i,j (f k ,l) can be solved by the bilateral transformation method.

[0075] Optionally, after obtaining the focusing matrix, the focusing covariance matrix shown in formula (7) can be further obtained. The focusing covariance matrices of multiple frequency bands within the sub-bandwidth can then be averaged to obtain the sub-bandwidth focusing covariance matrix, as shown in the following formula:

[0076]

[0077] Among them, J i is the total number of frequency bands in the i-th sub-bandwidth.

[0078] Optionally, when calculating the noise subspace, the noise subspace of the corresponding frequency band can be obtained by performing singular value decomposition on the focused covariance matrix of the corresponding frequency band, as shown in the following formula:

[0079]

[0080] Where P is the number of significantly large eigenvalues, V i (f k ,l),Σ i (f k ,l) and U i (f k ,l) are the left singular vector, singular value diagonal matrix and right singular matrix respectively. is the noise subspace, The singular value vector corresponding to the P+1th singular value is arranged from the largest direction.

[0081] Optionally, when calculating the sub-bandwidth focused frequency spatial spectrum, the focused frequency spatial spectrum of the i-th sub-bandwidth can be obtained based on formula (11):

[0082]

[0083] in, for The direction of the steering vector, is the steering vector search angle. Repeating this process for I sub-bandwidths yields a wideband frequency-focused spatial spectrum. The wideband frequency-focused spatial spectrum can be viewed as a concatenation of the narrowband frequency-focused spatial spectra corresponding to each of the I sub-bandwidths.

[0084] Step S208, determining an initial estimation result of the direction of arrival of the sound source of the observation signal according to the sub-bandwidth focused frequency spatial spectrum corresponding to the first sub-bandwidth;

[0085] There are generally two methods for estimating DOA based on broadband frequency focused spatial spectrum. The first method is to sum the narrowband frequency focused spatial spectrum of the broadband frequency focused spatial spectrum, and then estimate the DOA information of the sound source from the significant peak of the narrowband spatial spectrum obtained by the summation; the second method is to construct a peak set based on the significant peak of each narrowband frequency focused spatial spectrum, and then use a statistical method to estimate the final DOA information for the peak set. When using method 2, if Figure 3 As shown, the Gaussian distribution or von Mises distribution can be used to fit the histogram statistics, and the estimated DOA information θ=[θ1,θ2,...,θ Q ].

[0086] In some embodiments of the present application, the steps of using method 1 to determine the initial estimation result of the sound source arrival direction of the observation signal based on the sub-bandwidth focused frequency spatial spectrum include: summing the sub-bandwidth focused frequency spatial spectra to obtain the significant peak of the overall sub-bandwidth focused frequency spatial spectrum; and determining the initial estimation result based on the significant peak.

[0087] The advantage of using method 1 is its simplicity, but the summation process can lead to a decrease in the resolution of the spatial spectrum peaks, which can cause large errors in the DOA estimation results and even prevent the effective detection of some weaker sound sources. Method 2, on the other hand, is more computationally complex, but it can filter out the effects of many clutter signals and does not suffer from the problem of a decrease in the resolution of the spatial spectrum peaks due to summation.

[0088] In some embodiments of the present application, the steps of using method 2 to determine the initial estimation result of the sound source arrival direction of the observation signal based on the sub-bandwidth focused frequency spatial spectrum include: determining the significant peaks of each sub-bandwidth focused frequency spatial spectrum and the estimation results of the sound source arrival direction corresponding to the significant peaks; determining a set of significant peak-estimation results and determining the distribution information of the significant peak-estimation result set; and determining the initial estimation result based on the distribution information. The specific implementation details include the following steps:

[0089] 1. Obtain the corresponding DOA value based on the significant peak of the spatial spectrum of each narrowband frequency focus, and construct the peak DOA set from these DOA values;

[0090] 2. According to the peak DOA set, select the appropriate quantization accuracy (it is recommended to be between 1 and 2), and construct a statistical histogram, such as Figure 1 As shown;

[0091] 3. Use von Mises distribution to fit the histogram and obtain the DOA distribution curve;

[0092] 4. According to the DOA distribution curve, obtain the DOA interval corresponding to the significant peak value, and take the median of the interval as the estimated DOA information.

[0093] Step S210: Based on the initial estimation result, the observation signal is iteratively processed using a beamformer and a dereverberation filter to obtain a target estimation result of the direction of arrival of the sound source. In each round of iterative processing, based on the estimation result of the direction of arrival of the sound source obtained in the previous round of iterative processing, the reverberation signal in the observation signal is eliminated using a dereverberation filter, and the noise signal in the observation signal is eliminated using a beamformer. The estimation result of the direction of arrival of the sound source corresponding to the current iteration is determined based on the dereverberation and denoising signal obtained after eliminating the reverberation signal and the noise signal.

[0094] In some embodiments of the present application, Figure 3 As shown, the step of iteratively processing the observation signal using a beamformer and a dereverberation filter to obtain a target estimation result of the sound source arrival direction includes: in each round of iteration, using the beamformer to determine a time-varying variance based on the sound source arrival direction estimation result obtained in the previous round of iteration, wherein the time-varying variance is used to reflect the energy change of the observation signal within a preset frequency range and time period. In the first iteration, the beamformer is used to determine the time-varying variance based on the initial estimation result; using the dereverberation filter to eliminate the reverberation signal in the observation signal based on the time-varying variance to obtain a dereverberation signal; using the beamformer to eliminate the noise signal in the dereverberation signal to obtain a dereverberation and denoised signal, wherein the dereverberation and denoised signal includes multiple second sub-bandwidths, each second sub-bandwidth in the dereverberation and denoised signal contains multiple frequency band signals, and there is frequency overlap between adjacent second sub-bandwidths; determining a focused covariance matrix corresponding to each second sub-bandwidth, and obtaining a sub-bandwidth focused frequency spatial spectrum corresponding to each second sub-bandwidth based on the focused covariance matrix; and determining the sound source arrival direction estimation result of the current iteration based on the sub-bandwidth focused frequency spatial spectrum corresponding to the second sub-bandwidth. in Figure 3 The blue dotted line represents the transmission process of the initial estimation result, and the red solid line is the data transmission direction of the iterative process.

[0095] It should be noted that although the combined formulas (4) and (12) reduce the impact of reverberation on the sound source position, this process does not take into account the impact of noise. In order to further eliminate the noise components in the observation signal, a joint beamformer can be used to suppress the noise components. The noise reduction process requires that the multi-source signal be kept distortion-free, which makes the general beamformer unable to perform this task. To this end, a multi-target minimum variance distortion-free response (Multi-target MVDR, MT-MVDR) beamformer for multi-source denoising is modeled in the embodiment of the present application to ensure that the direct sound DOA signal of multiple sound sources is distortion-free. For ease of description, the dual-source scenario is used as an example in the subsequent process. The MT-MVDR beamformer can be modeled as:

[0096]

[0097] Where W(k,l) is the filter weight vector of beamforming, a(θ1,k,l) ​​and a(θ2,k,l) ​​are the steering vectors of sound source 1 and sound source 2, and θ1 and θ2 are the DOA information of sound source 1 and sound source 2 respectively. H Represents conjugate transpose. R zz (k,l) is the covariance matrix of the dereverberation signal, which can be expressed as follows:

[0098]

[0099] Based on the Lagrange multiplier method, the solution of equation (13) is:

[0100]

[0101] in,

[0102]

[0103] Based on the dereverberation signal of formula (4) and the beamforming denoising method of formula (15), the dereverberation and denoising signal can be obtained:

[0104] Y(k,l)=W H (k,l)Z(k,l) (20)

[0105] In some embodiments of the present application, the dereverberation filter includes a multi-channel linear prediction filter; using the dereverberation filter to eliminate the reverberation signal in the observation signal based on the time-varying variance to obtain the dereverberation signal includes: sequentially using each channel in the multi-channel linear prediction filter as a reference channel, and using the filter to eliminate the reverberation signal in the observation signal based on the reference channel and the time-varying variance to obtain a dereverberation signal matrix, wherein the dereverberation signal family includes dereverberation sub-signals obtained after eliminating the reverberation signal with each channel as the reference channel; and using the dereverberation signal matrix as the dereverberation signal.

[0106] In some embodiments of the present application, the beamformer includes a multi-objective minimum variance distortionless response beamforming filter; using the beamformer to eliminate the noise signal in the dereverberation signal to obtain the dereverberation and denoised signal includes: sequentially using each channel in the multi-objective minimum variance distortionless response beamforming filter as a reference channel, and using the multi-objective minimum variance distortionless response beamforming filter to eliminate the noise signal in the dereverberation signal according to the reference channel to obtain the dereverberation and denoised signal.

[0107] In some embodiments of the present application, by selecting different reference channels for a(θ1, k, l) and a(θ2, k, l), a multi-channel denoised and reverberated signal can be obtained, so that the signal can be subsequently Figure 3 As shown in Figure 1, a frequency-focused DOA estimation method based on phase transformation is repeatedly implemented to achieve more robust DOA estimation. However, while a cascaded WPE dereverberation filter (MCLP filter) and MT-MVDR beamformer denoising system can suppress the effects of reverberation and noise, the performance of the cascaded system is often suboptimal. Therefore, a convolutional beamformer is constructed based on the WPE dereverberation filter and MT-MVDR beamformer to achieve better performance.

[0108] As an optional implementation, in the convolutional beamformer, the WPE dereverberation filter and the MT-MVDR beamformer are often linked through time-varying variance, thereby achieving an alternating optimization process. First, the time-varying variance needs to be initialized. Let the lth channel be the reference channel. In order to make the algorithm converge faster, the WPE dereverberation and multi-target MVDR beamforming denoising system are connected in series to estimate the initial time-varying variance σ 1,0 (k, l), according to the result of formula (20), it can be expressed as follows:

[0109] σ 1,0 (k,l)≈|Y1(k,l)| 2 =|W H (k,l)Z1(k,l)| 2 (twenty one)

[0110] Among them, σ 1,0 The 1 in (k, l) represents the lth channel, 0 represents the initial time-varying variance, and Z1(k, l) represents the dereverberation signal with the lth channel as the reference channel. By selecting different reference channels, we can obtain σ0(k, l) = [σ 1,0 (k,l),σ 2,0 (k,l),...,σ M,0 (k,l)].

[0111] Optional, such as Figure 3 As shown, based on the time-varying variance of formula (21), the WPE dereverberation filter and MT-MVDR beamformer shown in formulas (3) and (15) can be jointly iterated. The specific iterative process is as follows:

[0112] (a) Clause WPE dereverberation in iterations

[0113] No. The dereverberation filter with the mth channel as the reference channel in the iteration can be expressed as follows:

[0114]

[0115] in, and Respectively The dereverberation filter and time-varying variance with the mth channel as the reference channel in the iteration.

[0116] Based on formula (22), the dereverberation signal can be calculated as follows:

[0117]

[0118] By selecting different reference channels, you can get the The dereverberation signal matrix of the iteration

[0119] (b) Multi-target beamforming denoising in iterations

[0120] No. The denoising beamformer with the mth channel as the reference channel in the iteration can be expressed as follows:

[0121]

[0122] Based on equations (23) and (24), the denoised signal can be calculated as follows:

[0123]

[0124] The corresponding time-varying variance can be updated as follows:

[0125]

[0126] By selecting different reference channels, you can get the The time-varying variance matrix of the iteration It should be noted that each channel will be selected as a reference channel. And the reference channel for the denoising process and the reference channel for the dereverberation process are the same reference channel.

[0127] Then, based on formula (26), the multi-channel dereverberation and denoising signal can be expressed as follows:

[0128]

[0129] Then you can Replace the initial dereverberation signal Z(k,l) and re-execute Figure 3 The frequency-focused DOA estimation process based on phase transformation is performed to obtain the estimation result of this iterative process.

[0130] As an optional implementation, the step of using a beamformer to eliminate noise signals in the dereverberation signal to obtain the dereverberation and denoised signal includes: when there are multiple sound sources, determining a steering vector for each sound source; and using a beamformer to eliminate noise signals in the dereverberation signal based on the steering vectors for each sound source to obtain the dereverberation and denoised signal.

[0131] It can be seen that Figure 3 As shown in Figure 1, each round of iteration includes three core steps, namely, frequency-focused DOA estimation based on non-uniform frequency band division, MCLP-based dereverberation, and multi-target MVDR beamforming-based denoising. The steps of each round of iteration are as follows:

[0132] (1) First, the initial DOA information θ0 is fed into the multi-target MVDR beamforming denoising module to obtain the updated time-varying variance

[0133] (2) Secondly, Update to MCLP filter, optimize the convolution beamforming enhancement system combining MCLP filtering and multi-target MVDR beamforming, and obtain the optimized enhanced signal

[0134] (3) Then, The DOA information is re-estimated by sending it to the frequency-focused DOA estimation module based on non-uniform frequency band division.

[0135] (4) Next, The result is sent to the multi-target MVDR beamforming denoising module for a new round of iteration.

[0136] When the estimated DOA converges or reaches the maximum number of iterations, the final estimated DOA information of multiple sound sources can be obtained.

[0137] An initial dereverberation signal is obtained by eliminating the late reverberation signal in the observation signal, wherein the number of sound sources of the sound source wave is one or more; the initial dereverberation signal is divided into multiple sub-bandwidth signals, wherein each sub-bandwidth signal corresponds to a sub-bandwidth, and there are overlapping frequencies between different sub-bandwidths; the covariance matrix corresponding to the multiple sub-bandwidth signals is determined, and the sub-bandwidth focused frequency spatial spectrum corresponding to the multiple sub-bandwidth signals is determined based on the covariance matrix; the initial estimation result of the sound source arrival direction of the observation signal is determined based on the sub-bandwidth focused frequency spatial spectrum; the initial estimation result and the observation signal are iteratively processed using a beamformer and a filter to obtain a target estimation result of the sound source arrival direction, wherein, in each round of iterative processing, based on the previous round The estimated result of the sound source arrival direction obtained in the iterative process uses a filter to eliminate the reverberation signal in the observation signal, and uses a beamformer to eliminate the noise signal in the observation signal, and determines the estimated result of the sound source arrival direction of this round of iteration based on the dedeverberation and denoised signal obtained after eliminating the reverberation signal and the noise signal. By using the filter and the beamformer to iteratively eliminate the noise signal and the reverberation signal in the observation signal, the purpose of efficiently removing the noise component and the reverberation component in the signal is achieved, thereby achieving the technical effect of improving the accuracy of the estimation result of the arrival direction of multiple sound sources, and thus solving the technical problem of inaccurate estimation results of the arrival direction of multiple sound sources caused by the inability to efficiently remove the noise component and the reverberation component of the signal in the related art.

[0138] The present invention provides a device for estimating the direction of arrival of multiple sound sources. Figure 4 It is a structural diagram of the device. Figure 4As can be seen from the figure, the device includes: a first processing module 40, which is used to eliminate the late reverberation signal in the observation signal to obtain an initial dereverberation signal, wherein the number of sound sources of the observation signal is one or more; a second processing module 42, which is used to non-uniformly divide the initial dereverberation signal into a plurality of first sub-bandwidths, wherein each first sub-bandwidth contains a plurality of frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; a third processing module 44, which is used to determine the focusing covariance matrix corresponding to each first sub-bandwidth, and obtain the sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth according to the focusing covariance matrix; a fourth processing module 46, which is used to obtain the sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth according to the sub-bandwidth corresponding to the first sub-bandwidth. The bandwidth-focused frequency spatial spectrum determines an initial estimation result of the sound source arrival direction of the observation signal; a fifth processing module 48 is used to iteratively process the observation signal using a beamformer and a dereverberation filter based on the initial estimation result to obtain a target estimation result of the sound source arrival direction, wherein, in each round of iterative processing, based on the estimation result of the sound source arrival direction obtained in the previous round of iterative processing, a dereverberation filter is used to eliminate the reverberation signal in the observation signal, and a beamformer is used to eliminate the noise signal in the observation signal, and the estimation result of the sound source arrival direction corresponding to the current round of iteration is determined based on the dereverberation and denoising signal obtained after eliminating the reverberation signal and the noise signal.

[0139] In some embodiments of the present application, the second processing module 42 non-uniformly divides the initial dereverberation signal into a plurality of first sub-bandwidths, including: discarding a portion of the initial dereverberation signal with a frequency lower than a first preset frequency; for a first signal portion of the initial dereverberation signal with a frequency not less than the first preset frequency and less than a second preset frequency, dividing the first signal portion into a plurality of first sub-bandwidths according to a first preset frequency interval, wherein a first preset number of frequency bands are overlapped between sub-bandwidths corresponding to two first sub-bandwidths with adjacent frequencies in the first signal portion; for a second signal portion of the initial dereverberation signal with a frequency not less than the second preset frequency and less than a third preset frequency, dividing the second signal portion into a plurality of first sub-bandwidths according to a second preset frequency interval, wherein the second signal portion a second preset number of frequency bands are overlapped between sub-generation widths corresponding to two first sub-bandwidths with adjacent frequencies in the portion; for a third signal portion whose frequency of the initial dereverberation signal is not less than a third preset frequency and less than a fourth preset frequency, the third signal portion is divided into a plurality of first sub-bandwidths according to a third preset frequency spacing, wherein a third preset number of frequency bands are overlapped between sub-generation widths corresponding to two first sub-bandwidths with adjacent frequencies in the third signal portion; for a fourth signal portion whose frequency of the initial dereverberation signal is not less than a fourth preset frequency and less than a fifth preset frequency, the fourth signal portion is divided into a plurality of first sub-bandwidths according to a fourth preset frequency spacing, wherein a fourth preset number of frequency bands are overlapped between sub-generation widths corresponding to two first sub-bandwidths with adjacent frequencies in the fourth signal portion.

[0140] In some embodiments of the present application, the step of determining the focused covariance matrix corresponding to each first sub-bandwidth by the third processing module 44 includes: determining the covariance matrix corresponding to each first sub-bandwidth and performing phase transformation processing on the covariance matrix; for each first sub-bandwidth, determining the focused covariance matrix corresponding to each first sub-bandwidth based on the covariance matrix after the phase transformation processing. The step of obtaining the sub-bandwidth focused frequency spatial spectrum corresponding to each first sub-bandwidth based on the focused covariance matrix includes: performing singular value decomposition on the focused covariance matrix to obtain the noise subspace corresponding to each first sub-bandwidth; and determining the sub-bandwidth focused frequency spatial spectrum corresponding to each first sub-bandwidth based on the noise subspace.

[0141] In some embodiments of the present application, the step of the fourth processing module 46 determining the initial estimation result of the sound source arrival direction of the observation signal based on the sub-bandwidth focused frequency spatial spectrum includes: summing the sub-bandwidth focused frequency spatial spectra to obtain the significant peak of the overall sub-bandwidth focused frequency spatial spectrum; and determining the initial estimation result based on the significant peak.

[0142] In some embodiments of the present application, the step of the fourth processing module 46 determining the initial estimation result of the sound source arrival direction of the observation signal based on the sub-bandwidth focused frequency spatial spectrum includes: determining the significant peaks of each sub-bandwidth focused frequency spatial spectrum, and the estimation result of the sound source arrival direction corresponding to the significant peak; determining the significant peak-estimation result set, and determining the distribution information of the significant peak-estimation result set; determining the initial estimation result based on the distribution information.

[0143] In some embodiments of the present application, the fifth processing module 48 uses a beamformer and a dereverberation filter to iteratively process the observation signal, and the step of obtaining a target estimation result of the sound source arrival direction includes: in each round of iteration, using a beamformer to determine a time-varying variance based on the estimation result of the sound source arrival direction obtained in the previous round of iteration, wherein the time-varying variance is used to reflect the energy change of the observation signal within a preset frequency range and time period, and in the first iteration, using a beamformer to determine the time-varying variance based on the initial estimation result; using a dereverberation filter to eliminate the reverberation signal in the observation signal based on the time-varying variance. signal to obtain a dereverberation signal; use a beamformer to eliminate the noise signal in the dereverberation signal to obtain a dereverberation and denoised signal, wherein the dereverberation and denoised signal includes multiple second sub-bandwidths, each second sub-bandwidth in the dereverberation and denoised signal contains multiple frequency band signals, and there is frequency overlap between adjacent second sub-bandwidths; determine a focused covariance matrix corresponding to each second sub-bandwidth, and obtain a sub-bandwidth focused frequency spatial spectrum corresponding to each second sub-bandwidth based on the focused covariance matrix; and determine an estimation result of the sound source arrival direction of this iteration based on the sub-bandwidth focused frequency spatial spectrum corresponding to the second sub-bandwidth.

[0144] In some embodiments of the present application, the dereverberation filter includes a multi-channel linear prediction filter; the fifth processing module 48 uses the dereverberation filter to eliminate the reverberation signal in the observation signal based on the time-varying variance, and the step of obtaining the dereverberation signal includes: taking each channel in the multi-channel linear prediction filter as a reference channel in turn, and using the filter to eliminate the reverberation signal in the observation signal based on the reference channel and the time-varying variance to obtain a dereverberation signal matrix, wherein the dereverberation signal family includes dereverberation sub-signals obtained after eliminating the reverberation signal with each channel as the reference channel; and using the dereverberation signal matrix as the dereverberation signal.

[0145] In some embodiments of the present application, the beamformer includes a multi-objective minimum variance distortionless response beamforming filter; the fifth processing module 48 uses the beamformer to eliminate the noise signal in the dereverberation signal, and the step of obtaining the dereverberation and denoised signal includes: sequentially using each channel in the multi-objective minimum variance distortionless response beamforming filter as a reference channel, and using the multi-objective minimum variance distortionless response beamforming filter to eliminate the noise signal in the dereverberation signal according to the reference channel to obtain the dereverberation and denoised signal.

[0146] In some embodiments of the present application, the fifth processing module 48 uses a beamformer to eliminate the noise signal in the dereverberation signal, and the steps of obtaining the dereverberation and denoised signal include: when there are multiple sound sources, determining the steering vector of each sound source; and using the beamformer to eliminate the noise signal in the dereverberation signal according to the steering vector of each sound source to obtain the dereverberation and denoised signal.

[0147] It should be noted that the various modules in the above-mentioned multi-sound source arrival direction estimation device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following form, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.

[0148] According to an embodiment of the present application, a non-volatile storage medium is provided, in which a program is stored, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to perform the following sound source wave arrival method estimation method: eliminating the late reverberation signal in the observation signal to obtain an initial dereverberation signal, wherein the number of sound sources of the observation signal is one or more; non-uniformly dividing the initial dereverberation signal into a plurality of first sub-bandwidths, wherein each first sub-bandwidth contains a plurality of frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; determining the focusing covariance matrix corresponding to each first sub-bandwidth, and obtaining the sub-bandwidth focusing frequency corresponding to each first sub-bandwidth according to the focusing covariance matrix. rate spatial spectrum; determining an initial estimation result of the direction of arrival of the sound source of the observation signal based on the sub-bandwidth focused frequency spatial spectrum corresponding to the first sub-bandwidth; iteratively processing the observation signal using a beamformer and a dereverberation filter based on the initial estimation result to obtain a target estimation result of the direction of arrival of the sound source, wherein, in each round of iterative processing, based on the estimation result of the direction of arrival of the sound source obtained in the previous round of iterative processing, a reverberation signal in the observation signal is eliminated using a dereverberation filter, and a noise signal in the observation signal is eliminated using a beamformer, and an estimation result of the direction of arrival of the sound source corresponding to the current iteration is determined based on a dereverberation and denoised signal obtained after eliminating the reverberation signal and the noise signal.

[0149] According to an embodiment of the present application, an electronic device is provided, comprising a memory and a processor, the processor being configured to run a program stored in the memory, wherein the program executes the following method for estimating the direction of arrival of multiple sound sources when running: eliminating a late reverberation signal in an observation signal to obtain an initial dereverberation signal, wherein the number of sound sources of the observation signal is one or more; non-uniformly dividing the initial dereverberation signal into a plurality of first sub-bandwidths, wherein each first sub-bandwidth contains a plurality of frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; determining a focusing covariance matrix corresponding to each first sub-bandwidth, and obtaining a sub-bandwidth focusing frequency space corresponding to each first sub-bandwidth based on the focusing covariance matrix; spectrum; determining an initial estimation result of the direction of arrival of the sound source of the observation signal based on the sub-bandwidth focused frequency spatial spectrum corresponding to the first sub-bandwidth; iteratively processing the observation signal using a beamformer and a dereverberation filter based on the initial estimation result to obtain a target estimation result of the direction of arrival of the sound source, wherein, in each round of iterative processing, based on the estimation result of the direction of arrival of the sound source obtained in the previous round of iterative processing, a reverberation signal in the observation signal is eliminated using a dereverberation filter, and a noise signal in the observation signal is eliminated using a beamformer, and an estimation result of the direction of arrival of the sound source corresponding to the current iteration is determined based on a dereverberation and denoised signal obtained after eliminating the reverberation signal and the noise signal.

[0150] According to an embodiment of the present application, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements the following method for estimating the direction of arrival of multiple sound sources: eliminating a late reverberation signal in an observation signal to obtain an initial dereverberation signal, wherein the number of sound sources of the observation signal is one or more; non-uniformly dividing the initial dereverberation signal into a plurality of first sub-bandwidths, wherein each first sub-bandwidth contains a plurality of frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; determining a focusing covariance matrix corresponding to each first sub-bandwidth, and obtaining a sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the focusing covariance matrix; and obtaining a sub-bandwidth focusing frequency spatial spectrum corresponding to each first sub-bandwidth based on the first sub-bandwidth. An initial estimation result of the direction of arrival of the sound source of the observation signal is determined using a sub-bandwidth focused frequency spatial spectrum corresponding to the sub-bandwidth; based on the initial estimation result, the observation signal is iteratively processed using a beamformer and a dereverberation filter to obtain a target estimation result of the direction of arrival of the sound source, wherein, in each round of iterative processing, based on the estimation result of the direction of arrival of the sound source obtained in the previous round of iterative processing, a dereverberation filter is used to eliminate the reverberation signal in the observation signal, and a beamformer is used to eliminate the noise signal in the observation signal, and based on the dereverberation and denoising signal obtained after eliminating the reverberation signal and the noise signal, the estimation result of the direction of arrival of the sound source corresponding to the current round of iteration is determined.

[0151] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0152] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0153] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0154] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0155] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0156] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for estimating the direction of arrival of multiple sound sources, characterized in that: include: Eliminating a late reverberation signal in an observation signal to obtain an initial dereverberation signal, wherein the observation signal has one or more sound sources; Non-uniformly dividing the initial dereverberation signal into a plurality of first sub-bandwidths, wherein each of the first sub-bandwidths contains a plurality of frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; Determine a focusing covariance matrix corresponding to each of the first sub-bandwidths, and obtain a sub-bandwidth focused frequency spatial spectrum corresponding to each of the first sub-bandwidths according to the focusing covariance matrix; Determining an initial estimation result of the direction of arrival of the sound source of the observation signal according to the sub-bandwidth focused frequency spatial spectrum corresponding to the first sub-bandwidth; Based on the initial estimation result, the observation signal is iteratively processed using a beamformer and a dereverberation filter to obtain a target estimation result of the direction of arrival of the sound source. In each round of iterative processing, based on the estimation result of the direction of arrival of the sound source obtained in the previous round of iterative processing, the dereverberation filter is used to eliminate the reverberation signal in the observation signal, and the beamformer is used to eliminate the noise signal in the observation signal. The estimation result of the direction of arrival of the sound source corresponding to the current round of iteration is determined based on the dereverberation and denoising signal obtained after eliminating the reverberation signal and the noise signal.

2. The method for estimating the direction of arrival of multiple sound sources according to claim 1, wherein: Determining the focusing covariance matrix corresponding to each of the first sub-bandwidths includes: Determining a covariance matrix corresponding to each of the first sub-bandwidths, and performing phase transformation processing on the covariance matrix; For each of the first sub-bandwidths, determining a focusing covariance matrix corresponding to each of the first sub-bandwidths according to the covariance matrix after phase transformation; Obtaining, according to the focusing covariance matrix, a sub-bandwidth focused frequency spatial spectrum corresponding to each of the first sub-bandwidths includes: Performing singular value decomposition on the focusing covariance matrix to obtain noise subspaces corresponding to the first sub-bandwidths; The sub-bandwidth focused frequency spatial spectrum corresponding to each of the first sub-bandwidths is determined according to the noise subspace.

3. The method for estimating the direction of arrival of multiple sound sources according to claim 1, wherein: Iteratively processing the observation signal using a beamformer and a dereverberation filter to obtain a target estimation result of the direction of arrival of the sound source includes: In each iteration, the beamformer is used to determine a time-varying variance based on the estimation result of the direction of arrival of the sound source obtained in the previous iteration, wherein the time-varying variance is used to reflect the energy change of the observation signal within a preset frequency range and time period. In the first iteration, the beamformer is used to determine the time-varying variance based on the initial estimation result; Eliminate the reverberation signal in the observation signal according to the time-varying variance using the dereverberation filter to obtain a dereverberation signal; using the beamformer to eliminate a noise signal in the dereverberation signal to obtain the dereverberation and denoised signal, wherein the dereverberation and denoised signal includes a plurality of second sub-bandwidths, each second sub-bandwidth in the dereverberation and denoised signal contains a plurality of frequency band signals, and there is frequency overlap between adjacent second sub-bandwidths; Determine a focusing covariance matrix corresponding to each second sub-bandwidth, and obtain a sub-bandwidth focused frequency spatial spectrum corresponding to each second sub-bandwidth according to the focusing covariance matrix; An estimation result of the sound source arrival direction of this iteration is determined according to the sub-bandwidth focused frequency spatial spectrum corresponding to the second sub-bandwidth.

4. The method for estimating the direction of arrival of multiple sound sources according to claim 3, wherein: The dereverberation filter includes a multi-channel linear prediction filter; Eliminating the reverberation signal in the observation signal according to the time-varying variance using the dereverberation filter to obtain the dereverberation signal includes: sequentially using each channel in the multi-channel linear prediction filter as a reference channel, and using the dereverberation filter to eliminate the reverberation signal in the observation signal according to the reference channel and the time-varying variance, to obtain a dereverberation signal matrix, wherein the dereverberation signal family includes dereverberation sub-signals obtained by eliminating the reverberation signal with each channel as the reference channel; The dereverberation signal matrix is ​​used as the dereverberation signal.

5. The method for estimating the direction of arrival of multiple sound sources according to claim 3, wherein: The beamformer includes a multi-objective minimum variance distortionless response beamforming filter; Eliminating the noise signal in the dereverberation signal using the beamformer to obtain the dereverberation and noise-removed signal includes: Each channel in the multi-objective minimum variance distortionless response beamforming filter is sequentially used as a reference channel, and the multi-objective minimum variance distortionless response beamforming filter is used to eliminate the noise signal in the dereverberation signal according to the reference channel to obtain the dereverberation and denoising signal.

6. The method for estimating the direction of arrival of multiple sound sources according to claim 3, wherein: Eliminating the noise signal in the dereverberation signal using the beamformer to obtain the dereverberation and noise-removed signal includes: In the case where there are multiple sound sources, determining a steering vector for each sound source; The beamformer is used to eliminate noise signals in the dereverberation signal according to the steering vectors of the respective sound sources to obtain the dereverberation and noise-removed signal.

7. The method for estimating the direction of arrival of multiple sound sources according to claim 1, wherein: Non-uniformly dividing the initial dereverberation signal into a plurality of first sub-bandwidths includes: discarding a portion of the initial dereverberation signal whose frequency is lower than a first preset frequency; For a first signal portion of the initial dereverberation signal having a frequency not less than a first preset frequency and less than a second preset frequency, dividing the first signal portion into a plurality of first sub-bandwidths at a first preset frequency interval, wherein a first preset number of frequency bands are overlapped between sub-bandwidths corresponding to two first sub-bandwidths with adjacent frequencies in the first signal portion; For a second signal portion of the initial dereverberation signal having a frequency not less than the second preset frequency and less than a third preset frequency, dividing the second signal portion into a plurality of first sub-bandwidths at a second preset frequency interval, wherein a second preset number of frequency bands are overlapped between sub-bandwidths corresponding to two first sub-bandwidths with adjacent frequencies in the second signal portion; For a third signal portion of the initial dereverberation signal having a frequency not less than a third preset frequency and less than a fourth preset frequency, dividing the third signal portion into a plurality of first sub-bandwidths at a third preset frequency interval, wherein a third preset number of frequency bands are overlapped between sub-bandwidths corresponding to two first sub-bandwidths having adjacent frequencies in the third signal portion; For a fourth signal portion of the initial dereverberation signal having a frequency not less than a fourth preset frequency and less than a fifth preset frequency, the fourth signal portion is divided into a plurality of first sub-bandwidths at a fourth preset frequency interval, wherein a fourth preset number of frequency bands are overlapped between sub-bandwidths corresponding to two first sub-bandwidths with adjacent frequencies in the fourth signal portion.

8. The method for estimating the direction of arrival of multiple sound sources according to claim 1, wherein: Determining an initial estimation result of the sound source arrival direction of the observation signal based on the sub-bandwidth focused frequency spatial spectrum includes: Summing the spatial spectra of the focused frequencies of the sub-bandwidths to obtain a significant peak value of the overall spatial spectra of the focused frequencies of the sub-bandwidths; The initial estimation result is determined according to the significant peak.

9. The method for estimating the direction of arrival of multiple sound sources according to claim 1, wherein: Determining an initial estimation result of the sound source arrival direction of the observation signal based on the sub-bandwidth focused frequency spatial spectrum includes: Determining a significant peak of each of the sub-bandwidth focused frequency spatial spectra, and an estimation result of the sound source arrival direction corresponding to the significant peak; Determining a significant peak-estimation result set, and determining distribution information of the significant peak-estimation result set; The initial estimation result is determined according to the distribution information.

10. A device for estimating the direction of arrival of multiple sound sources, characterized in that: include: a first processing module, configured to eliminate a late reverberation signal in an observation signal to obtain an initial dereverberation signal, wherein the observation signal has one or more sound sources; a second processing module, configured to non-uniformly divide the initial dereverberation signal into a plurality of first sub-bandwidths, wherein each of the first sub-bandwidths contains a plurality of frequency band signals, and there is frequency overlap between adjacent first sub-bandwidths; A third processing module is configured to determine a focusing covariance matrix corresponding to each of the first sub-bandwidths, and obtain a sub-bandwidth focused frequency spatial spectrum corresponding to each of the first sub-bandwidths according to the focusing covariance matrix; a fourth processing module, configured to determine an initial estimation result of the direction of arrival of the sound source of the observation signal according to the sub-bandwidth focused frequency spatial spectrum corresponding to the first sub-bandwidth; a fifth processing module, configured to iteratively process the observation signal using a beamformer and a dereverberation filter based on the initial estimation result to obtain a target estimation result of the direction of arrival of the sound source; wherein, in each round of iterative processing, based on the estimation result of the direction of arrival of the sound source obtained in the previous round of iterative processing, the dereverberation filter is used to eliminate the reverberation signal in the observation signal, and the beamformer is used to eliminate the noise signal in the observation signal; and the estimation result of the direction of arrival of the sound source corresponding to the current round of iteration is determined based on the dereverberation and denoised signal obtained after eliminating the reverberation signal and the noise signal.

11. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the method for estimating the direction of arrival of multiple sound sources according to any one of claims 1 to 10.

12. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the method for estimating the direction of arrival of multiple sound sources according to any one of claims 1 to 10 is executed when the program is run.

13. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method for estimating the directions of arrival of multiple sound sources according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Double sound source localization method based on consistent focusing transform least square method

    CN105301563A

  • Broadband DOA estimation method based on time-varying mixed signal blind separation

    CN112565119A

  • Multi-sound-source arrival direction estimation method and device based on frequency focusing spatial spectrum

    CN119936786A

  • Voice de-reverberation method, device, equipment and medium

    CN119993178A

  • Online dereverberation algorithm based on weighted prediction error for noisy time-varying environments

    US20180182410A1