Sound source positioning method, related apparatus and medium
By performing frame-segmentation, windowing, and noise suppression processing on the audio signals received by the microphone array, combined with an adaptive beamforming algorithm, the problem of decreased sound source localization accuracy in noisy environments is solved, achieving higher localization accuracy and real-time performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ESWIN COMPUTING TECH CO LTD
- Filing Date
- 2022-12-16
- Publication Date
- 2026-07-28
AI Technical Summary
Existing sound source localization algorithms are easily interfered with in noisy environments, leading to a decrease in localization accuracy.
By performing frame-by-frame windowing and short-time Fourier transform on the audio signal received by the microphone array, noise suppression is performed frame by frame, frequency domain frame data is reconstructed, and an adaptive beamforming algorithm is used to search for spectral peaks to estimate the direction of arrival of the sound source.
It improves the accuracy and real-time performance of sound source localization, reduces computational load, minimizes noise interference with localization, and enhances the signal-to-noise ratio of the microphone array received signals.
Smart Images

Figure CN116106826B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of voice localization technology, specifically to a sound source localization method, related apparatus and medium. Background Technology
[0002] With the development of artificial intelligence, intelligent voice devices are increasingly widely used in daily life, such as smart speakers, intelligent conference audio pickup devices, and intelligent robots. Intelligent voice devices can capture sounds from surrounding sources, and locating the sound source is essential for obtaining clear speech. Intelligent voice devices equipped with microphone arrays can use sound source localization algorithms to process the audio signals received by the microphone array and locate the direction of the sound source relative to the microphone array, i.e., the direction of arrival (DOA). Currently, sound source localization algorithms such as time-delay estimation-based, high-resolution spectrum estimation-based, and controllable beamforming-based algorithms can be used to locate the DOA of a sound source. However, since microphone arrays are often in noisy environments, even highly noisy ones, the aforementioned sound source localization algorithms are easily affected by environmental noise, thus reducing the accuracy of sound source localization. Summary of the Invention
[0003] To address the aforementioned technical problems, this disclosure provides a sound source localization method, related apparatus, and medium, which can reduce noise interference with sound source localization and improve the accuracy of sound source localization.
[0004] According to a first aspect of this disclosure, a sound source localization method is provided, comprising:
[0005] Time-domain framed data is obtained by performing frame-by-frame windowing processing on the time-domain data of the audio signal received by the microphone array.
[0006] By performing a short-time Fourier transform on the time-domain framed data, frequency-domain framed data is obtained.
[0007] Enhanced frequency domain frame data is reconstructed by suppressing noise in the frequency domain frame data frame by frame.
[0008] Based on the enhanced frequency domain frame data, an adaptive beamforming algorithm is used to perform spectral peak search to estimate the direction of arrival of the sound source frame by frame.
[0009] Optionally, the frequency domain framing data includes frequency domain framing data from multiple channels, and the enhanced frequency domain framing data, which reconstructs the frequency domain framing data by performing noise suppression on the frequency domain framing data frame by frame, includes:
[0010] Calculate the estimated noise covariance of the frequency domain frame data of the first channel of the plurality of channels;
[0011] Using the estimated noise covariance, the gain function of the frequency domain frame data of the first channel before and after noise suppression is calculated.
[0012] Based on the amplitude values of the frequency domain segmented frame data of the multiple channels in the current frame and the gain function, the enhanced frequency domain segmented frame data of the current frame after noise suppression is reconstructed.
[0013] Optionally, after performing frame-segmentation and windowing processing on the time-domain data of the audio signal received by the microphone array to obtain time-domain framed data, the sound source localization method further includes:
[0014] Using speech activity detection technology, speech activity is detected frame by frame in the time-domain frame data to obtain the speech activity detection result of each frame of the time-domain frame data, wherein the presence of speech in each frame of the time-domain frame data is determined based on the speech activity detection result.
[0015] Optionally, calculating the estimated noise covariance of the frequency domain frame data of the first channel of the plurality of channels includes:
[0016] Based on the voice activity detection results, the frequency domain representation of the frequency domain frame data of the current frame of the first channel is determined;
[0017] Based on the frequency domain representation of the frequency domain segmented frame data of the current frame of the first channel, calculate the estimated noise covariance of the frequency domain segmented frame data of the current frame of the first channel.
[0018] Specifically, if there is speech in the current frame of the first channel, the estimated noise covariance of the frequency domain subframe data of the previous frame of the first channel is used as the estimated noise covariance of the frequency domain subframe data of the current frame of the first channel. If there is no speech in the current frame of the first channel, the weighted sum of the estimated noise covariance of the frequency domain subframe data of the previous frame of the first channel and the power of the frequency domain subframe data of the current frame of the first channel is used as the estimated noise covariance of the frequency domain subframe data of the current frame of the first channel.
[0019] Optionally, the step of using the estimated noise covariance to calculate the gain function of the frequency domain frame data of the first channel before and after noise suppression includes:
[0020] Based on the estimated noise covariance and the power of the frequency domain segmented frame data of the current frame of the first channel, the posterior signal-to-noise ratio of the frequency domain segmented frame data of the current frame of the first channel is calculated.
[0021] Based on the decision-guided estimation method, the weighted average of the prior signal-to-noise ratio of the frequency domain segmented frame data of the previous frame of the first channel and the posterior signal-to-noise ratio of the frequency domain segmented frame data of the current frame of the first channel is used as the prior signal-to-noise ratio of the frequency domain segmented frame data of the current frame of the first channel.
[0022] Based on the posterior and prior signal-to-noise ratios of the frequency domain segmented data of the current frame of the first channel, and using the minimum mean square error of the logarithmic spectrum amplitude estimation as the distortion metric, the gain function of the frequency domain segmented data of the current frame of the first channel before and after noise suppression is calculated.
[0023] Optionally, the adaptive beamforming algorithm includes a minimum variance distortionless response beamforming algorithm, and the step of using the adaptive beamforming algorithm to perform spectral peak search to estimate the direction of arrival of the sound source frame by frame based on the enhanced frequency domain framed data includes:
[0024] The enhanced frequency domain framing data of the current frame of the multiple channels is used as the input of the minimum variance distortionless response beamforming algorithm. Under the condition of satisfying the constraints, the optimal weight vector of the current frame of the minimum variance distortionless response beamforming algorithm is solved by the Lagrange multiplier method.
[0025] Based on the optimal weight vector of the current frame and the enhanced frequency domain framing data of the current frame of the multiple channels, the average power of the output beam of the minimum variance distortionless response beamforming algorithm in the current frame is calculated.
[0026] Within the angular scanning range, a spectral peak search is performed on the average power of the output beam of the current frame, and the signal incident angle corresponding to the peak point is used as the estimated direction of arrival of the sound source in the current frame.
[0027] Optionally, the constraints include the requirement that the signal in the incident direction is expected to pass through completely, while suppressing interference sources and noise from other directions to the greatest extent possible.
[0028] Optionally, after estimating the direction of arrival of the sound source frame by frame using an adaptive beamforming algorithm to perform spectral peak search based on the enhanced frequency domain framed data, the sound source localization method further includes:
[0029] Based on the speech activity detection results, the estimated direction of arrival (DOA) of the sound source in the current frame is smoothed. If speech is present in the temporal framing data of the current frame, the estimated DOA of the sound source in the current frame is determined as the final DOA of the sound source in the current frame. If speech is not present in the temporal framing data of the current frame, the estimated DOA of the sound source in the previous frame is determined as the final DOA of the sound source in the current frame.
[0030] According to a second aspect of this disclosure, a sound source localization device is provided, comprising:
[0031] The frame-segmentation and windowing processing unit is configured to perform frame-segmentation and windowing processing on the time-domain data of the audio signal received by the microphone array to obtain time-domain framed data.
[0032] The short-time Fourier transform unit is configured to obtain frequency-domain frame data by performing a short-time Fourier transform on the time-domain frame data.
[0033] The frequency domain frame data enhancement unit is configured to reconstruct enhanced frequency domain frame data by performing noise suppression on the frequency domain frame data frame by frame.
[0034] The direction-of-arrival estimation unit is configured to perform spectral peak search using an adaptive beamforming algorithm based on the enhanced frequency domain frame data to estimate the direction of arrival of the sound source frame by frame.
[0035] Optionally, the sound source localization device further includes:
[0036] The voice activity detection unit is configured to use voice activity detection technology to perform voice activity detection on the time-domain frame data frame by frame, and obtain the voice activity detection result of the time-domain frame data for each frame, wherein the presence of voice in the time-domain frame data for each frame is determined based on the voice activity detection result.
[0037] According to a third aspect of this disclosure, a sound source localization system is provided, comprising:
[0038] The microphone array is configured to receive time-domain data of audio signals from the sound source;
[0039] The sound source localization device is configured to perform the method described in any of the above-described embodiments.
[0040] According to a fourth aspect of this disclosure, an electronic device is provided, comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method described above.
[0041] According to a fifth aspect of this disclosure, a storage medium is provided that stores a computer program or instructions, which, when executed by a processor, implement the steps of the method described above.
[0042] According to embodiments of this disclosure, time-domain framed data is obtained by performing frame-by-frame windowing on the time-domain data of the audio signal received by the microphone array. Frequency-domain framed data is then obtained by performing a short-time Fourier transform on the time-domain framed data. Noise suppression is applied frame-by-frame to reconstruct enhanced frequency-domain framed data. Based on this enhanced frequency-domain framed data, an adaptive beamforming algorithm is used to search for spectral peaks to estimate the direction of arrival (DOA) of the sound source frame-by-frame. Thus, before sound source localization, the frequency-domain framed data of the audio signals from multiple channels received by the microphone array is preprocessed using noise suppression technology, and the noise-suppressed enhanced frequency-domain framed data is reconstructed. This improves the signal-to-noise ratio of the audio signal received by the microphone array, thereby increasing the accuracy of sound source localization.
[0043] By using the enhanced frequency domain frame data of the current frame from multiple channels as input to the adaptive beamforming algorithm, and performing spectral peak search to estimate the direction of arrival (DOA) of the sound source frame by frame, the DOA of the sound source can be estimated by performing only one short-time Fourier transform on the time-domain frame data of the audio signals from multiple channels received by the microphone array. This reduces the computational load and improves the real-time performance of sound source localization.
[0044] By calculating the estimated noise covariance of the frequency domain frame data of the first channel in multiple channels, and using the estimated noise covariance, the gain function of the frequency domain frame data of the first channel before and after noise suppression is calculated. Based on the amplitude values of the frequency domain frame data of multiple channels in the current frame and the gain function, the enhanced frequency domain frame data of the current frame of the frequency domain frame data of multiple channels after noise suppression is reconstructed. In this way, it is only necessary to perform noise suppression on the frequency domain frame data of one channel and calculate the gain function of the frequency domain frame data of that channel. This gain function is then used as the gain function of all channels to reconstruct the enhanced frequency domain frame data of the frequency domain frame data of multiple channels after noise suppression, which reduces the amount of computation and improves the real-time performance of sound source localization.
[0045] By utilizing speech activity detection technology, speech activity is detected frame by frame in the temporal segmented data. After obtaining the speech activity detection results for each frame, the frequency domain representation of the current frame's frequency domain segmented data for the first channel among multiple channels received by the microphone array is determined based on the speech activity detection results. According to the frequency domain representation of the current frame's frequency domain segmented data for the first channel received by the microphone array, the estimated noise covariance of the current frame's frequency domain segmented data for the first channel is calculated. The speech activity detection results are applied during noise estimation, effectively reducing the estimation error of the noise covariance and thus improving the accuracy of sound source localization. Furthermore, based on the speech activity detection results, the estimated direction of arrival (DOA) of the sound source in the current frame is smoothed, effectively avoiding abnormal abrupt changes in the estimated DOA caused by noise, improving the accuracy of the estimated DOA, and thus improving the accuracy of sound source localization.
[0046] It should be noted that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this disclosure. Attached Figure Description
[0047] Figure 1 This diagram illustrates the structure of a sound source localization system provided according to an embodiment of the present disclosure.
[0048] Figure 2 A schematic flowchart of a sound source localization method provided according to an embodiment of the present disclosure is shown;
[0049] Figure 3 This diagram illustrates a process for reconstructing enhanced frequency domain frame data by performing noise suppression, according to an embodiment of the present disclosure.
[0050] Figure 4 This diagram illustrates a process for estimating the direction of arrival of a sound source frame by frame using a minimum variance distortionless response beamforming algorithm according to an embodiment of the present disclosure.
[0051] Figure 5 This diagram illustrates the structure of a sound source localization device provided according to an embodiment of the present disclosure.
[0052] Figure 6 A schematic diagram of the structure of an electronic device provided according to an embodiment of the present disclosure is shown. Detailed Implementation
[0053] To facilitate understanding of this disclosure, a more complete description will now be given with reference to the accompanying drawings, which illustrate preferred embodiments of the present disclosure. However, this disclosure may be implemented in various forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure.
[0054] The following concepts are used in this article:
[0055] A microphone array is a collection of sound-collecting devices arranged in a predetermined topology to sample and process audio signals from different directions in space. Each sound-collecting device in a microphone array can be called an element, and each microphone array includes at least two elements. As an example, the sound-collecting devices can be acoustic sensors or microphones. Based on their topology, microphone arrays can be classified as linear microphone arrays, planar microphone arrays, and stereo microphone arrays, etc.
[0056] Sound source localization: When the location of the sound source is unknown, array elements are arranged according to a set topology to form a microphone array. The audio signals collected by the microphone array are processed by a corresponding algorithm to obtain the direction of the sound source relative to the microphone array, that is, the direction of arrival (DOA) of the sound source.
[0057] Voice Activity Detection (VAD): also known as speech endpoint detection and speech boundary detection, refers to the detection of the presence or absence of speech in a noisy environment.
[0058] Noise suppression is the filtering process used to process audio signals. This involves filtering out noise from the audio signal to extract the audio signal in the direction of the target sound source, and suppressing or eliminating audio signals from other directions, including those generated by the speaker or background noise, in order to obtain or recover a clearer audio signal.
[0059] Minimum variance distortionless response (MVDR) beamforming algorithm is an adaptive beamforming algorithm based on the maximum signal-to-noise ratio (SNR) criterion. It can adaptively minimize the power of the microphone array output beam in the desired direction while maximizing the SNR.
[0060] Beamforming is a process that applies time delay or phase compensation and amplitude weighting to the output of each element of a microphone array to form a beam in a specific direction. This can pick up signals within the beam and eliminate noise outside the beam, thereby achieving the purpose of speech enhancement.
[0061] Figure 1 A schematic diagram of the structure of a sound source localization system provided according to an embodiment of the present disclosure is shown. Figure 1As shown, the sound source localization system 100 includes a smart voice device 110 and a server 120. In some embodiments, the smart voice device 110 and the server 120 can be connected via a network, for example, the smart voice device 110 and the server 120 can be connected to the network via GPRS, 4G, WIFI, etc.
[0062] The intelligent voice device 110 is a device that supports voice interaction (such as a ticket machine, microphone, voice recorder, smart speaker, and mobile communication device), and it is equipped with a microphone array 111 consisting of multiple microphone elements 112. The microphone array 111 can receive audio signals from sound sources in different directions. The server 120 is a server that provides business services to the intelligent voice device 110. It has voice recognition capabilities and is equipped with a sound source localization device 121. The sound source localization device 121 can obtain the direction of the sound source relative to the microphone array 111, that is, the direction of arrival (DOA) of the sound source, thereby clearly acquiring the audio signals received by the microphone array 111. The server 120 can obtain command information by performing voice recognition on the clearly acquired audio signals, thereby providing relevant business services to the users of the intelligent voice device 110. Since the specific steps for estimating the direction of arrival of the sound source based on the audio signals received by the microphone array 111 will be detailed below, they will not be repeated here.
[0063] It should be noted that, Figure 1 The microphone array 111 in the example is just one example. Multiple microphone array elements 112 can also be configured in other topologies. The microphone array 111 can be a linear microphone array, a planar microphone array, or a stereo microphone array, etc.
[0064] It should be noted that server 120 can be a server deployed in the cloud, a server deployed locally specifically for speech recognition processing of smart voice device 110, or a server deployed within smart voice device 110 for speech recognition processing. For example, if smart voice device 110 is a ticket vending machine, it can be deployed in scenarios such as subways and train stations, and server 120 can be deployed in the cloud, providing business services to smart voice device 110 in the form of a data center. If smart voice device 110 is a microphone, voice recorder, or smart speaker, it can be deployed at a meeting venue, and server 120 can be deployed locally as a dedicated server for speech recognition processing of smart voice device 110. If smart voice device 110 is a mobile communication device, server 120 can be deployed on smart voice device 110 as a dedicated processor.
[0065] Figure 2A schematic flowchart of a sound source localization method according to an embodiment of this disclosure is shown. (See also...) Figure 2 The sound source localization method provided in this embodiment includes steps S210 to S240.
[0066] In step S210, time-domain framed data is obtained by performing frame-segmentation and windowing processing on the time-domain data of the audio signal received by the microphone array.
[0067] In some embodiments, a microphone array is composed of multiple microphone elements arranged in a predetermined topology to sample and process audio signals from sound sources in different spatial directions. Based on the topology, microphone arrays can be classified as linear microphone arrays, planar microphone arrays, and stereo microphone arrays, etc. Audio signals can be understood as carriers of sound wave frequency and amplitude variation information. Due to the presence of noise, the audio signals received by the microphone array are noisy signals, including both pure speech and noise. The audio signals received by the microphone array include audio signals from multiple channels arriving at the multiple microphone elements from the sound source. Each microphone element samples its received audio signals using a preset sampling frequency to convert the audio signals into processable digital signals, thereby obtaining time-domain data for multiple channels.
[0068]
[0069] Where y(t) is the time-domain data of the audio signals from multiple channels received by the microphone array, y i (t) is the time-domain data of the audio signal of the i-th channel received by the i-th microphone element of the microphone array, 0≤i<m, where i and m are integers and m is the number of microphone elements.
[0070] It should be noted that audio signals are time-varying and non-stationary, but they possess short-term stationarity. Generally, the characteristics of audio signals are considered relatively stable within a time range of 10-30ms. Therefore, by performing frame-by-frame windowing processing on continuous time-domain data from multiple channels, each frame can be approximated as a stationary signal. To increase the coherence between two frames, the framing method can be an overlapping segmentation method, where the overlap between frames is called frame shift. In some embodiments, the frame shift can be half the frame length. Furthermore, the window function used can be a rectangular window, Hanning window, Hamming window, triangular window, etc. For example, a Hamming window can be selected, as its low-pass characteristics are relatively smooth and can well reflect the frequency characteristics of short-time signals like speech.
[0071] In some embodiments, after acquiring continuous time-domain data from multiple channels received by the microphone array, the time-domain data of each channel can be processed by framing and windowing using a window of finite length, thereby obtaining time-domain framed data:
[0072]
[0073] Where y(k,n) is the time-domain frame data of the audio signals from multiple channels received by the microphone array, y i (k,n) is the time-domain framed data of the audio signal of the i-th channel received by the i-th microphone element of the microphone array. y(k,n) can be defined as s(k,n)+d(k,n), where s(k,n) is the time-domain framed data of the clean speech signal received by the microphone array, and d(k,n) is the time-domain framed data of the noise received by the microphone array. 0≤i<m, i, m, n, and k are integers, m is the number of microphone elements, n is the frame index, and k is the frequency index.
[0074] In some embodiments, in real-world speech environments, the audio signals from multiple channels received by a microphone array typically contain a significant amount of noise. After acquiring the temporal frame data of the audio signals from multiple channels received by the microphone array, speech activity detection technology is used to perform speech activity detection frame by frame on the temporal frame data, obtaining the speech activity detection result for each frame of the temporal frame data. Based on the speech activity detection result, it is determined whether speech exists in each frame of the temporal frame data. In some embodiments, the temporal frame data of the audio signals from multiple channels received by the microphone array can be fed into a neural network-based speech activity detection model. The neural network-based speech activity detection model can then classify whether speech exists in each frame of the temporal frame data, thereby separating speech frames from noise frames.
[0075]
[0076] Wherein, VAD(·) represents the process of speech activity detection, n is the frame index and is an integer. When the result of speech activity detection is 1, it means that there is speech in the time-domain frame data of the nth frame and the nth frame is a speech frame. When the result of speech activity detection is 0, it means that there is no speech in the time-domain frame data of the nth frame and the nth frame is a noise frame.
[0077] Since neural network-based speech activity detection models and their training methods are relatively mature existing technologies, they will not be elaborated on here.
[0078] It should be noted that, in addition to the neural network-based speech activity detection model, other speech activity detection techniques can also be used in the embodiments of this disclosure to identify whether there is speech in each frame of temporal segmented data. Any speech activity detection technique that can separate speech frames and noise frames is within the protection scope of this disclosure.
[0079] In step S220, frequency domain frame data is obtained by performing a short-time Fourier transform on the time-domain frame data.
[0080] Short-Time Fourier Transform (STFT) is the most widely used method for studying non-stationary signals. The basic idea of STFT is to divide the signal into many small time intervals and then use Fourier transform to analyze each time interval in order to determine the frequencies present in each time interval. In some embodiments, by performing STFT on the aforementioned time-domain framed data, frequency-domain framed data of the audio signals from multiple channels received by the microphone array can be obtained.
[0081] Y(k,n)=STFT(y(k,n)) (4)
[0082] Where Y(k,n) is the frequency domain frame data of the audio signals from multiple channels received by the microphone array, y(k,n) is the time domain frame data of the audio signals from multiple channels received by the microphone array, Y(k,n) can be defined as S(k,n)+D(k,n), where S(k,n) is the frequency domain frame data of the clean speech signal received by the microphone array, D(k,n) is the frequency domain frame data of the noise received by the microphone array, STFT(·) represents the short-time Fourier transform process, n and k are integers, n is the frame index, and k is the frequency point index.
[0083] In some embodiments, y(k,n) is the time-domain framed data of the audio signals from multiple channels received by the microphone array. i (k,n) is the time-domain framed data of the audio signal from the i-th channel received by the i-th microphone element of the microphone array. Therefore, performing a short-time Fourier transform on y(k,n) can be understood as performing a short-time Fourier transform on y(k,n) respectively. i Perform a short-time Fourier transform on (k,n). The frequency domain frame data Y(k,n) of the audio signals from multiple channels received by the microphone array can be expressed as:
[0084]
[0085] Where Y(k,n) is the frequency domain frame data of the audio signals from multiple channels received by the microphone array, Y i(k,n) is the frequency domain frame data of the audio signal of the i-th channel received by the i-th microphone element of the microphone array, 0≤i<m, where i and m are integers and m is the number of microphone elements.
[0086] In step S230, the enhanced frequency domain frame data of the frequency domain frame data is reconstructed by performing noise suppression on the frequency domain frame data frame by frame.
[0087] Sound source localization involves arranging array elements in a predetermined topology to form a microphone array when the location of the sound source is unknown. The audio signals acquired by the microphone array are then processed using a beamforming algorithm to obtain the direction of the sound source relative to the array, i.e., the direction of arrival (DOA). However, the audio signals received by the microphone array from multiple channels typically contain a significant amount of noise. This noise interferes with the processing of the audio signals by the beamforming algorithm, thus reducing the accuracy of sound source localization. Therefore, before sound source localization, it is necessary to preprocess the frequency domain frame data of the audio signals received by the microphone array using noise suppression techniques. This reconstructs the enhanced frequency domain frame data after noise suppression, improving the signal-to-noise ratio of the audio signals received by the microphone array and thereby increasing the accuracy of sound source localization.
[0088] Figure 3 This diagram illustrates a process for reconstructing enhanced frequency domain frame data by performing noise suppression, according to an embodiment of the present disclosure. (See reference...) Figure 3 The method for reconstructing enhanced frequency domain frame data by performing noise suppression provided in this embodiment includes steps S310 to S330.
[0089] In step S310, the estimated noise covariance of the frequency domain frame data of the first channel of the plurality of channels is calculated.
[0090] Since the accuracy of noise estimation determines the effectiveness of noise suppression, and considering that noise in real-world environments typically has a non-uniform impact on the speech spectrum, this disclosure employs a time-recursive smoothing method based on the probability of speech presence for noise estimation. In some embodiments, speech activity detection technology is used to perform speech activity detection frame-by-frame on the time-domain framed data. After obtaining the speech activity detection result for each frame of time-domain framed data, the frequency domain representation of the current frame of the first channel among the multiple channels received by the microphone array is determined based on the speech activity detection result. Based on the speech activity detection result, the frequency domain representation of the nth frame of the i-th channel is as follows:
[0091]
[0092] Where H0(k,n) represents the state where there is no speech in the nth frame of the i-th channel's temporal framing data, and H1(k,n) represents the state where there is speech in the nth frame of the i-th channel's temporal framing data. i (k,n) represents the nth frame of the audio signal from the i-th channel received by the microphone array in the frequency domain. In the H0(k,n) state, Y... i (k,n) is defined as D i (k,n), in state H1(k,n), Y i (k,n) is defined as S i (k,n)+D i (k,n), S i (k,n) represents the nth frame of the clean speech signal from the i-th channel received by the microphone array, D. i (k,n) represents the nth frame of frequency domain data of the noise of the i-th channel received by the microphone array, where 0 ≤ i < m, i, n, k, and m are integers, m is the number of microphone array elements, n is the frame index, and k is the frequency point index.
[0093] Next, based on the frequency domain representation of the frequency domain segmented frame data of the current frame of the first channel received by the microphone array, the estimated noise covariance of the frequency domain segmented frame data of the current frame of the first channel is calculated. In some embodiments, if there is speech in the current frame of the first channel, the estimated noise covariance of the frequency domain segmented frame data of the previous frame of the first channel is used as the estimated noise covariance of the frequency domain segmented frame data of the current frame of the first channel; if there is no speech in the current frame of the first channel, the weighted sum of the estimated noise covariance of the frequency domain segmented frame data of the previous frame of the first channel and the power of the frequency domain segmented frame data of the current frame of the first channel is used as the estimated noise covariance of the frequency domain segmented frame data of the current frame of the first channel.
[0094] In some embodiments, noise estimation is performed as follows:
[0095]
[0096] in, It is the estimated noise covariance of the nth frame of the frequency domain data of the audio signal from the i-th channel received by the microphone array. It is the estimated noise covariance of the (n-1)th frame of the frequency domain data of the audio signal of the i-th channel received by the microphone array, |Y i (k,n)| 2 α is the power of the nth frame of frequency domain data of the audio signal from the i-th channel received by the microphone array. d It is a smoothing factor, 0 < α d <1.
[0097] In state H0(k,n), i.e., when there is no speech in the nth frame of the time-domain segmented data of the i-th channel, the estimated noise covariance of the nth frame of the frequency-domain segmented data of the audio signal of the i-th channel received by the microphone array. This represents the estimated noise covariance of the (n-1)th frame of the frequency domain data of the audio signal from the i-th channel received by the microphone array. The weighted sum of the power of the nth frame frequency domain sub-frame data of the audio signal from the i-th channel received by the microphone array. The estimated noise covariance of the nth frame frequency domain sub-frame data of the audio signal from the i-th channel received by the microphone array in state H1(k,n), i.e., when speech is present in the nth frame time domain sub-frame data of the i-th channel. This represents the estimated noise covariance of the (n-1)th frame of the frequency domain data of the audio signal from the i-th channel received by the microphone array.
[0098] In step S320, the gain function of the frequency domain frame data of the first channel before and after noise suppression is calculated using the estimated noise covariance.
[0099] In some embodiments, after obtaining the estimated noise covariance, the prior signal-to-noise ratio (SNR) and posterior signal-to-noise ratio (SNR) can be further estimated, thereby obtaining the gain function of the frequency domain framed data of the current frame of the first channel received by the microphone array before and after noise suppression. In some embodiments, based on the estimated noise covariance and power of the nth frame frequency domain framed data of the audio signal of the i-th channel received by the microphone array, the formula for estimating the posterior SNR of the nth frame frequency domain framed data of the audio signal of the i-th channel received by the microphone array is as follows:
[0100]
[0101] Where, γ k (n) is the posterior signal-to-noise ratio of the nth frame of the frequency domain segmented data of the audio signal of the i-th channel received by the microphone array. It is the estimated noise covariance of the nth frame of the frequency domain data of the audio signal of the i-th channel received by the microphone array, |Y i (k,n)| 2 It is the power of the nth frame of frequency domain data of the audio signal of the i-th channel received by the microphone array.
[0102] In some embodiments, the prior signal-to-noise ratio (SNR) estimation of the nth frame frequency domain segmented data of the audio signal of the i-th channel received by the microphone array adopts a decision-guided estimation method. The weighted average of the estimated prior SNR of the (n-1)th frame frequency domain segmented data of the audio signal of the i-th channel received by the microphone array and the posterior SNR of the nth frame frequency domain segmented data of the audio signal of the i-th channel received by the microphone array is used as the estimated prior SNR of the nth frame frequency domain segmented data of the audio signal of the i-th channel received by the microphone array. In some embodiments, the formula for estimating the prior SNR of the nth frame frequency domain segmented data of the audio signal of the i-th channel received by the microphone array is:
[0103]
[0104] Where, ξ k (n) is the a priori signal-to-noise ratio of the nth frame of the audio signal from the i-th channel received by the microphone array. It is the estimated noise covariance of the (n-1)th frame of the frequency domain data of the audio signal of the i-th channel received by the microphone array, |X(k,n-1)|. 2 γ is the power of the enhanced frequency domain frame data of the (n-1)th frame of the audio signal received by the microphone array after noise suppression. k (n) is the posterior signal-to-noise ratio of the nth frame of the frequency domain segmented data of the audio signal of the i-th channel received by the microphone array, max[γ k [(n)-1,0] indicates the selection of γ k The larger of (n)-1 and 0, where a is the weighting factor.
[0105] In some embodiments, after estimating the prior and posterior signal-to-noise ratios, the gain function of the frequency domain frame data of the first channel before and after noise suppression is obtained using statistical knowledge, with the minimum mean square error of the logarithmic spectrum amplitude estimation as the distortion criterion. In some embodiments, the formula for calculating the gain function of the frequency domain frame data of the nth frame of the audio signal of the i-th channel before and after noise suppression is as follows:
[0106]
[0107] Wherein G(ξ) k (n),v k (n) is the gain function of the nth frame of frequency domain data of the audio signal of the i-th channel received by the microphone array before and after noise suppression, ξ k (n) is the prior signal-to-noise ratio (SNR) of the nth frame of the audio signal from the i-th channel received by the microphone array, γ. k(n) is the posterior signal-to-noise ratio of the nth frame of the frequency domain segmented data of the audio signal of the i-th channel received by the microphone array.
[0108] In step S330, based on the amplitude values of the frequency domain segmented frame data of the multiple channels of the current frame and the gain function, the enhanced frequency domain segmented frame data of the current frame of the multiple channels after noise suppression is reconstructed.
[0109] In some embodiments, the gain function G(ξ) is based on the nth frame frequency domain segmentation data Y(k,n) of the audio signals from multiple channels received by the microphone array and the nth frame frequency domain segmentation data of the audio signals from the i-th channel received by the microphone array before and after noise suppression. k (n),v k (n)), the formula for calculating the amplitude value of the enhanced frequency domain frame data of the nth frame of the audio signal after noise suppression is as follows:
[0110] X k,n =G(ξ k (n),v k (n))Y k,n (11)
[0111] Among them, X k,n G(ξ) is the amplitude value of the nth frame of enhanced frequency domain frame data of the audio signals from multiple channels received by the microphone array after noise suppression. k (n),v k Y(n) is the gain function of the nth frame of frequency domain data of the audio signal of the i-th channel received by the microphone array before and after noise suppression. k,n It is the amplitude value of the nth frame of frequency domain data of the audio signals from multiple channels received by the microphone array.
[0112] In some embodiments, the formula for calculating the enhanced frequency domain frame data of the nth frame of the audio signal after noise suppression is as follows:
[0113] X(ω k,n ) = X k,n e jθ′(k,n) (12)
[0114] Where X(ω) k,n X is the enhanced frequency domain frame data of the nth frame of the audio signals from multiple channels received by the microphone array after noise suppression. k,n It is the amplitude value of the nth frame of the enhanced frequency domain frame data of the audio signals from multiple channels received by the microphone array after noise suppression.
[0115] In step S240, based on the enhanced frequency domain frame data, an adaptive beamforming algorithm is used to perform spectral peak search to estimate the direction of arrival of the sound source frame by frame.
[0116] Beamforming involves weighted summation of the audio signals received by each array element and adjusting the microphone array's receiving direction to a specified sound source direction according to a pre-set angle. This achieves directional selection of the audio signal by the microphone array, enhancing the audio signal in the specified direction and suppressing interfering sound sources and noise from other directions. In this embodiment, an adaptive beamforming (ABF) algorithm (e.g., minimum variance distortionless response beamforming algorithm) is used to estimate the direction of arrival of the sound source frame by frame.
[0117] In some embodiments, when using an adaptive beamforming algorithm to perform beamforming processing in a desired direction according to a preset angle, the weighting coefficients of the output signal of each array element are not fixed. Instead, they can be automatically adjusted according to changes in the environment and signal, using an adaptive algorithm, to ensure that the formed beam always points to the preset angle, thereby suppressing interference sources and noise and enhancing the useful signal within the beam. Adaptive beamforming algorithms include, for example, algorithms based on the maximum signal-to-noise ratio (SNR) criterion, algorithms based on the minimum mean square error (MMSE) criterion, and algorithms based on the linearly constrained minimum variance (LCMV) criterion. This disclosure does not specifically limit the specific adaptive beamforming algorithm used. The following detailed explanation uses the minimum variance distortionless response beamforming algorithm to estimate the direction of arrival of a sound source frame by frame as an example.
[0118] Figure 4 This diagram illustrates a flowchart of a frame-by-frame estimation of the direction of arrival (DOA) of a sound source using a minimum variance distortionless response beamforming algorithm according to an embodiment of this disclosure. (See reference...) Figure 4 The method for estimating the direction of arrival of a sound source frame by frame using the minimum variance distortionless response beamforming algorithm provided in this embodiment includes steps S410 to S430.
[0119] In step S410, the enhanced frequency domain framing data of the current frame of the multiple channels is used as the input to the minimum variance distortionless response beamforming algorithm. Under the condition that the constraints are met, the optimal weight vector of the current frame of the minimum variance distortionless response beamforming algorithm is solved using the Lagrange multiplier method. The constraints include the expectation that the signal in the incident direction will pass through completely, and the maximum suppression of interference sources and noise in other directions.
[0120] In some embodiments, based on the idea of adaptive beamforming algorithms, in order to obtain audio signals in the desired direction and suppress interference sources and noise from other directions, it is necessary to design a weight vector for the minimum variance distortionless response beamforming algorithm. In some embodiments, the nth frame enhanced frequency domain frame data of the audio signals from multiple channels received by the noise-suppressed microphone array are weighted, and the formula for obtaining the nth frame output beam of the minimum variance distortionless response beamforming algorithm is as follows:
[0121]
[0122] Where a(θ) is the steering vector of the microphone array, which can be expressed as X(ω k,n S(ω) is the enhanced frequency domain frame data of the nth frame of the audio signals from multiple channels received by the microphone array after noise suppression. k,n D(ω) is the nth frame of the clean speech signal from multiple channels received by the microphone array after noise suppression, in the frequency domain. k,n ) is the frequency domain data of the nth frame of noise, ω is the weight vector, and m is the number of microphone array elements.
[0123] The power of the nth frame output beam of the minimum variance distortionless response beamforming algorithm can be calculated using formula (13):
[0124]
[0125] Where P(θ) is the power of the nth frame output beam of the minimum variance distortionless response beamforming algorithm, a(θ) is the steering vector of the microphone array, ω is the weight vector, and R = X(ω) k,n )X(ω k,n ) H , is the autocorrelation matrix, X(ω) k,n () is the enhanced frequency domain frame data of the nth frame of the audio signals from multiple channels received by the microphone array after noise suppression.
[0126] Assuming the desired incident direction is θ, in order for the signal with the desired incident direction θ to pass completely, constraint (1) must be satisfied: ω Ha(θ)=1; In order to suppress noise and interference signals from other directions to the greatest extent, it is necessary to satisfy constraint (2) on the basis of satisfying constraint (1): P(θ) is minimized, and then the optimal weight vector of the nth frame of the minimum variance distortionless response beamforming algorithm can be solved by using the Lagrange multiplier method:
[0127]
[0128] Where, ω opt is the optimal weight vector for the nth frame of the minimum variance distortionless response beamforming algorithm, a(θ) is the steering vector of the microphone array, and R is the autocorrelation matrix.
[0129] In step S420, based on the optimal weight vector of the current frame and the enhanced frequency domain framing data of the current frame of the multiple channels, the average power of the output beam of the minimum variance distortionless response beamforming algorithm for the current frame is calculated.
[0130] In some embodiments, the optimal weight vector ω of the nth frame of the minimum variance distortionless response beamforming algorithm is used. opt Substituting into formula (13), the average power of the output beam in the nth frame of the minimum variance distortionless response beamforming algorithm can be calculated:
[0131]
[0132] Where P'(θ) is the average power of the output beam of the nth frame of the minimum variance distortionless response beamforming algorithm, a(θ) is the steering vector of the microphone array, and R is the autocorrelation matrix.
[0133] In step S430, within the angle scanning range, a spectral peak search is performed on the average power of the output beam of the current frame, and the signal incident angle corresponding to the peak point is used as the estimated direction of arrival of the sound source in the current frame.
[0134] In some embodiments, the angle scanning range is determined according to the topology of the microphone array and the actual environment. Then, the average power P'(θ) of the current frame output beam of the minimum variance distortionless response beamforming algorithm is calculated successively according to a certain step size. Spectral peak search is performed, and the signal incident angle θ corresponding to the peak point is used as the predicted direction of arrival of the sound source in the current frame.
[0135] When environmental noise is high, the minimum variance distortionless response beamforming algorithm (MINRR) is prone to misjudgment, causing the direction of arrival (DOA) of the sound source estimated frame-by-frame using MINRR, which deviates significantly from the actual sound source location. Therefore, in some embodiments, after using an adaptive beamforming algorithm to perform spectral peak search to estimate the DOA of the sound source frame-by-frame based on the enhanced frequency domain frame data of the current frame from multiple channels, the estimated DOA of the sound source in the current frame is smoothed based on the speech activity detection results. In some embodiments, speech activity detection technology is used to determine whether speech exists in the temporal domain frame data of the current frame. If speech exists in the temporal domain frame data of the current frame, the estimated DOA of the sound source in the current frame is determined as the final DOA of the sound source in the current frame. If no speech exists in the temporal domain frame data of the current frame, the estimated DOA of the sound source in the previous frame is determined as the final DOA of the sound source in the current frame. In this way, based on the speech activity detection results, the estimated direction of arrival of the sound source in the current frame is smoothed, which effectively avoids abnormal abrupt changes in the estimated direction of arrival of the sound source caused by noise and improves the accuracy of the estimated direction of arrival of the sound source.
[0136] Furthermore, this disclosure also provides a sound source localization device for implementing the aforementioned sound source localization method. (See reference...) Figure 5 The sound source localization device 500 disclosed in this embodiment includes: a frame windowing processing unit 510, a short-time Fourier transform unit 520, a frequency domain frame data enhancement unit 530, a direction of arrival estimation unit 540, and a speech activity detection unit 550.
[0137] The system includes a frame-segmentation and windowing processing unit 510, configured to perform frame-segmentation and windowing processing on the time-domain data of the audio signal received by the microphone array to obtain time-domain frame-segmented data. A short-time Fourier transform unit 520 is configured to perform a short-time Fourier transform on the time-domain frame-segmented data to obtain frequency-domain frame-segmented data. A frequency-domain frame-segmented data enhancement unit 530 is configured to reconstruct enhanced frequency-domain frame-segmented data by performing noise suppression on the frequency-domain frame-segmented data frame by frame. A direction-of-arrival estimation unit 540 is configured to estimate the direction of arrival of the sound source frame by frame by performing spectral peak search using an adaptive beamforming algorithm based on the enhanced frequency-domain frame-segmented data. A speech activity detection unit 550 is configured to perform speech activity detection on the time-domain frame-segmented data frame by frame using speech activity detection technology to obtain a speech activity detection result for each frame of the time-domain frame-segmented data, wherein the presence of speech in each frame of the time-domain frame-segmented data is determined based on the speech activity detection result.
[0138] In practical implementation, each module / unit in the sound source localization device can be implemented as an independent entity, or it can be arbitrarily combined to be implemented as the same or several entities. Furthermore, the specific implementation of each module / unit in the sound source localization device described above can be found in the aforementioned sound source localization method embodiments, and will not be repeated here.
[0139] This disclosure also provides a sound source localization system for implementing the aforementioned sound source localization method. A microphone array is configured to receive time-domain data of audio signals from a sound source. A sound source localization device is configured to execute the aforementioned sound source localization method. Specific implementations of each module / unit in the sound source localization system described above can be found in the aforementioned sound source localization method embodiments, and will not be repeated here.
[0140] This disclosure also provides an electronic device, such as... Figure 6 As shown, it includes a memory 620, a processor 610, and a program stored in the memory 620 and executable on the processor 610. When the program is executed by the processor 610, it can implement the various processes of the above-described sound source localization methods and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0141] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. Therefore, this disclosure also provides a storage medium storing a computer program or instructions, which, when executed by a processor, can implement the various processes of the embodiments of the above sound source localization methods.
[0142] Since the instructions stored in the storage medium can execute the steps in the sound source localization method provided in the embodiments of this disclosure, the beneficial effects achievable by the sound source localization method provided in the embodiments of this disclosure can be realized, as detailed in the preceding embodiments, and will not be repeated here. The specific implementation of each of the above operations can be found in the preceding embodiments, and will not be repeated here.
[0143] In summary, according to the embodiments of this disclosure, by performing frame-by-frame windowing processing on the time-domain data of the audio signal received by the microphone array, time-domain framed data is obtained. By performing short-time Fourier transform on the time-domain framed data, frequency-domain framed data is obtained. By performing noise suppression on the frequency-domain framed data frame by frame, enhanced frequency-domain framed data is reconstructed. Based on the enhanced frequency-domain framed data, an adaptive beamforming algorithm is used to perform spectral peak search to estimate the direction of arrival of the sound source frame by frame. In this way, before sound source localization, the frequency-domain framed data of the audio signals from multiple channels received by the microphone array is preprocessed by noise suppression technology, and the enhanced frequency-domain framed data after noise suppression is reconstructed, which improves the signal-to-noise ratio of the audio signal received by the microphone array, thereby improving the accuracy of sound source localization.
[0144] By using the enhanced frequency domain frame data of the current frame from multiple channels as input to the adaptive beamforming algorithm, and performing spectral peak search to estimate the direction of arrival (DOA) of the sound source frame by frame, the DOA of the sound source can be estimated by performing only one short-time Fourier transform on the time-domain frame data of the audio signals from multiple channels received by the microphone array. This reduces the computational load and improves the real-time performance of sound source localization.
[0145] By calculating the estimated noise covariance of the frequency domain frame data of the first channel in multiple channels, and using the estimated noise covariance, the gain function of the frequency domain frame data of the first channel before and after noise suppression is calculated. Based on the amplitude values of the frequency domain frame data of multiple channels in the current frame and the gain function, the enhanced frequency domain frame data of the current frame of the frequency domain frame data of multiple channels after noise suppression is reconstructed. In this way, it is only necessary to perform noise suppression on the frequency domain frame data of one channel and calculate the gain function of the frequency domain frame data of that channel. This gain function is then used as the gain function of all channels to reconstruct the enhanced frequency domain frame data of the frequency domain frame data of multiple channels after noise suppression, which reduces the amount of computation and improves the real-time performance of sound source localization.
[0146] By utilizing speech activity detection technology, speech activity is detected frame by frame in the temporal segmented data. After obtaining the speech activity detection results for each frame, the frequency domain representation of the current frame's frequency domain segmented data for the first channel among multiple channels received by the microphone array is determined based on the speech activity detection results. According to the frequency domain representation of the current frame's frequency domain segmented data for the first channel received by the microphone array, the estimated noise covariance of the current frame's frequency domain segmented data for the first channel is calculated. The speech activity detection results are applied during noise estimation, effectively reducing the estimation error of the noise covariance and thus improving the accuracy of sound source localization. Furthermore, based on the speech activity detection results, the estimated direction of arrival (DOA) of the sound source in the current frame is smoothed, effectively avoiding abnormal abrupt changes in the estimated DOA caused by noise, improving the accuracy of the estimated DOA, and thus improving the accuracy of sound source localization.
[0147] Finally, it should be noted that the above embodiments are merely examples for clearly illustrating this disclosure and are not intended to limit the implementation. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of this disclosure.
Claims
1. A method for locating a sound source, characterized in that, include: Time-domain framed data is obtained by performing frame-by-frame windowing processing on the time-domain data of the audio signal received by the microphone array. By performing a short-time Fourier transform on the time-domain framed data, frequency-domain framed data is obtained. Enhanced frequency domain frame data is reconstructed by suppressing noise in the frequency domain frame data frame by frame. Based on the enhanced frequency domain framed data, an adaptive beamforming algorithm is used to perform spectral peak search to estimate the direction of arrival of the sound source frame by frame. The frequency domain framing data includes frequency domain framing data from multiple channels. The enhanced frequency domain framing data, which reconstructs the frequency domain framing data by performing noise suppression on the frequency domain framing data frame by frame, includes: Calculate the estimated noise covariance of the frequency domain frame data of the first channel of the plurality of channels; Using the estimated noise covariance, the gain function of the frequency domain frame data of the first channel before and after noise suppression is calculated. Based on the amplitude values of the frequency domain segmented frame data of the multiple channels in the current frame and the gain function, the enhanced frequency domain segmented frame data of the current frame after noise suppression is reconstructed.
2. The sound source localization method according to claim 1, characterized in that, After performing frame-segmentation and windowing processing on the time-domain data of the audio signal received by the microphone array to obtain time-domain framed data, the sound source localization method further includes: Using speech activity detection technology, speech activity is detected frame by frame in the time-domain frame data to obtain the speech activity detection result of each frame of the time-domain frame data, wherein the presence of speech in each frame of the time-domain frame data is determined based on the speech activity detection result.
3. The sound source localization method according to claim 2, characterized in that, The calculation of the estimated noise covariance of the frequency domain frame data of the first channel in the plurality of channels includes: Based on the voice activity detection results, the frequency domain representation of the frequency domain frame data of the current frame of the first channel is determined; Based on the frequency domain representation of the frequency domain segmented frame data of the current frame of the first channel, calculate the estimated noise covariance of the frequency domain segmented frame data of the current frame of the first channel. Specifically, if there is speech in the current frame of the first channel, the estimated noise covariance of the frequency domain subframe data of the previous frame of the first channel is used as the estimated noise covariance of the frequency domain subframe data of the current frame of the first channel. If there is no speech in the current frame of the first channel, the weighted sum of the estimated noise covariance of the frequency domain subframe data of the previous frame of the first channel and the power of the frequency domain subframe data of the current frame of the first channel is used as the estimated noise covariance of the frequency domain subframe data of the current frame of the first channel.
4. The sound source localization method according to claim 3, characterized in that, The step of using the estimated noise covariance to calculate the gain function of the frequency domain frame data of the first channel before and after noise suppression includes: Based on the estimated noise covariance and the power of the frequency domain segmented frame data of the current frame of the first channel, the posterior signal-to-noise ratio of the frequency domain segmented frame data of the current frame of the first channel is calculated. Based on the decision-guided estimation method, the weighted average of the prior signal-to-noise ratio of the frequency domain segmented frame data of the previous frame of the first channel and the posterior signal-to-noise ratio of the frequency domain segmented frame data of the current frame of the first channel is used as the prior signal-to-noise ratio of the frequency domain segmented frame data of the current frame of the first channel. Based on the posterior and prior signal-to-noise ratios of the frequency domain segmented data of the current frame of the first channel, and using the minimum mean square error of the logarithmic spectrum amplitude estimation as the distortion metric, the gain function of the frequency domain segmented data of the current frame of the first channel before and after noise suppression is calculated.
5. The sound source localization method according to claim 3, characterized in that, The adaptive beamforming algorithm includes a minimum variance distortionless response beamforming algorithm. The step of using the adaptive beamforming algorithm to perform spectral peak search to estimate the direction of arrival of the sound source frame-by-frame based on the enhanced frequency domain frame data includes: The enhanced frequency domain framing data of the current frame of the multiple channels is used as the input of the minimum variance distortionless response beamforming algorithm. Under the condition of satisfying the constraints, the optimal weight vector of the current frame of the minimum variance distortionless response beamforming algorithm is solved by the Lagrange multiplier method. Based on the optimal weight vector of the current frame and the enhanced frequency domain framing data of the current frame of the multiple channels, the average power of the output beam of the minimum variance distortionless response beamforming algorithm in the current frame is calculated. Within the angular scanning range, a spectral peak search is performed on the average power of the output beam of the current frame, and the signal incident angle corresponding to the peak point is used as the estimated direction of arrival of the sound source in the current frame.
6. The sound source localization method according to claim 5, characterized in that, The constraints include the requirement that the signal in the incident direction should pass through completely, while suppressing interference sources and noise from other directions to the greatest extent possible.
7. The sound source localization method according to claim 5, characterized in that, After estimating the direction of arrival of the sound source frame by frame using an adaptive beamforming algorithm to perform spectral peak search based on the enhanced frequency domain framed data, the sound source localization method further includes: Based on the speech activity detection results, the estimated direction of arrival (DOA) of the sound source in the current frame is smoothed. If speech is present in the temporal framing data of the current frame, the estimated DOA of the sound source in the current frame is determined as the final DOA of the sound source in the current frame. If speech is not present in the temporal framing data of the current frame, the estimated DOA of the sound source in the previous frame is determined as the final DOA of the sound source in the current frame.
8. A sound source localization device, characterized in that, include: The frame-segmentation and windowing processing unit is configured to perform frame-segmentation and windowing processing on the time-domain data of the audio signal received by the microphone array to obtain time-domain frame data. The short-time Fourier transform unit is configured to obtain frequency-domain frame data by performing a short-time Fourier transform on the time-domain frame data. The frequency domain frame data enhancement unit is configured to reconstruct enhanced frequency domain frame data by performing noise suppression on the frequency domain frame data frame by frame. The direction of arrival estimation unit is configured to perform spectral peak search using an adaptive beamforming algorithm based on the enhanced frequency domain frame data to estimate the direction of arrival of the sound source frame by frame. The frequency domain framing data includes frequency domain framing data from multiple channels. The enhanced frequency domain framing data, which reconstructs the frequency domain framing data by performing noise suppression on the frequency domain framing data frame by frame, includes: Calculate the estimated noise covariance of the frequency domain frame data of the first channel of the plurality of channels; Using the estimated noise covariance, the gain function of the frequency domain frame data of the first channel before and after noise suppression is calculated. Based on the amplitude values of the frequency domain segmented frame data of the multiple channels in the current frame and the gain function, the enhanced frequency domain segmented frame data of the current frame after noise suppression is reconstructed.
9. The sound source localization device according to claim 8, characterized in that, The sound source localization device further includes: The voice activity detection unit is configured to use voice activity detection technology to perform voice activity detection on the time-domain frame data frame by frame, and obtain the voice activity detection result of the time-domain frame data for each frame, wherein the presence of voice in the time-domain frame data for each frame is determined based on the voice activity detection result.
10. A sound source localization system, characterized in that, include: The microphone array is configured to receive time-domain data of audio signals from the sound source; A sound source localization device is configured to perform the method according to any one of claims 1 to 7.
11. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method as described in any one of claims 1 to 7.
12. A storage medium, characterized in that, The storage medium stores a computer program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 7.