A sound source positioning method and device, electronic equipment and storage medium
By combining a four-element microphone array and an FIR bandpass filter bank with the GCC-PHAT-ργ algorithm, the problem of natural noise influence in outdoor sound source localization is solved, and accurate sound source localization is achieved at low signal-to-noise ratios and long distances.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU UNIVERSITY
- Filing Date
- 2023-07-19
- Publication Date
- 2026-04-28
AI Technical Summary
In outdoor sound source localization tasks, natural noise affects the sound source localization accuracy of microphone arrays. Existing algorithms have poor robustness in low signal-to-noise ratio and long-distance scenarios, making it difficult to accurately estimate time delay and locate sound sources.
By employing a four-element microphone array and an FIR bandpass filter bank, the microphone array signal is preprocessed, power spectral density analyzed, bandpass filtered, and time delay estimated. Combined with the GCC-PHAT-ργ algorithm, the main energy components are extracted and time delay is estimated, ultimately determining the azimuth angle of the sound source.
Under conditions of low signal-to-noise ratio and long distance, it improves the accuracy and robustness of sound source localization, effectively resists natural noise interference, and achieves accurate sound source localization.
Smart Images

Figure CN117031400B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and in particular to a sound source localization method, apparatus, electronic device, and storage medium. Background Technology
[0002] Passive acoustic monitoring provides valuable information for conservation biology and biogeography, serving as a crucial tool for collecting ecological data, locating wildlife, and monitoring conservation efforts. Audio files collected during acoustic monitoring can provide vital data for wildlife conservation or research on bird behavior and other activities during migration. New non-invasive technologies, such as microphone array-based passive acoustic monitoring, can acquire specific directional and locational information of bird communities over a given period of time.
[0003] Researchers have proposed numerous methods for locating sound sources of target species in the wild. However, the wild environment is highly complex, with various natural noises, obstructions, and multipath effects caused by sound wave reflection from dense vegetation, among many other factors, affecting the accuracy of sound source localization using microphone arrays. When recording sound signals in the wild, natural noise such as wind noise is an unavoidable problem, and this type of noise is one of the significant factors influencing the accuracy of sound source localization techniques using microphone arrays.
[0004] Currently, sound source localization in the wild still faces many challenges, and researchers have proposed numerous solutions. For example, high-resolution spectral estimation-based MUSIC localization algorithms are used to estimate temporally overlapping bird calls, and time delay difference (TDD)-based localization algorithms are employed to remotely monitor and even decipher elephant social interactions. However, TDD-based localization algorithms are highly sensitive to noise. Natural noise, although low in frequency, has high energy and occupies the majority of the time domain, causing cluttered peaks in the cross-correlation algorithms used to obtain time delay estimates, thus posing significant difficulties for time delay estimation. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a sound source localization method, apparatus, electronic device, and storage medium, which can accurately locate sound sources.
[0006] On one hand, embodiments of the present invention provide a sound source localization method, including:
[0007] The sound signals from multiple channels of the target sound source are acquired from the microphone array, and the sound signals are preprocessed to obtain the amplitude spectrum of each frame of each channel.
[0008] The power spectral density of the audio signal in each channel is determined based on the amplitude spectrum, and the target power spectral density is obtained through averaging.
[0009] The target power spectral density is logarithmically divided and first-order differencing is performed to obtain the target difference array. The local maxima and the corresponding frequency values of the target power spectral density are extracted from the target difference array. The local maxima are stored in the first array and the frequency values are stored in the second array.
[0010] Based on preset conditions, the target local maximum value and target frequency value are obtained from the first array and the second array, and then a bandpass filter bank is constructed.
[0011] The filtered audio signal is obtained by using a bandpass filter bank, and the time delay of the filtered signal is estimated to obtain the time delay estimation matrix.
[0012] Conditional operations are performed on the time delay estimation matrix to obtain the estimated parameters, and then the azimuth angle of the target sound source is determined based on the estimated parameters.
[0013] Optionally, the microphone array includes multiple microphone elements, which are distributed around the target sound source. The sound signal is preprocessed to obtain the amplitude spectrum of each frame of signal for each channel, including:
[0014] The number of frames is determined based on the preset frame length and frame shift;
[0015] By performing a framing operation, the audio signals of each channel are divided into frames to obtain frame signals; the number of frames in the frame signal corresponding to each channel is the number of frames.
[0016] The signal is windowed after framing based on a preset window length, and then a Discrete Fourier Transform is performed using the window length as the number of points for the Discrete Fourier Transform to obtain the amplitude spectrum of each frame signal.
[0017] Optionally, the power spectral density of the audio signal in each channel is determined based on the amplitude spectrum, and the target power spectral density is obtained through averaging, including:
[0018] Based on the amplitude spectrum of each frame of signal, the power spectral density of each frame of signal is calculated using the power spectral density formula.
[0019] The amplitude spectrum is obtained based on a discrete Fourier transform with a preset number of points; the power spectral density formula is as follows:
[0020]
[0021] In the formula, P m (k,l) represents the power spectral density of the l-th frame signal in the m-th channel; X m (k,l) represents the amplitude spectrum of the l-th frame signal in the m-th channel; k represents the frequency index of the signal; N F = N / 2 + 1, where N represents the number of points in the Discrete Fourier Transform;
[0022] The power spectral density of each frame of signal in each channel is accumulated and averaged to obtain the power spectral density of the audio signal in each channel, and the target power spectral density is obtained by averaging the power spectral densities of the audio signals in all channels.
[0023] Optionally, the target power spectral density is logarithmically calculated and first-order differencing is performed to obtain a target difference array. Local maxima and their corresponding frequencies are then extracted from the target power spectral density based on this array, including:
[0024] The logarithm of the target power spectral density is taken to obtain the decibel value;
[0025] Perform a first-order difference operation on the decibel value to obtain the first difference array;
[0026] Construct a second sign array based on the positive and negative values of the elements in the first difference array;
[0027] Perform a first-order difference operation on the second symbol array to obtain the target difference array;
[0028] Extract the indices of elements less than 0 from the target difference array, increment each element index by 1, and obtain the local maximum index;
[0029] Based on the local maximum index, all local maximum values of the target power spectral density are obtained and stored in the first array, and the frequency values corresponding to all local maximum values are obtained and stored in the second array; then, the frequency values in the second array that exceed the preset target frequency band are removed, and the corresponding local maximum values in the first array are also removed.
[0030] Optionally, based on preset conditions, the target local maximum value and target frequency value are obtained from the first array and the second array, and then a bandpass filter bank is constructed, including:
[0031] Get the maximum value in the first array, and iterate through all the local maximum values in the first array. Remove the local maximum values in the first array whose difference from the maximum value is greater than a preset threshold, and retain a preset number of local maximum values as the target local maximum value. Then, only retain the frequency values in the second array that correspond to the target local maximum value as the target frequency value.
[0032] The frequency values in the target frequency value are divided into frequency bands. The frequency values within the target frequency value are taken as the center frequency, and one-third octave of the center frequency is taken as the upper boundary frequency and the lower boundary frequency, which are then used as the passband boundary frequencies to construct the filter.
[0033] Based on all the constructed filters, a bandpass filter bank is obtained.
[0034] Optionally, a filtered audio signal is obtained through a bandpass filter bank, and a time delay estimation is performed on the filtered signal to obtain a time delay estimation matrix, including:
[0035] The filtered audio signal for each channel is obtained by using a bandpass filter bank;
[0036] Autocorrelation processing is performed on the filtered signal of each channel to obtain the autocorrelation result of the corresponding channel;
[0037] The autocorrelation results of each channel are subjected to discrete Fourier transform to obtain the autopower spectrum. The autopower spectra are then summed by taking the modulus to obtain the summation power spectrum. The peak-to-average power ratio of the summation power spectrum is then determined.
[0038] Logarithmic transformation of the peak-to-average ratio yields the joint logarithmic peak-to-average ratio;
[0039] Based on a preset threshold value, the joint logarithmic peak-to-mean ratio is conditionally determined to obtain weighting coefficients for different joint logarithmic peak-to-mean ratios.
[0040] Cross-correlation processing is performed on the filtered signals of each channel to obtain the cross-correlation results between the channels;
[0041] The cross-correlation results between the channels are subjected to discrete Fourier transform to obtain the cross-power spectrum, and then the coherence function of each cross-power spectrum is determined.
[0042] The sound signals from all channels are averaged, and the energy ratio is determined based on the average signal. Adjustment parameters are set according to the range of the energy ratio data.
[0043] The weighting function is obtained based on the cross-power spectrum, coherence function, and adjustment parameters;
[0044] Multiply the weighting coefficients and the weighting function, and then perform an inverse discrete Fourier transform to obtain the cross-correlation sequence;
[0045] Peak search is performed on the cross-correlation sequence to obtain the peak index value, which is then used as the index vector of the cross-correlation sequence.
[0046] Based on the peak index value, the time delay estimate value of each filtered signal is determined; then, the time delay estimate matrix is obtained.
[0047] Optionally, the time delay estimation matrix includes three rows of data. Conditional operations are performed on the time delay estimation matrix to obtain estimated parameters. Then, the azimuth angle of the target sound source is determined based on the estimated parameters, including:
[0048] The first estimated parameter is obtained by averaging the data in the first row of the time delay estimation matrix.
[0049] The second row of data in the time delay estimation matrix is subjected to mean or modulo operation and then minimum operation to obtain the second estimated parameter.
[0050] The third estimated parameter is obtained by averaging the data in the third row of the time delay estimation matrix.
[0051] The azimuth of the target sound source is calculated using the azimuth formula based on the first, second, and third estimated parameters.
[0052] The expression for the azimuth formula is as follows:
[0053]
[0054] In the formula, denoted by azimuth, arctan(·) represents the arctangent function, Dealy12 represents the first estimated parameter, Dealy13 represents the second estimated parameter, and Delay14 represents the third estimated parameter.
[0055] On the other hand, embodiments of the present invention provide a sound source localization device, comprising:
[0056] The first module is used to acquire sound signals from multiple channels of the target sound source from the microphone array, and to preprocess the sound signals to obtain the amplitude spectrum of each frame of each channel.
[0057] The second module is used to determine the power spectral density of the audio signal in each channel based on the amplitude spectrum, and to obtain the target power spectral density through averaging.
[0058] The third module performs logarithmic and first-order difference operations on the target power spectral density to obtain a target difference array. Based on this array, it extracts the local maxima and their corresponding frequencies. The local maxima are stored in the first array, and the frequencies are stored in the second array.
[0059] The fourth module is used to obtain the target local maximum value and target frequency value from the first array and the second array based on preset conditions, and then construct a bandpass filter bank.
[0060] The fifth module is used to obtain the filtered signal of the sound signal through the bandpass filter bank, estimate the time delay of the filtered signal, and obtain the time delay estimation matrix.
[0061] The sixth module is used to perform conditional operations on the time delay estimation matrix to obtain the estimated parameters, and then determine the azimuth angle of the target sound source based on the estimated parameters.
[0062] On the other hand, embodiments of the present invention provide an electronic device, including a processor and a memory;
[0063] Memory is used to store programs;
[0064] The processor executes the program as described above.
[0065] On the other hand, embodiments of the present invention provide a computer-readable storage medium storing a program that is executed by a processor to implement the method described above.
[0066] This invention also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned method.
[0067] This invention first acquires sound signals from multiple channels of a target sound source from a microphone array, preprocesses the sound signals to obtain the amplitude spectrum of each frame of each channel; determines the power spectral density of the sound signals in each channel based on the amplitude spectrum, and obtains the target power spectral density through averaging; performs logarithmic and first-order difference operations on the target power spectral density to obtain a target difference array, and extracts the local maxima and the corresponding frequency values of the target power spectral density based on the target difference array; wherein, the local maxima are stored in a first array, and the frequency values are stored in a second array; based on preset conditions, the target local maxima and target frequency values are obtained from the first and second arrays, and a bandpass filter bank is constructed; the filtered sound signal is obtained through the bandpass filter bank, the time delay of the filtered signal is estimated to obtain a time delay estimation matrix; conditional operations are performed on the time delay estimation matrix to obtain estimation parameters, and then the azimuth angle of the target sound source is determined based on the estimation parameters. This invention employs a sound source localization method based on a microphone array and filter bank. By analyzing the acquired sound signals, a bandpass filter bank is designed, and time delay estimation is introduced to determine the estimation parameters and thus the azimuth angle. This method can effectively estimate time delay and provide localization results in scenarios with low signal-to-noise ratios and long distances, and it exhibits high robustness and localization performance against natural noise in real-world scenarios. This invention can accurately locate sound sources. Attached Figure Description
[0068] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0069] Figure 1This is a flowchart illustrating a sound source localization method provided in an embodiment of the present invention;
[0070] Figure 2 This is a schematic diagram of the overall process of the sound source localization method provided in the embodiments of the present invention;
[0071] Figure 3 This is a schematic diagram of the microphone array provided in an embodiment of the present invention;
[0072] Figure 4 This is a schematic diagram of the structure of a sound source localization device provided in an embodiment of the present invention;
[0073] Figure 5 This is a schematic diagram of the frame of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0074] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0075] First, it should be noted that the relevant technologies include: localization algorithms based on high-resolution spectral estimation. The MUSIC algorithm is the most representative high-resolution spatial spectral estimation localization algorithm. It deals with stationary narrowband signals from mutually independent signal sources. After sampling the array signals, its covariance matrix is calculated, and then eigenvalue decomposition is performed. Based on the orthogonality between the decomposed sub-signals, the location of the target sound source is calculated by combining array signal processing and mathematical methods. Because this type of method is not subject to the Rayleigh constraint and has very high resolution, it is called a high-resolution spectral estimation localization method.
[0076] An energy-based localization algorithm for input acoustic signals is employed. This algorithm uses a microphone array to collect ambient noise as a silent segment, then collects a signal containing the sound source, and subtracts the two signals using spectral subtraction. Next, a mathematical model of sound energy attenuation is established to calculate the path loss index of each element in the microphone array under the current environment. A weighted average of the path loss indices for each element is then used to obtain the path loss index for that environment. Finally, the location of the sound source is determined based on the energy and attenuation index of the signals received by different elements.
[0077] Deep learning-based localization algorithms primarily utilize deep neural networks to learn the characteristics of sound source signals. By constructing network models with different frameworks, deep neural networks can automatically learn the inherent patterns of sound source signals, enabling more accurate sound source localization. Typically, multi-channel input signals collected by a microphone array are used as input data for the localization model. A feature extraction module automatically extracts features from the input signals, which are then fed into DNNs (Deep Neural Networks). The DNNs provide estimates of the sound source's direction or location.
[0078] However, the related technologies have the following drawbacks: The positioning algorithm based on high-resolution spectral estimation deals with stationary narrowband signals from mutually independent sources, resulting in high computational complexity. When extending to broadband signal estimation, the non-overlapping signal space within the frequency band needs to be transformed and focused to a reference frequency point to obtain the data covariance matrix at that frequency point. Then, narrowband processing methods are used for parameter estimation, requiring multiple iterations to obtain a relatively accurate estimation result, further increasing the computational load. Furthermore, this method has poor robustness to noise, thus limiting its application in practical scenarios.
[0079] The disadvantages of energy localization algorithms based on input acoustic signals are: existing acoustic energy attenuation models are not accurate, and the attenuation coefficient conditions of the established acoustic model need to meet some idealized assumptions; the attenuation coefficient is usually set to a fixed value, but in reality, the value of the attenuation coefficient will be different when the sound source type and distance are different; in addition, the localization accuracy will also be affected when the algorithm is affected by various environmental noises or when the sampling data is too small.
[0080] The drawbacks of deep learning-based localization algorithms include: Currently, deep learning-based detection algorithms require massive amounts of training data and have high dataset requirements. When the amount of available training data is insufficient, the model typically fails to achieve good generalization performance, resulting in poor final model performance. Furthermore, deep learning algorithms are computationally intensive, demanding on device memory and have long training times, placing high demands on the performance of computing devices. This makes it difficult to guarantee the normal operation of the running equipment in real-world field scenarios.
[0081] In view of this, on the one hand, such as Figure 1 As shown, an embodiment of the present invention provides a sound source localization method, including:
[0082] S100: Acquire sound signals from multiple channels of the target sound source from the microphone array, and preprocess the sound signals to obtain the amplitude spectrum of each frame of each channel.
[0083] It should be noted that the microphone array includes multiple microphone elements, which are distributed around the target sound source. In some embodiments, the preprocessing of the sound signal to obtain the amplitude spectrum of each frame signal of each channel may include: determining the number of frames based on a preset frame length and frame shift; performing frame processing on the sound signal of each channel to obtain frame signals; wherein the number of frames of the frame signal corresponding to each channel is the number of frames; and windowing each frame signal after frame division based on a preset window length, and then performing a Discrete Fourier Transform using the window length as the number of discrete Fourier Transform points to obtain the amplitude spectrum of each frame signal.
[0084] In some specific embodiments, such as Figure 2 As shown, the first step is to preprocess the acquired signal, as follows:
[0085] 1. A microphone array with four elements can be used to collect the sound signal emitted by the sound source S. The frequency range of the signal is denoted as (f... low ,f high Microphone array structure as follows: Figure 3 As shown, Mic1 represents channel one microphone, Mic2 represents channel two microphone, Mic3 represents channel three microphone, and Mic4 represents channel four microphone. Let d be the azimuth angle of the sound source, and d be the distance from the array element to the origin O.
[0086] 2. The signals collected by each microphone are sampled and recorded as x. m (n), with length Len (determined based on the characteristics of the sound source S), and Fs is the signal sampling frequency. m represents the microphone number (1≤m≤4), and n represents the signal sampling point number (0≤n≤Len-1).
[0087] 3. For x m (n) Perform framing operation. Set the frame length to FrameSize and the frame shift to Inc. Obtain the number of frames FrameNums from equation (1). Each frame signal after framing is denoted as x. m (n',l), where n' is the sampling point number (0≤n'≤FrameSize-1), and l is the frame number (l=1,…,FrameNums). Equation (1):
[0088]
[0089] 4. Apply a Hamming window function with a window size of FrameSize to x. m Windowing is applied to (n',l), denoted as x'. m (n',l).
[0090] 5. Set N = FrameSize to the number of points in the DFT (Discrete Fourier Transform), and then calculate x'. m The modulus value of the N-point DFT operation result of (n', l) is used to obtain the amplitude spectrum of each frame of the signal, denoted as X. m (k,l), k=0,1,2,…,N F -1 represents the frequency point number, N F = N / 2 + 1. X m The calculation process of (k,l) is shown in equation (2), where abs(·) represents the modulo operation and -j represents the imaginary unit. Equation (2):
[0091]
[0092] S200: Determine the power spectral density of the sound signal in each channel based on the amplitude spectrum, and obtain the target power spectral density through averaging.
[0093] It should be noted that in some embodiments, step S200 may include: calculating the power spectral density of each frame of signal based on the amplitude spectrum of each frame using the power spectral density formula; wherein, the amplitude spectrum is obtained based on a discrete Fourier transform with a preset number of discrete Fourier transform points; the expression of the power spectral density formula is:
[0094]
[0095] In the formula, P m (k,l) represents the power spectral density of the l-th frame signal in the m-th channel; X m (k,l) represents the amplitude spectrum of the l-th frame signal in the m-th channel; k represents the frequency index of the signal; N F = N / 2 + 1, where N represents the number of points in the Discrete Fourier Transform;
[0096] The power spectral density of each frame of signal in each channel is accumulated and averaged to obtain the power spectral density of the audio signal in each channel, and the target power spectral density is obtained by averaging the power spectral densities of the audio signals in all channels.
[0097] In some specific embodiments, such as Figure 2 As shown, power spectral density estimation is performed as follows:
[0098] For each frame of signal, the power spectral density is calculated, and the amplitude spectrum X obtained from equation (2) is used. m Substituting (k,l) into equation (3), calculate the power spectral density P of each frame of the signal. m (k,l):
[0099]
[0100] The power spectral density P of the signal is obtained. m (k), as shown in the following formula:
[0101]
[0102] Power spectral density P of the four channels m (k) Take the average, as shown in the following formula:
[0103]
[0104] S300. Perform logarithmic and first-order difference operations on the target power spectral density to obtain the target difference array. Extract the local maximum value and the frequency value corresponding to the local maximum value of the target power spectral density based on the target difference array.
[0105] Local maxima are stored in the first array, and frequency values are stored in the second array.
[0106] It should be noted that in some embodiments, step S300 may include: taking the logarithm of the target power spectral density to obtain a decibel value; performing a first-order difference operation on the decibel value to obtain a first difference array; constructing a second symbol array based on the positive and negative values of the elements in the first difference array; performing a first-order difference operation on the second symbol array to obtain a target difference array; extracting the indices of elements less than 0 from the target difference array, incrementing each element index by 1 to obtain a local maximum index; storing all local maximum values of the target power spectral density in the first array based on the local maximum index, and obtaining the frequency values corresponding to all local maximum values in the second array; and then removing frequency values in the second array that exceed the preset target frequency band, and removing the corresponding local maximum values in the first array.
[0107] In some specific embodiments, such as Figure 2 As shown, the local maxima and their corresponding frequency values are extracted as follows:
[0108] The logarithm of P(k) is taken to convert it into a decibel value, as shown in equation (6).
[0109] y(k) = 10log 10 |P(k)|, k=0,1,…,N F -1 (1)
[0110] Perform a first-order difference operation on y(k) to obtain a difference array a1 of length NF-1.
[0111] a1(s) = y(k+1) - y(k), 0 ≤ s ≤ N F -twenty two)
[0112] Determine the sign of each element in the difference array a1 and generate a sequence of length N. F The symbolic array a2 is a set of elements -1. If an element of a1 is positive, the corresponding element of a2 is 1; if an element of a1 is 0, the corresponding element of a2 is 0; if an element of a1 is negative, the corresponding element of array a2 is -1.
[0113] Performing a first-order difference operation on the symbol array a2 yields a result of length N. F Find the difference array a3 of -2. Then find the indices of the elements in a3 that are less than 0, and increment each index by 1 to get the indices of all local maxima.
[0114] a3(z)=a2(k+1)-a2(k),0≤z≤N F -3 (3)
[0115] All local maxima are obtained by indexing and stored in the array Locmax. The frequency values corresponding to the local maxima are stored in the array freqs.
[0116] For each frequency value in the array freqs, determine whether it is within the target frequency band (f low ,f high If it exists, keep it; otherwise, remove it from freqs, and also remove the corresponding local maximum value in Locmax.
[0117] S400: Based on preset conditions, obtain the target local maximum value and target frequency value from the first array and the second array, and then construct a bandpass filter bank;
[0118] It should be noted that in some embodiments, step S400 may include: obtaining the maximum value in the first array, and traversing all local maximum values in the first array, removing local maximum values in the first array whose difference from the maximum value is greater than a preset threshold, and retaining a preset number of local maximum values as target local maximum values; then retaining only the frequency values in the second array corresponding to the target local maximum values as target frequency values; dividing the frequency values in the target frequency values into frequency bands, taking the frequency values within the target frequency values as center frequencies, taking one-third octaves of the center frequencies as upper boundary frequencies and lower boundary frequencies, and then using them as passband boundary frequencies to construct filters; and obtaining a bandpass filter bank based on all constructed filters.
[0119] In some specific embodiments, such as Figure 2 As shown, an FIR bandpass filter bank is designed based on local maxima, and implemented as follows:
[0120] Find the maximum value in the array Locmax to get the maximum value maxVal; traverse the array Locmax and keep only the local maximum values within 15dB of maxVal. If there are more than 4 local maximum values, keep only the first 4. Also keep only the corresponding frequency values in freqs.
[0121] Divide the frequency values in the freqs array into frequency bands. Then, use each frequency value in the freqs array as a center frequency f. c Take one-third of its octave band as the upper boundary frequency f. h With the lower boundary frequency f l As shown in equations (9) and (10):
[0122]
[0123]
[0124] Design an FIR bandpass filter bank. Assume the `freqs` array contains `I` elements, where `I` has a maximum value of 4. The filter bank consists of `I` bandpass filters with the following parameters: order N / 2, window function type (hamming), and the `H` parameter for the `i`th (1 ≤ i ≤ I) filter. i The passband boundary frequency is calculated by equations (9) and (10): Lower boundary frequency f l (i), upper boundary frequency f h (i).
[0125] S500: Obtain the filtered audio signal through a bandpass filter bank, estimate the time delay of the filtered signal, and obtain the time delay estimation matrix.
[0126] It should be noted that, in some embodiments, step S500 may include: obtaining the filtered signal of the audio signal for each channel through a bandpass filter bank; performing autocorrelation processing on the filtered signal of each channel to obtain the autocorrelation result of the corresponding channel; performing discrete Fourier transform on the autocorrelation result of each channel to obtain the autopower spectrum, and then performing modulo summation on the autopower spectrum to obtain the summed power spectrum; and determining the peak-to-average power ratio (PAPR) of the summed power spectrum; performing a logarithmic transform on the PAPR to obtain the joint logarithmic PAPR; performing conditional judgment on the joint logarithmic PAPR based on a preset threshold value to obtain the weighting coefficients of different joint logarithmic PAPRs; and performing cross-correlation processing on the filtered signals of each channel to obtain the signal for each channel. The process involves: obtaining cross-correlation results between channels; performing a Discrete Fourier Transform (DFT) on the cross-correlation results of each channel to obtain the cross-power spectrum, and then determining the coherence function of each cross-power spectrum; averaging the audio signals of all channels to determine the energy ratio based on the average signal; setting adjustment parameters based on the data range corresponding to the energy ratio; obtaining a weighting function based on the cross-power spectrum, coherence function, and adjustment parameters; multiplying the weighting coefficients and weighting function, and then performing an Inverse Discrete Fourier Transform (IFT) to obtain the cross-correlation sequence; performing a peak search on the cross-correlation sequence to obtain the peak index value as the index vector of the cross-correlation sequence; determining the time delay estimate value of each filtered signal based on the peak index value; and finally, assembling the results to obtain the time delay estimation matrix.
[0127] In some specific embodiments, such as Figure 2 As shown, the time delay estimation is implemented as follows:
[0128] Use a filter bank to sample the signal x m After filtering (n)(1≤m≤4), let it be denoted as x. (i,m) (n), where i is the filter index (1≤i≤I).
[0129] Calculate the autocorrelation of each filtered channel, m = 1, 2, 3, 4; τ represents the number of lag points, -(Len-1)≤τ≤(Len-1), when τ<0, 0≤n≤Len+τ-1; when τ≥0, τ≤n≤Len-1.
[0130]
[0131] Perform a DFT on the autocorrelation results to obtain the autopower spectrum. After taking the modulus of the power spectra of each channel, sum them to obtain the summation power spectrum Asum. (i) (w), 0≤w≤2*Len-1.
[0132]
[0133] Calculate the peak-to-average power ratio (PAR) of the summation power spectrum. (i) (w).
[0134]
[0135] Take the base-10 logarithm transformation to the logarithmic domain to obtain the joint logarithmic peak-to-average power ratio \(P_{log}\). (i) (\(\omega\)).
[0136] \(P_{log}\) (i) (\(\omega\)) = log 10 (PAR (i) (\(\omega\))) (9)
[0137] Calculate the weighting coefficient \(W\) (i) (\(\omega\)), where \(a\) is the threshold for suppressing noise, with a value of half of the average value of \(P_{log}\) (i) (\(\omega\)), \(c\) is the bias coefficient (\(c\geq1\)), and \(g\) is the given gain coefficient, usually taken as half of \(c\).
[0138]
[0139] Calculate the cross-correlation \(R\) (i,1q) (\(\tau\)) between channel 1 and channel \(q\), where \(q = 2, 3, 4\); \(R\) (i,12) (\(\tau\)) is the cross-correlation between channel 1 and channel 2, \(R\) (i,13) (\(\tau\)) is the cross-correlation between channel 1 and channel 3, \(R\) (i,14) (\(\tau\)) is the cross-correlation between channel 1 and channel 4. When \(\tau < 0\), \(0\leq n\leq Len+\tau - 1\); when \(\tau\geq0\), \(\tau\leq n\leq Len - 1\).
[0140]
[0141]
[0142]
[0143] Perform DFT on the cross-correlation result to obtain the cross-power spectrum where \(q = 2, 3, 4\).
[0144] Calculate the coherence function \(\gamma\) (i,1q) (\(\omega\)), where \(q = 2, 3, 4\).
[0145]
[0146] Calculate the average of the four channels of \(x\) m (\(n\)) to obtain Calculate its energy ratio \(ER\), and use \(ER\) to set the adjustment parameter \(\rho\). When \(ER\leq0\), \(0.7 < \rho\leq0.9\), it is recommended to take \(\rho = 0.8\); when \(0 < ER\leq2\), \(0.5 < \rho\leq0.7\), it is recommended to take \(\rho = 0.6\); when \(2 < ER\), \(0.2 < \rho\leq0.5\), it is recommended to take \(\rho = 0.5\).
[0147]
[0148] Calculate the weighting function ψ (i,1q) (w), q = 2, 3, 4.
[0149]
[0150] The weighting coefficient W (i) (w) and weighting function ψ (i,1q) Multiply (w) together, then perform IDFT to obtain the final cross-correlation sequence Rcc. (i,1q) (τ), q = 2, 3, 4.
[0151] Rcc (i,1q) (τ)=IDFT[W (i) (w)*ψ (i,1q) (w)] (17)
[0152] For Rcc (i,1q) (τ) Perform peak search to obtain the corresponding peak index value Lag (i,1q) Lag (i,1q) This is the index vector of the cross-correlation sequence, with values ranging from -(Len-1) to (Len-1).
[0153] Use Lag (i,1q) Divide by the sampling rate Fs to obtain the signal x (i,1) (n) and x (i,q) (n) time delay (i,1q) , q = 2, 3, 4.
[0154]
[0155] In summary, cross-correlation of the signals filtered by each filter Hi yields three time delay estimates, which are delay values. (i,12) delay (i,13) delay (i,14) (1≤i≤I), write it as a column vector and arrange it according to the filter order; the time delay estimation matrix Delay with 3 rows and i columns (1≤i≤I) is obtained as shown in the following formula:
[0156]
[0157] S600. Perform conditional operations on the time delay estimation matrix to obtain the estimation parameters, and then determine the azimuth angle of the target sound source based on the estimation parameters.
[0158] The raw data includes information about the target object, behavioral data, and item ratings;
[0159] It should be noted that, in some embodiments, step S600 may include: performing a mean operation on the first row of data in the time delay estimation matrix to obtain a first estimated parameter; performing a mean operation or modulo operation on the second row of data in the time delay estimation matrix and performing a minimum value operation to obtain a second estimated parameter; performing a mean operation on the third row of data in the time delay estimation matrix to obtain a third estimated parameter; and calculating the azimuth angle of the target sound source using the azimuth angle formula based on the first estimated parameter, the second estimated parameter, and the third estimated parameter.
[0160] The expression for the azimuth formula is as follows:
[0161]
[0162] In the formula, denoted as azimuth, arctan(·) represents the arctangent function, Delay12 represents the first estimated parameter, Delay13 represents the second estimated parameter, and Delay14 represents the third estimated parameter.
[0163] In some specific embodiments, such as Figure 2 As shown, firstly, based on the conditions, the corresponding operations are performed on the Delay matrix. The first row of the delay estimation matrix Delay, containing all columns Delay(1,:), is averaged to obtain Delay12. The second row, containing all columns Delay(2,:), is then averaged (either by taking the minimum value) to obtain Delay13. The third row, containing all columns Delay(3,:), is averaged to obtain Delay14. mean(·) is the mean operation, min(·) is the minimum value operation, and abs(·) is the modulo operation, as shown in the following formula:
[0164] Delay12=mean(Delay(1,:)) (20)
[0165]
[0166] Delay14=mean(Delay(3,:)) (22)
[0167] Finally, the azimuth angle is estimated. The calculation is shown in the following formula:
[0168]
[0169] In summary, by introducing an FIR bandpass filter bank into the localization algorithm based on arrival delay difference, a sound source localization method based on a four-element microphone array and an FIR filter bank is proposed, which exhibits high robustness to natural wind noise. Considering the significant impact of natural noise on the time delay estimation of the cross-correlation algorithm, an FIR bandpass filter bank is introduced. Furthermore, the location of the main energy distribution is found based on the local maxima of the power spectral density curve of the sampled signal, and different frequency bands are divided at the main energy locations to design the FIR bandpass filter bank. The effective signal within a specific frequency band is extracted using the FIR bandpass filter bank, and a GCC-PHAT-ργ algorithm based on logarithmic peak-to-average power ratio (PAPR) weighting coefficient is proposed. This algorithm can effectively estimate the time delay between signals of different array elements of the microphone array under low signal-to-noise ratio and long-distance conditions, ultimately estimating the accurate azimuth angle of the sound source. Experiments verify that this method has accurate localization performance in scenarios with natural noise interference. This addresses the problem of low localization accuracy or inability to estimate the sound source direction caused by the inability to obtain time delay estimation due to natural noise or long distance when using microphone array localization algorithms based on time delay estimation in field environments. This invention extracts the main energy component of the acquired signal by filtering out local maxima in the power spectral density curve; designs an FIR filter bank to filter the acquired data; and proposes a GCC-PHAT-ργ algorithm based on logarithmic peak-to-average power ratio weighting coefficients, effectively improving the accuracy and robustness of the cross-correlation algorithm based on time delay estimation, and realizing sound source localization under low signal-to-noise ratio conditions. The beneficial effects of this invention include at least:
[0170] 1. In response to the low signal-to-noise ratio of the acquired signal due to the presence of natural noise in the microphone array, this invention proposes a sound source localization method based on a four-element microphone array and an FIR filter bank, which can achieve sound source localization under low signal-to-noise ratio conditions.
[0171] 2. This invention proposes a GCC-PHAT-ργ algorithm based on logarithmic peak-to-average ratio weighting coefficients, which can effectively improve the estimation error of cross-correlation algorithms for long-distance signals.
[0172] 3. The relationship between the adjustment parameter ρ and the energy ratio is given, and suggestions are provided for the value of the adjustment parameter ρ under different energy ratios.
[0173] 4. The frequency band division method based on the local maximum value of the power spectral density curve proposed in this invention can extract the main energy part of the audio signal under the condition of noise interference and successfully estimate the time delay from these effective data.
[0174] On the other hand, such as Figure 4As shown, an embodiment of the present invention provides a sound source localization device 700, comprising: a first module 710, used to acquire sound signals from multiple channels of a target sound source from a microphone array, and preprocess the sound signals to obtain the amplitude spectrum of each frame of each channel; a second module 720, used to determine the power spectral density of the sound signals of each channel based on the amplitude spectrum, and obtain the target power spectral density through averaging; a third module 730, used to perform logarithmic operation and first-order difference operation on the target power spectral density to obtain a target difference array, and extract the local maximum value and the frequency value corresponding to the local maximum value of the target power spectral density according to the target difference array; wherein, the local maximum value is stored in the first array, and the frequency value is stored in the second array; a fourth module 740, used to acquire the target local maximum value and the target frequency value from the first array and the second array based on preset conditions, and then construct a bandpass filter bank; a fifth module 750, used to obtain the filtered signal of the sound signal through the bandpass filter bank, perform time delay estimation on the filtered signal, and obtain a time delay estimation matrix; and a sixth module 760, used to perform conditional operation on the time delay estimation matrix to obtain estimation parameters, and then determine the azimuth angle of the target sound source according to the estimation parameters.
[0175] The content of the method embodiments of the present invention is applicable to the device embodiments. The specific functions implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above methods.
[0176] like Figure 5 As shown, another aspect of the present invention provides an electronic device 800, including a processor 810 and a memory 820;
[0177] Memory 820 is used to store programs;
[0178] The processor 810 executes programs as described above.
[0179] The content of the method embodiments of the present invention is applicable to the embodiments of the present electronic device. The specific functions implemented by the embodiments of the present electronic device are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above methods.
[0180] Another aspect of this invention provides a computer-readable storage medium storing a program that is executed by a processor to implement the method described above.
[0181] The content of the method embodiments of the present invention is applicable to the computer-readable storage medium embodiments. The specific functions implemented by the computer-readable storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above methods.
[0182] This invention also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned method.
[0183] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this invention are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.
[0184] Furthermore, although the invention has been described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the invention. Rather, given the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of conventional skill of an engineer. Therefore, those skilled in the art can implement the invention as set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of the invention, which is determined by the full scope of the appended claims and their equivalents.
[0185] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0186] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution means, apparatus, or device (such as a computer-based device, a processor-including device, or other means that can fetch and execute instructions from, or in conjunction with, an instruction execution means, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution means, apparatus, or device.
[0187] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0188] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution device. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0189] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0190] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
[0191] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of the present invention.
Claims
1. A method for locating a sound source, characterized in that, include: The sound signals of multiple channels of the target sound source are acquired from the microphone array, and the sound signals are preprocessed to obtain the amplitude spectrum of each frame of the signal of each channel. The power spectral density of the sound signal in each of the channels is determined based on the amplitude spectrum, and the target power spectral density is obtained by averaging. The target power spectral density is logarithmically divided and first-order differencing is performed to obtain a target difference array. Local maxima and corresponding frequency values of the target power spectral density are extracted from the target difference array. The local maxima are stored in a first array and the frequency values are stored in a second array. Based on preset conditions, the target local maximum value and target frequency value are obtained from the first array and the second array, and then a bandpass filter bank is constructed. The filtered signal of the sound signal is obtained through the bandpass filter bank, and the time delay of the filtered signal is estimated to obtain the time delay estimation matrix. Conditional operations are performed on the time delay estimation matrix to obtain estimation parameters, and then the azimuth angle of the target sound source is determined based on the estimation parameters.
2. The sound source localization method according to claim 1, characterized in that, The microphone array includes multiple microphone elements, which are distributed around the target sound source. The preprocessing of the sound signal to obtain the amplitude spectrum of each frame of each channel includes: The number of frames is determined based on the preset frame length and frame shift; By performing a framing operation, the audio signals of each channel are processed into frames to obtain frame signals; wherein, the number of frames in the frame signal corresponding to each channel is the number of frames in the framing operation. Based on a preset window length, each frame of the signal after framing is windowed, and then the discrete Fourier transform is performed using the window length as the number of discrete Fourier transform points to obtain the amplitude spectrum of each frame of the signal.
3. The sound source localization method according to claim 1, characterized in that, The step of determining the power spectral density of the audio signal in each of the channels based on the amplitude spectrum, and obtaining the target power spectral density through averaging, includes: Based on the amplitude spectrum of each frame of signal, the power spectral density of each frame of signal is calculated using the power spectral density formula. The amplitude spectrum is obtained based on a discrete Fourier transform with a preset number of discrete Fourier transform points; the expression for the power spectral density formula is: In the formula, Indicates the first The first channel Power spectral density of the frame signal; Indicates the first The first channel The amplitude spectrum of the frame signal; Indicates the frequency point number of the signal; , This represents the number of points in the Discrete Fourier Transform. The power spectral density of each frame signal in each channel is accumulated and averaged to obtain the power spectral density of the audio signal in each channel, and the target power spectral density is obtained by averaging the power spectral densities of the audio signals in all channels.
4. The sound source localization method according to claim 1, characterized in that, The process of performing a logarithmic operation and a first-order difference operation on the target power spectral density to obtain a target difference array, and extracting the local maxima of the target power spectral density and the frequency values corresponding to the local maxima based on the target difference array, includes: The target power spectral density is logarithmically calculated to obtain a decibel value. Perform a first-order difference operation on the decibel value to obtain the first difference array; Based on the positive and negative values of the elements in the first difference array, a second sign array is constructed; Perform the first-order difference operation on the second symbol array to obtain the target difference array; Extract the indices of elements less than 0 from the target difference array, and increment each element index by 1 to obtain the local maximum index; Based on the local maximum index, all local maximum values of the target power spectral density are indexed and stored in the first array, and the frequency values corresponding to all local maximum values are obtained and stored in the second array; then, the frequency values in the second array that exceed the preset target frequency band are removed, and the corresponding local maximum values in the first array are removed.
5. The sound source localization method according to claim 1, characterized in that, The step of obtaining the target local maximum value and target frequency value from the first array and the second array based on preset conditions, and then constructing a bandpass filter bank, includes: Obtain the maximum value in the first array, and iterate through all the local maximum values in the first array. Remove the local maximum values in the first array whose difference from the maximum value is greater than a preset threshold, and retain a preset number of local maximum values as target local maximum values. Then, retain only the frequency values in the second array that correspond to the target local maximum values as target frequency values. The frequency values in the target frequency value are divided into frequency bands. The frequency values within the target frequency value are taken as the center frequency, and one-third octave of the center frequency is taken as the upper boundary frequency and the lower boundary frequency, which are then used as the passband boundary frequencies to construct a filter. Based on all the constructed filters, a bandpass filter bank is obtained.
6. The sound source localization method according to claim 1, characterized in that, The process of obtaining the filtered audio signal through the bandpass filter bank, and estimating the time delay of the filtered signal to obtain a time delay estimation matrix includes: The filtered signal of the audio signal for each of the channels is obtained through the bandpass filter bank; Autocorrelation processing is performed on the filtered signal of each channel to obtain the autocorrelation result for the corresponding channel; The autocorrelation results of each channel are subjected to discrete Fourier transform to obtain the autopower spectrum. The autopower spectra are then summed by taking the modulus to obtain the summation power spectrum. The peak-to-average power ratio of the summation power spectrum is then determined. Perform a logarithmic transformation on the peak-to-average power ratio (PAPR) to obtain the joint logarithmic PAPR. Based on a preset threshold value, the joint logarithmic peak-to-average ratio is conditionally determined to obtain weighting coefficients for different joint logarithmic peak-to-average ratios; The filtered signals of each channel are cross-correlated to obtain the cross-correlation results between the channels. The cross-correlation results between the channels are subjected to discrete Fourier transform to obtain the cross-power spectrum, and then the coherence function of each cross-power spectrum is determined. The sound signals from all the channels are averaged, and the energy ratio is determined based on the average signal; adjustment parameters are set according to the range of the energy ratio data. The weighting function is obtained based on the cross-power spectrum, the coherence function, and the adjustment parameters; Multiply the weighting coefficients and the weighting function, and then perform an inverse discrete Fourier transform to obtain a cross-correlation sequence; A peak search is performed on the cross-correlation sequence to obtain the peak index value, which is then used as the index vector of the cross-correlation sequence. Based on the peak index value, the time delay estimate value of each of the filtered signals is determined; then, the time delay estimate matrix is obtained.
7. The sound source localization method according to claim 1, characterized in that, The time delay estimation matrix includes three rows of data. The step of performing conditional operations on the time delay estimation matrix to obtain estimation parameters, and then determining the azimuth angle of the target sound source based on the estimation parameters, includes: The first estimated parameter is obtained by averaging the data in the first row of the time delay estimation matrix. The second row of data in the time delay estimation matrix is subjected to mean or modulo operation and then to minimum value operation to obtain the second estimated parameter; The third estimated parameter is obtained by averaging the data in the third row of the time delay estimation matrix. The azimuth of the target sound source is calculated using the azimuth formula based on the first estimated parameter, the second estimated parameter, and the third estimated parameter. The expression for the azimuth formula is as follows: In the formula, Indicates azimuth. Represents the arctangent function. Indicates the first estimated parameter. This represents the second estimated parameter. This represents the third estimated parameter.
8. A sound source localization device, characterized in that, include: The first module is used to acquire sound signals from multiple channels of the target sound source from the microphone array, and to preprocess the sound signals to obtain the amplitude spectrum of each frame of the signal of each channel. The second module is used to determine the power spectral density of the sound signal in each of the channels based on the amplitude spectrum, and to obtain the target power spectral density through averaging. The third module is used to perform logarithmic and first-order difference operations on the target power spectral density to obtain a target difference array, and to extract the local maximum value and the frequency value corresponding to the local maximum value of the target power spectral density based on the target difference array; wherein, the local maximum value is stored in a first array and the frequency value is stored in a second array; The fourth module is used to obtain the target local maximum value and the target frequency value from the first array and the second array based on preset conditions, and then construct a bandpass filter bank; The fifth module is used to obtain a filtered signal of the sound signal through the bandpass filter bank, perform time delay estimation on the filtered signal, and obtain a time delay estimation matrix; The sixth module is used to perform conditional operations on the time delay estimation matrix to obtain estimation parameters, and then determine the azimuth angle of the target sound source based on the estimation parameters.
9. An electronic device, characterized in that, Including the processor and memory; The memory is used to store programs; The processor executes the program to implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a program that is executed by a processor to implement the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Line spectrum extraction method based on target characteristics
CN111665489A
Birdsong recognition method and system based on spatial orientation, computer equipment and medium
CN113314127A