A sound source localization method based on SRP-PHAT spatial spectrum and GCC
By combining GCC and SRP-PHAT algorithms, the sound source angle range is determined and accurate search is carried out therein, and the problems of large calculation and low positioning accuracy in the prior art are solved, and high-precision sound source positioning is achieved.
Patent Information
- Application Number
- CN202211654433.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-22
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-12-22
AI Technical Summary
While existing sound source positioning technologies reduce the calculation amount, it is difficult to ensure the accuracy of sound source positioning.
The sound source positioning method based on SRP-PHAT spatial spectrum and GCC is adopted, and the time delay between the microphones is calculated through the GCC algorithm, the sound source angle range is determined, and then the SRP-PHAT algorithm is used to accurately search within this range to determine the sound source direction.
While reducing the calculation amount, the accuracy of sound source positioning is improved and the requirements for sampling frequency and array element spacing are reduced.
Smart Images

Figure CN115951305B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sound source localization technology, and in particular to a sound source localization method based on SRP-PHAT (Steerable Response Power-Phase Transformation) spatial spectrum and GCC (Generalized Cross-Correlation). Background Art
[0002] Sound source localization technology has broad application prospects in speech recognition, intelligent voice systems, front-end processing of human-computer interaction, video conferencing, smart homes and other fields. Traditional sound source localization technology is based on microphone array technology, which uses data collected by different microphones to estimate the direction of sound. There are three common DOA algorithms:
[0003] 1. Based on the time delay of arrival (TDOA) of sound: Since the delay of the sound source signal received by each array element on the array is different, we can estimate the delay difference between each array element through generalized cross-correlation (GCC), and then determine the direction of the sound source by combining the geometric array relationship. This method has low computational complexity and strong noise resistance, but has high requirements for sampling accuracy and array element spacing.
[0004] 2. Beamforming-based algorithms: This type of algorithm performs angle compensation phase on each element in the array, and then performs weighted summation on each signal. The direction with the maximum beam output power is the direction of the target sound source. Common beamforming algorithms include the Steerable Response Power-Phase Transformation (SRP-PHAT) algorithm and the Minimum Variance Distortionless Response (MVDR) algorithm. Since this type of method uses angle compensation to scan the direction with the maximum output power, when there are high requirements for positioning accuracy, the scanning resolution will increase the amount of calculation, and the algorithm's anti-noise ability is weak.
[0005] 3. Estimation based on high-resolution spectrum: The direction of the sound source is estimated by calculating the correlation matrix of the spatial spectrum by acquiring the signal of the microphone array, such as the eigenvalue decomposition method (MUSIC algorithm). This type of algorithm has high accuracy, but it involves matrix calculation and requires a lot of calculation, and is sensitive to environmental noise. It is usually used for narrowband signals and single-frequency signals.
[0006] Therefore, how to ensure the accuracy of sound source positioning while reducing the amount of calculation is a technical problem that needs to be solved urgently in the industry. Summary of the invention
[0007] The technical problem to be solved by the present invention is to propose a sound source localization method based on SRP-PHAT spatial spectrum and GCC, which can reduce the amount of calculation and ensure the accuracy of sound source localization.
[0008] The technical solution adopted by the present invention to solve the above technical problems is:
[0009] A sound source localization method based on SRP-PHAT spatial spectrum and GCC locates the direction angle of the sound source based on a microphone array structure. The method comprises the following steps:
[0010] S1. According to the observation signal collected by the microphone in the microphone array structure, the time delay between other microphones and the reference microphone is calculated by the GCC algorithm;
[0011] S2. Determine the range of the sound source angle based on the time delay calculated in step S1 according to the geometric structure of the microphone array;
[0012] S3. Use the SRP-PHAT algorithm to search within the range of the sound source angle determined in step S2 to determine the direction angle of the sound source.
[0013] Furthermore, the microphone array structure adopts a 4-microphone linear microphone array structure.
[0014] Furthermore, in step S1, the time delay between other microphones and the reference microphone is calculated by the GCC algorithm, which specifically includes: first, performing fast Fourier transform on the observation signals of the two microphones; then, conjugate multiplying the two signals after the fast Fourier transform and performing PHAT weighting; finally, performing inverse Fourier transform on the PHAT-weighted signal to obtain a generalized cross-correlation sequence, and the time delay corresponding to the peak in the generalized cross-correlation sequence is the signal delay between the two microphones.
[0015] Further, in step S2, according to the geometric structure of the microphone array, the range of the sound source angle is determined based on the time delay calculated in step S1, specifically including:
[0016] Assuming that the time delay calculated in step S1 is T, the range of the sound source angle θ is:
[0017]
[0018] Among them, f s is the sampling frequency, c is the speed of sound, and L is the spacing between microphone array elements.
[0019] Further, in step S3, the use of the SRP-PHAT algorithm to search within the range of the sound source angle determined in step S2 to determine the direction angle of the sound source specifically includes:
[0020] First, the signal collected by each microphone is Fourier transformed; then, the amplitude frequency of different frequencies is phase compensated according to the scanning angle; then, the frequency domain signals of all microphones are summed to obtain the summed frequency domain signal after phase compensation; finally, the summed frequency domain signal after phase compensation is inverse Fourier transformed to obtain the compensated time series signal at this angle, and within the sound source angle range determined in step S2, the angle that maximizes the energy of the compensated time series signal is determined as the sound source direction angle.
[0021] The beneficial effects of the present invention are:
[0022] The present invention first uses the generalized cross-correlation GCC sound source localization algorithm to calculate the time delay between two observation signals, then determines the angular range of the sound source according to the time delay deviation, and then uses the SRP-PHAT spatial spectrum algorithm to scan within the direction range to determine the direction of the sound source. That is, the present invention uses the advantages of the GCC sound source localization algorithm of small calculation amount and high speed, first determines the angular range of the sound source, and combines the advantages of the SRP-PHAT spatial spectrum algorithm of high precision to directly search the direction of the sound source within the determined angular range of the sound source, thereby reducing the amount of calculation while ensuring the accuracy of sound source localization. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is an overall flow chart of the sound source localization method in the present invention;
[0024] Figure 2 is a schematic diagram of a microphone array in an embodiment of the present invention;
[0025] Figure 3 It is the GCC algorithm flow chart;
[0026] Figure 4 is the generalized cross-correlation sequence diagram;
[0027] Figure 5 This is the flow chart of the SRP-PHAT algorithm. DETAILED DESCRIPTION
[0028] Sound source localization usually uses multiple microphones to collect sound signals at different measurement positions, and uses the differences in the sounds collected by different microphones to determine the location of the sound. For sound source localization, two pieces of information are usually required: the direction of arrival (DOA) of the sound source and the distance estimation of the sound source, and the present invention mainly studies DOA estimation. The present invention aims to propose a sound source localization method based on SRP-PHAT spatial spectrum and GCC, which reduces the amount of calculation while ensuring the accuracy of sound source localization. The core idea is: first, use the GCC algorithm to make a rough estimate of the sound source direction, obtain the sound source angle range, and thus narrow the scanning range of the beam, and then use the SRP-PHAT algorithm to perform more accurate angle compensation on the angles within the range to calculate the output power of the wave technique and determine the final sound source direction. The overall process is as follows: Figure 1 shown.
[0029] The solution of the present invention combines the advantages of the small computational complexity of the GCC algorithm and the high precision of the SRP-PHAT algorithm. Compared with the traditional single-use GCC algorithm, it has higher positioning accuracy and does not require high sampling frequency and array element spacing. Compared with the traditional single-use SRP-PHAT algorithm, the search angle range is small, the calculation time is shorter, and the calculation complexity is smaller.
[0030] Example:
[0031] This embodiment adopts a 4-microphone linear microphone array structure, the spacing between each array element is l, assuming that the sound field is a far-field source, the array structure is as follows Figure 2 shown.
[0032] Assume that the signal collected by the i-th microphone is x i (t), the sampling frequency is f s , the spacing between array elements is l. First, the signal is subjected to the GCC (generalized cross correlation) algorithm to determine a rough direction range.
[0033] The generalized cross-correlation algorithm process is as follows Figure 3 As shown, firstly, the two observation signals are fast Fourier transformed, then the two are conjugate multiplied and PHAT weighted, and finally the inverse Fourier transform is performed to obtain the generalized cross-correlation sequence, which is described as follows:
[0034] The generalized cross-correlation algorithm between two observed signals x1(t) and x2(t) is defined as:
[0035]
[0036] in, is the cross spectrum of the two signals, X i (f) Through fast Fourier transform, we get:
[0037]
[0038] θ(f) is a weighting function. There are many common weighting functions. This embodiment adopts a phase transform PHAT weighting function:
[0039]
[0040] The phase transformation weighting function is essentially a filter. When the microphone array is used in actual situations, the peak of the generalized cross-correlation function is not obvious due to the presence of reverberation and environmental noise, which reduces the accuracy of the calculated delay. Therefore, PHAT is needed to highlight the peak and suppress the interference of noise and reverberation. Finally, the corresponding peak of the generalized cross-correlation function of x1(t) and x2(t) is the delay T between the two channels. 12 The sequence obtained by the GCC algorithm is as follows Figure 4 As shown, the delay corresponding to the peak value is T 12 , and then the incident angle θ of the sound source can be calculated 12 :
[0041]
[0042] Where c is the speed of sound and L is the distance between array elements. Similarly, the four microphones can obtain the incident angle θ of the sound source 12 ,θ 23 ,θ 34 ,θ 41 , and then the direction angle θ of the sound source can be determined using the geometric relationship of the microphone array:
[0043]
[0044] Although the GCC algorithm can find the angle of the sound source, the principle of the algorithm itself causes errors. When GCC calculates the delay, the delay accuracy is 1 / f s , then the possible error ΔT in calculating the delay is equal to 1 / f s . Then the error caused by the orientation angle is
[0045] err = arccos(c / Lf s ) (6)
[0046] Therefore, in practical applications, when the calculated delay is T, the actual angle range should be:
[0047]
[0048] SRP-PHAT positioning algorithm:
[0049] Controllable beam response is a method based on beamforming. It finds the direction with the highest energy by compensating for delay and then accumulating. Since speech is a broadband signal, a broadband beam thread is required. After the GCC algorithm determines the approximate direction of the sound source, the SRP-PHAT algorithm can traverse a smaller angle range when traversing the angle and use a higher traversal accuracy, thereby reducing the amount of calculation and improving accuracy.
[0050] The process of SRP-PHAT algorithm is as follows Figure 5 As shown, the signal collected by each microphone is first Fourier transformed, and then the amplitude frequency of different frequencies is phase compensated according to the scanning angle. Then the frequency domain of all microphones is summed to obtain the summed frequency domain signal after phase compensation. Finally, the inverse Fourier transform can obtain the compensated time series signal at this angle, which is described as follows:
[0051] Assume that the signal collected by mic1 is x1(t). Since mic2~mic4 are linear arrays, there will be a linear delay relative to the signal collected by mic1. The data collected by the i-th mic
[0052] X i (ω) = a 1i X1(ω)e -j2πωτ(i-1) (8)
[0053] Ignoring the influence of noise, each mic will have a delay and attenuation, and the delay is related to the distance between microphone mic1.
[0054] If X i (ω) performs phase compensation τ(i-1) and sums all mic spectra:
[0055]
[0056] The phase corresponding to the maximum energy and value is the direction angle of the voice. When the incident angle is θ, the phase compensation of the i-th mic is γ i for
[0057]
[0058] Since the speech signal is a broadband signal, it is necessary to perform phase compensation and calculate the energy sum in all frequency ranges. Perform phase compensation on the frequency domain signals of all microphones and sum them according to the frequency scale to obtain the compensated frequency domain sum:
[0059]
[0060] Y θ Perform inverse Fourier transform to obtain the phase-compensated signal and y θ(t). According to the θ range obtained by the previous GCC algorithm (Equation (7)), within this range y θ (t) The angle with the maximum energy is the final sound source positioning direction angle.
[0061] Finally, it should be noted that the above embodiments are only preferred implementations and are not intended to limit the present invention. It should be pointed out that for those skilled in the art, several modifications, equivalent replacements, improvements, etc. can be made without departing from the scope of the present invention and the scope of protection of the claims, and all of these should be included in the protection scope of the present invention.
Claims
1. A sound source localization method based on SRP-PHAT spatial spectrum and GCC, characterized in that: The method comprises the following steps: S1. According to the observation signal collected by the microphone in the microphone array structure, the time delay between other microphones and the reference microphone is calculated by the GCC algorithm; S2. Determine the range of the sound source angle based on the time delay calculated in step S1 according to the geometric structure of the microphone array: Assuming the time delay calculated in step S1 is T, the sound source angle The range is: in, is the sampling frequency, c is the speed of sound, and L is the spacing between microphone array elements; S3, using the SRP-PHAT algorithm to search within the range of the sound source angle determined in step S2, to determine the direction angle of the sound source: First, the signal collected by each microphone is Fourier transformed; then, the amplitude frequency of different frequencies is phase compensated according to the scanning angle; then, the frequency domain signals of all microphones are summed to obtain the summed frequency domain signal after phase compensation; finally, the summed frequency domain signal after phase compensation is inverse Fourier transformed to obtain the compensated time series signal at this angle, and within the sound source angle range determined in step S2, the angle that maximizes the energy of the compensated time series signal is determined as the sound source direction angle.
2. A sound source localization method based on SRP-PHAT spatial spectrum and GCC as claimed in claim 1, characterized in that: In step S1, the time delay between other microphones and the reference microphone is calculated by the GCC algorithm, which specifically includes: first, performing fast Fourier transform on the observation signals of the two microphones; then, conjugate multiplying the two signals after fast Fourier transform and performing PHAT weighting; finally, performing inverse Fourier transform on the PHAT-weighted signal to obtain a generalized cross-correlation sequence, and the time delay corresponding to the peak in the generalized cross-correlation sequence is the signal delay between the two microphones.
3. A sound source localization method based on SRP-PHAT spatial spectrum and GCC as claimed in claim 1 or 2, characterized in that: The microphone array structure adopts a 4-microphone linear microphone array structure.
Citation Information
Patent Citations
Double-microphone speech enhancement method based on improved power difference noise estimation algorithm
CN111951818A
Methods and systems for estimating the location of sound sources using azimuth-frequency expression and convolution neural network model
KR102199158B1