Sound source localization method using time delay and dead reckoning

By combining wavelet transform noise reduction and an improved generalized cross-correlation algorithm, along with inertial sensors and virtual sound rotation processing, the accuracy and multiple solutions of sound source localization in complex environments are solved, achieving low-cost and highly portable sound source localization suitable for ordinary smartphones.

CN121027996AActive Publication Date: 2025-11-28EAST CHINA JIAOTONG UNIVERSITY

Patent Information

Application Number
CN202511565768.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2025-11-28
Estimated Expiration
2045-10-30

AI Technical Summary

Technical Problem

Existing sound source localization technology is not accurate enough in complex environments. Traditional dual-microphone systems cannot provide a unique sound source direction and are costly, not portable, and difficult to operate in real time on ordinary smartphones.

Method used

By utilizing the common dual microphones and inertial sensors built into a smartphone, combined with wavelet transform noise reduction and an improved generalized cross-correlation algorithm, high-precision estimation of sound source direction and three-dimensional spatial positioning are achieved through dead reckoning and virtual sound rotation processing.

Benefits of technology

It significantly improves the estimation accuracy of sound source arrival time difference, solves the problem of multiple solutions in traditional methods, realizes low-cost and portable sound source localization, and expands the application potential of consumer-grade scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121027996A_ABST
    Figure CN121027996A_ABST
Patent Text Reader

Abstract

The invention provides a sound source localization method using time delay and dead reckoning. The method comprises the following steps: acquiring an original sound signal and carrying out noise reduction preprocessing; virtual sound rotation processing is carried out on the two sound source directions, interpolation simulation rotation is carried out on effective sound signals after noise reduction, and angle differences before and after rotation are compared to eliminate false angles; and finally calculating the position of the sound source through a geometric positioning algorithm by combining the unique sound source direction, the movement distance and the movement direction of the equipment and the time difference between the sound source and the two microphones. According to the method, acoustic orientation and a smart phone inertial navigation technology are deeply fused, a displacement vector is obtained by utilizing dead reckoning, and multi-position TDOA measurement information is combined, so that sound source three-dimensional space positioning only depending on single mobile equipment is realized, and the inherent defect that a traditional pure acoustic positioning method cannot measure distance is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound source localization technology, and in particular to a sound source localization method that utilizes time delay and dead reckoning. Background Technology

[0002] In recent years, target location technology has attracted much attention from the academic community, becoming a research hotspot and finding widespread application in reality. Target location technology also plays an important role in the security field. Police use this technology to track criminal suspects or find missing persons, while fire departments use it for firefighting and rescue operations. In daily life, applications such as mobile phone positioning and car navigation are also important uses of target location technology, providing convenience and security for our lives.

[0003] Common target localization technologies mainly include video localization, wireless signal localization, and sound localization. Video-based target localization systems use cameras to capture video data and determine the target's position and trajectory through target detection, tracking algorithms, and motion estimation. Wireless signal-based localization technologies use signal characteristics and localization algorithms to determine the target's location. By using a signal source and receiver, signal characteristics are extracted, and localization algorithms are applied to calculate the position. However, video positioning is limited by the camera's finite field of view; when the target is outside the field of view, the system may be unable to continue tracking. Furthermore, lighting conditions and environmental factors can affect image quality and the accuracy of target detection. Wireless signal positioning technology also faces challenges such as signal attenuation and multipath effects, which may affect the accuracy and stability of the positioning system.

[0004] Sound-based localization technology boasts advantages such as multi-directional positioning, adaptability to complex environments, independence from light limitations, high real-time performance, and low cost. These characteristics make sound localization systems stand out in low-light conditions, when objects obstruct the view, or when real-time positioning is required, and they typically have low implementation costs. Therefore, sound source targeting based on sound signals has significant development potential and a large market. For example, in fire and earthquake scenes, sound source localization can pinpoint the location of those seeking help. However, this technology also has several limitations, including the need for customized sensors, the requirement for multiple sensors working together, and the large size of the equipment. Summary of the Invention

[0005] In view of the above, the main objective of this invention is to propose a sound source localization method that utilizes time delay and dead reckoning to solve the aforementioned technical problems.

[0006] This invention proposes a sound source localization method utilizing time delay and dead reckoning, the method comprising the following steps: Step 1: Acquire the original sound signal and perform noise reduction preprocessing to obtain the effective sound signal after noise reduction; Step 2: For the effective sound signal after noise reduction, calculate the time delay of the sound source to reach the two array elements in the two microphone arrays using the improved generalized cross-correlation algorithm, and calculate the arrival angle of the sound source based on the time delay. Based on the symmetry of the device, obtain the directions of the two sound sources. Step 3: The two sound source directions are processed by virtual sound rotation. The effective sound signal after noise reduction is interpolated to simulate rotation. The angle difference before and after rotation is compared to eliminate false angles and determine the unique sound source direction. Step 4: Set the position when the device starts moving as the initial position and obtain the coordinates of the initial position; when the device starts moving, use the device's inertial sensors to collect the corresponding acceleration and angular velocity data, and combine them with the coordinates of the initial position to calculate the moving distance and direction of the device using a dead reckoning algorithm. Step 5: Combining the unique sound source direction, the device's moving distance and direction, and the time difference between the sound source's arrival at the two microphones, the location of the sound source is finally calculated using a geometric positioning algorithm.

[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention utilizes the common dual microphones and inertial sensors built into smartphones, combined with wavelet transform noise reduction and an improved generalized cross-correlation algorithm, to effectively suppress environmental noise and reverberation interference, significantly improving the estimation accuracy of sound source arrival time difference (TDOA) in complex acoustic environments, thereby achieving highly robust sound source direction estimation.

[0008] 2. This invention innovatively solves the problem of ambiguous sound source direction caused by the symmetry of the device structure by proposing virtual sound rotation processing technology. It can accurately determine the direction of a unique sound source without adding hardware, breaking through the limitation of traditional dual-microphone systems that can only provide symmetrical dual solutions.

[0009] 3. This invention deeply integrates acoustic orientation with smartphone inertial navigation technology, uses dead reckoning to obtain displacement vectors, and combines multi-location TDOA measurement information to achieve three-dimensional spatial positioning of sound sources relying solely on a single mobile device, overcoming the inherent defect of traditional pure acoustic positioning methods that cannot measure distances.

[0010] 4. By optimizing the algorithm structure and data processing flow, this invention enables the entire positioning system to run in real time on ordinary smartphones. It has significant advantages such as low cost, high portability, and easy deployment, greatly expanding the application potential of sound source localization technology in consumer scenarios.

[0011] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by means of embodiments of the invention. Attached Figure Description

[0012] Figure 1 This is a flowchart illustrating the steps of a sound source localization method based on time delay and dead reckoning proposed in this invention.

[0013] Figure 2 This diagram illustrates the wavelet denoising process of a sound source localization method that utilizes time delay and dead reckoning proposed in this invention.

[0014] Figure 3 This is a schematic diagram of the sound wave arrival angle of a sound source localization method that utilizes time delay and dead reckoning proposed in this invention.

[0015] Figure 4 This is a flowchart of the cross-correlation algorithm for a sound source localization method that utilizes time delay and dead reckoning proposed in this invention.

[0016] Figure 5 This is a flowchart of the generalized weighted cross-correlation and quadratic cross-correlation combined algorithm of a sound source localization method that utilizes time delay and dead reckoning proposed in this invention.

[0017] Figure 6 This is a regional division diagram of the sound source location method proposed in this invention, which utilizes time delay and dead reckoning to locate the sound source.

[0018] Figure 7 This is a schematic diagram illustrating the sound rotation effect of a sound source localization method based on time delay and dead reckoning proposed in this invention.

[0019] Figure 8 This is a schematic diagram of the interpolated sound signal effect of a sound source localization method that utilizes time delay and dead reckoning proposed in this invention.

[0020] Figure 9 This is a schematic diagram of a dead reckoning algorithm for a sound source localization method that utilizes time delay and dead reckoning proposed in this invention.

[0021] Figure 10 This is a schematic diagram illustrating the combination of dead reckoning and TDOA in a sound source localization method that utilizes time delay and dead reckoning proposed in this invention. Detailed Implementation

[0022] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0023] These and other aspects of the embodiments of the present invention will become clear from the following description and accompanying drawings. In these descriptions and drawings, some specific embodiments of the present invention are specifically disclosed to provide some ways of implementing the principles of the embodiments of the present invention; however, it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0024] Please see Figure 1 This invention proposes a sound source localization method using time delay and dead reckoning, which includes the following steps: Step 1: Acquire the original sound signal and perform noise reduction preprocessing to obtain the effective sound signal after noise reduction.

[0025] Please see Figure 2 In step 1, the original sound signal is acquired and subjected to noise reduction preprocessing to obtain the noise-reduced effective sound signal, which specifically includes the following steps: The original sound signal is acquired, and multi-level discrete wavelet decomposition is performed on the original sound signal. Low-pass and high-pass filters are used to extract the low-frequency approximation coefficients and high-frequency detail coefficients of the signal, respectively. The low-frequency approximation coefficients and high-frequency detail coefficients of the signal are processed using preset thresholds. Coefficients below the preset thresholds are set to zero to achieve noise suppression. At the same time, speech components above the thresholds are retained to obtain the processed approximation coefficients and processed detail coefficients. The processed approximation coefficients and processed detail coefficients are used to reconstruct the signal using a multi-resolution reconstruction algorithm to achieve noise reduction, resulting in a noise-reduced effective sound signal.

[0026] For further details, please refer to Figure 2 Given that the original signal contains both the desired sound source signal and environmental noise, this invention employs wavelet transform, a particularly effective method for filtering out noise. Specifically, for a given signal S with 1000 data values, low-pass and high-pass filters are first used to obtain approximate coefficients (low-frequency components) and detail coefficients (high-frequency components). Then, multi-level discrete wavelet decomposition is performed, and thresholding techniques are used to process the detail coefficients (the wavelet coefficients obtained after wavelet transform contain important time and frequency domain information; the wavelet coefficients of the true signal are larger, while those of noise are smaller. By selecting a suitable threshold, larger coefficients are retained, and smaller ones are set to 0, thus preserving the noise). Finally, the signal is reconstructed using a multi-resolution reconstruction algorithm to achieve noise reduction. This method has potentially significant implications for indoor sound source localization, as it can improve the accuracy and timeliness of localization algorithms.

[0027] Step 2: For the effective sound signal after noise reduction, calculate the time delay of the sound source to reach the two array elements in the two microphone arrays using the improved generalized cross-correlation algorithm, and calculate the arrival angle of the sound source based on the time delay. Based on the symmetry of the device, obtain the directions of the two sound sources.

[0028] Please see Figure 3 , Figure 4 and Figure 5 In step 2, for the noise-reduced effective sound signal, the time delay of the sound source arriving at the two array elements in the two microphone arrays is calculated using an improved generalized cross-correlation algorithm. Based on the time delay, the arrival angle of the sound source is calculated, and according to the device symmetry, the directions of the two sound sources are obtained. Specifically, this includes: The device uses its microphones to receive the noise-reduced effective sound signals, thus obtaining a dual-channel microphone reception signal. The cross-correlation function of the received signals from the dual-channel microphones is calculated, and peak search is performed to obtain the estimated time delay of the sound source reaching the two microphones. The estimated time delay of the sound source to reach the two microphones, the known speed of sound, and the distance between the microphones are substituted into the formula for calculating the angle of arrival. At the same time, based on the symmetry of the equipment, the directions of the two sound sources are finally obtained.

[0029] The cross-correlation function of the received signals from the dual-channel microphones is calculated, and a peak search is performed to obtain the estimated time delay of the sound source arriving at the two microphones. Specifically, this includes: Perform a discrete Fourier transform on the received signals from the dual-channel microphones to obtain the frequency domain representation of the dual-channel signals; Based on the frequency domain representation of the dual-channel signal, the cross-power spectrum between the two signals is calculated and obtained; The cross-power spectrum between the two signals is weighted using the PHAT weighting function to obtain the first weighted cross-power spectrum. Perform an inverse Fourier transform on the first weighted cross-power spectrum to obtain the first generalized cross-correlation function; An improved weighting function is applied to the cross-power spectrum between the two signals to obtain a second weighted cross-power spectrum; wherein, the improved weighting function introduces a noise cross-power spectrum estimation term and an adjustable weighting factor into the denominator of the PHAT weighting function. Perform an inverse Fourier transform on the second weighted cross-power spectrum to obtain the second generalized cross-correlation function; A joint analysis of the first and second generalized cross-correlation functions is performed, and a peak search is conducted. By finding the time offset corresponding to the global maximum value of the peak, the estimated time delay of the sound source reaching the two microphones is obtained.

[0030] The device uses its microphones to receive the noise-reduced effective sound signals, resulting in dual-channel microphone received signals. The corresponding relationship in this process is as follows: ; in, This indicates the signal received by the first microphone. Represents the independent variable, time. This indicates the signal received by the second microphone. This indicates the first transmission gain of the signal reaching the microphone. This indicates the second transmission gain of the signal reaching the microphone. Indicates the sound source. Indicates the time it takes for the sound to reach the first microphone. Indicates the time it takes for the sound to reach the second microphone. This represents the noise signal received by the first microphone. This indicates the noise signal received by the second microphone; It should be noted that the device in this embodiment of the invention is a mobile phone.

[0031] In the process of substituting the estimated time delay of the sound source to the two microphones, the known speed of sound, and the distance between the microphones into the formula for calculating the angle of arrival, and considering the symmetry of the equipment, the relationship between the two sound sources is as follows: ; in, Indicates the angle at which the sound source arrives. Represents inverse trigonometric functions, Indicates the speed of sound. This indicates the distance between the two microphones.

[0032] It should be noted that the right side of the formula is calculated based on the relationship between the triangle sides. This The corresponding radians, and then the inverse solution from the radians. That is, the angle at which the sound arrives.

[0033] Performing a Discrete Fourier Transform on the received signals from the dual-channel microphones yields the frequency domain representation of the dual-channel signals. The corresponding relationship in this process is as follows: ; in, This indicates the signal received by the first microphone. The signal received by the second microphone cross power spectral density, Represents a complex exponential function. This indicates the signal received by the second microphone. Representation in the frequency domain The complex conjugate, Represents angular frequency. This indicates the signal received by the first microphone. The representation in the frequency domain; It should be noted that, Used to handle the conjugate symmetry of complex signals, in frequency domain analysis, combined with the Fourier transform of the original signal, it can reveal the energy, correlation and other characteristics of the signal.

[0034] In the step of applying the PHAT weighting function to the cross-power spectrum between the two signals to obtain the first weighted cross-power spectrum, the corresponding relationship in the process is as follows: ; in, This indicates the signal received by the first microphone. The signal received by the second microphone A weighted function for the numerical values ​​of the correlation functions between them; In the step of applying an improved weighting function to the cross-power spectrum between the two signals to obtain the second weighted cross-power spectrum, the corresponding relationship in the process is as follows: ; in, This represents the weighting function used in the first calculation of the relevant function values. This represents the weighting function used when calculating the correlation function value for the second time. This indicates the signal received by the first microphone. The self-power spectral density, This represents the cross-power spectral density calculated during the second cross-correlation of the two signals after weighting function processing. The cross-power spectral density represents the noise. Indicates the weighting factor; In the step of performing an inverse Fourier transform on the second weighted cross-power spectrum to obtain the second generalized cross-correlation function, the corresponding relationship in the process is as follows: ; in, This indicates the signal received by the first microphone. The autocorrelation function, This represents the autocorrelation function when calculating the autocorrelation function value for the second time.

[0035] Furthermore, after noise reduction processing of the signal, the TDOA (Time Delay Aspect) of the data can be estimated. Since the signals collected by the two microphones are signals arriving at different times from the same sound source, the relative time delay between the two signals can be determined by calculating the peak value of the cross-correlation function (a function used to describe the correlation or similarity between two signals, where the function ranges from -1 to 1; a value of -1 indicates that the two signals are completely synchronized, 0 indicates that the two signals are unrelated, and -1 indicates that the two signals are completely opposite) and thus solving for the location information of the sound source.

[0036] Furthermore, the cross-correlation function of the signals received by the dual-channel microphones is calculated, and the corresponding relationship is as follows: ; in, Represents the cross-correlation function. This represents the calculation of mathematical expectation. Indicates a time delay; By inputting signal data from different times into the cross-correlation function, countless points are obtained. These points form a signal data image, and the correlation between the two images is then observed. The cross-correlation function is then defined as follows: ; in, Indicates the time it takes for the sound to reach the first microphone. This indicates the time it takes for the sound to reach the second microphone.

[0037] It should be noted that the input and output remain unchanged during the derivation process; This indicates a time delay operation on the sound source signal, shifting the signal along the time axis. ; This indicates that the sound source signal is subjected to two superposition time delay operations, which shifts the signal on the time axis. and ; This indicates a time delay operation on the noisy signal, shifting the signal along the time axis. ; This indicates that the sound source signal is subjected to two superposition time delay operations, which shifts the signal on the time axis. and .

[0038] Assumption It is white noise and is related to the sound source. If the cross-correlation function is independent (its value is 0, so the function containing the noise and the cross-correlation with the sound source can be directly eliminated as 0), then the cross-correlation function is: ; Assumption and There is no correlation between them (containing) and If the cross-correlation function takes a value of 0, then the cross-correlation function is: ; in, As the sound source The autocorrelation function, Indicates the time it takes for the sound to reach the first microphone. Time of sound reaching the second microphone The time interval.

[0039] Because the autocorrelation function of the signal has an arbitrary time delay The values ​​below the threshold do not exceed the value at which the time delay is 0. In other words, the correlation between the sound signals emitted by a sound source is strongest at the exact moment of sound emission; afterwards, the correlation gradually weakens and disperses as the sound propagates. Therefore, when... At that time, cross-correlation function The mathematical expression for obtaining the maximum value is: ; in, This refers to a Python function used to find the index of the maximum value in a given function or array, particularly in mathematics and computer science, especially in optimization problems and machine learning. It can help determine at which point a function or array reaches its maximum value.

[0040] When calculating the time delay of a signal arriving at two elements of a microphone array using the cross-correlation algorithm, environmental noise and reverberation interference cause the calculated cross-correlation peak to be mixed with spurious peaks formed by noise and other interference factors, making accurate judgment impossible. To improve the estimation accuracy, a combination of the basic cross-correlation algorithm and a frequency domain weighting function is used to suppress environmental interference and thus sharpen the peak. The cross-power spectrum after PHAT weighting is then subjected to an inverse Fourier transform to the time domain to obtain the time delay function of the signal arriving at the two microphone sensors. , At this point, the cross-correlation function reaches its maximum value. After a Fourier transform, the cross-correlation function is then weighted.

[0041] To address interference issues that cannot be resolved in the time domain, we can weight the cross-power spectrum in the frequency domain based on prior knowledge of the signal and noise, thereby suppressing noise and reverberation interference. Finally, an inverse Fourier transform is performed to obtain the generalized cross-correlation function. This approach combines cross-correlation analysis of two signals in the time domain with operations on the frequency components of the signals in the frequency domain, multiplying, integrating, and then transforming back to the time domain. This leverages the convenience of linear operations in the frequency domain. The algorithm flow is as follows: Figure 4 As shown.

[0042] like Figure 5 As shown, noise and reverberation interference still exist, and the experimental results are still significantly affected. To increase resistance to noise and reverberation interference, a quadratic cross-correlation algorithm is used to estimate the signal delay. Compared with the performance of the traditional generalized cross-correlation algorithm, the noise resistance of the quadratic cross-correlation algorithm is further enhanced. Although the noise resistance of the time delay estimation algorithm of the quadratic cross-correlation algorithm is improved compared with the single generalized cross-correlation algorithm, the relationship between its result error and the signal-to-noise ratio is still inversely proportional. Therefore, this invention proposes to combine the two algorithms to enhance its noise resistance by combining the generalized cross-correlation algorithm based on the PHAT weighting function with the quadratic cross-correlation algorithm. Compared with the direct generalized quadratic cross-correlation algorithm, this algorithm first selects an appropriate weighting function to perform a weighting process on the signals before calculating the delay of the two signals through the quadratic cross-correlation. This further improves the signal-to-noise ratio (a measure of signal strength and stability, which is the ratio of the power of the useful sound signal to the power of the noise signal) of the two input signals in the quadratic cross-correlation operation. Then, the inverse Fourier transform is performed to obtain the cross-correlation function and the delay estimate. In real-world environments, it is difficult to achieve the assumed condition that the signal and noise are uncorrelated, and it is also difficult to ensure that the received noise is uncorrelated. The presence of noise and reverberation can interfere with the time delay results. To reduce the impact of noise and reverberation, when improving the PHAT weighting function, the cross-correlation function of the noise is first subtracted from the denominator to weaken the influence of noise. Secondly, a weighting factor is introduced. The optimal value of the weighting factor is estimated through multiple experiments under different signal-to-noise ratio conditions. Based on experience, an appropriate weighting factor is selected to reduce the time delay estimation error of the PHAT weighting function under low signal-to-noise ratio conditions and improve the noise and reverberation resistance under high reverberation and high noise conditions, thereby improving the accuracy of time delay estimation.

[0043] After calculating the frequency domain using the quadratic cross-correlation function and then performing an inverse Fourier transform to the time domain, the time delay can be calculated using the formula. Based on this mathematical model, the angle of arrival of the sound source can then be determined. For example... Figure 3 As shown in the figure, the coordinates of the two microphones are respectively... and Since sound waves travel in a straight line, the difference in propagation distance between the two microphones when the sound source emits sound is _____. .

[0044] Step 3: The two sound source directions are processed by virtual sound rotation. The effective sound signal after noise reduction is interpolated to simulate rotation. The angle difference before and after rotation is compared to eliminate false angles and determine the unique sound source direction.

[0045] Please see Figure 6 , Figure 7 and Figure 8 In step 3, the two sound source directions are processed by virtual sound rotation. The effective sound signal after noise reduction is interpolated to simulate rotation. The angle difference before and after rotation is compared to eliminate false angles and determine the unique sound source direction. The specific steps include the following: Based on the central symmetry of the device structure, the surrounding space of the two sound source directions is divided into four regions. The sound source arrival angles calculated for the two opposite regions are the same, thus obtaining two symmetrical candidate values ​​for the sound source directions. Interpolation processing is performed on the signals received by the dual-channel microphone. By changing the signal sample points, the change in the arrival time difference of the sound waves is simulated, which is equivalent to realizing the virtual rotation of the sound signal, and thus obtaining the interpolated sound signal. The arrival angle of the sound source is recalculated on the interpolated sound signal to determine the direction of the sound source after virtual rotation. The virtual rotated sound source direction is compared with two symmetrical sound source direction candidate values. Based on the difference in the change of angle in different regions, false sound source directions are identified and eliminated, and the unique true sound source direction is determined.

[0046] Furthermore, after obtaining the estimated time delay, an angle value can be calculated using a formula. However, the angle obtained at this point is not unique because the mobile phone is a centrally symmetric object, so the angle has two possible values, such as... Figure 6 As shown, when the phone is held horizontally, we divide the possible locations of the sound source into... (First area) (Second area) (The third area) and (The fourth area), in which, and , and The calculated angle values ​​are the same, and the combined cross-correlation algorithm cannot distinguish the angles of the upper and lower sound sources. Therefore, this invention interpolates the recorded sound signals during calculation to obtain a virtual rotation effect. By comparing the direction of the rotated sound source, the exact angle of sound arrival can be obtained.

[0047] To address this problem, this invention uses virtual sound effects to rotate the sound signals received by the microphone during computation, thereby distinguishing the sound source. A virtual sound source is the creation and perception of a sound source that does not physically exist in a given space. This is commonly used in audio engineering, virtual reality, and augmented reality to create immersive and realistic auditory experiences. The goal is to allow listeners to perceive sound from a specific direction, distance, and location, even if the actual sound-producing object does not exist. One method to achieve this is using multiple speakers placed around the listener to create the auditory illusion of sound coming from different directions. This is commonly used in home theater systems and immersive audio setups. The mathematical formulas used to calculate the delays and intensities required to create a specific auditory perception can be complex. Another approach is through binaural audio. Binaural recording or processing uses two microphones positioned similarly to the human ear. By recording or simulating sound in this way and then playing it back through headphones, listeners can perceive a 3D auditory environment because the brain processes the differences in sound arrival time and intensity between the two ears to locate the sound. The scenario encountered in this invention does not require setting a separate HRTF function (Head-Related Transfer Function (HRTF) is a filter that represents the way sound changes as it travels from the sound source to the ear, taking into account the shape of the head, ears, and torso); it only requires signal processing methods related to it.

[0048] Furthermore, such as Figure 7 As shown, suppose there is a sound source whose angle of arrival has been obtained after initial calculation. At this point, the calculated sound source direction is symmetrical, and it is necessary to eliminate an incorrect sound source direction. This can be achieved by processing the received signal and rotating it by 90° mathematically. Figure 7 In the 'a' of the sound register, then... The angle calculated after rotation is... And the vocal range The calculation result after rotation is still the same. If at this time, originally in If the sound source in the area is real, then the actual sound source should be located in... The area and its symmetrical area ;exist Figure 7 In b, if If the sound source in the area is real, then the actual sound source should be located in... The area and its symmetrical area By simply determining the sound source location after rotating it 90° and the area where its mirror image is located, the original correct sound source location can be determined. Based on this principle, incorrect sound source directions can be eliminated.

[0049] This invention can distinguish which sound region the sound source is emitting from based on the different results after rotation. Figure 7 Similarly, for vocal registers and Calculated sound arrival angle , vocal range After rotating 90° clockwise, the angle at which the sound source arrives remains the same. And the vocal range The angle after rotation is .like Figure 8 As shown, according to step 2, the relationship between angle and time delay in physical space is as follows: With the speed of sound and the microphone spacing being fixed, this invention can determine that there is a certain linear relationship between the time difference and the angle of arrival, so the angle value can be changed by changing the value of the time difference.

[0050] To reduce the computational burden of experiments, this invention performs inference based on cross-correlation calculations. For example, if a PCM signal is interpolated to zero without changing its properties, such as... Figure 8 As shown, the cross-correlation function will calculate some additional zero values, and the peak value of the cross-correlation function, which is used to calculate the time delay, will be shifted several interpolated values ​​later, resulting in a delay a few microseconds longer than the original delay. By using the mathematical relationship between the time delay and the angle of arrival, the change in the angle of arrival can be calculated, thus achieving the effect of virtual sound rotation.

[0051] Step 4: Set the position when the device starts moving as the initial position and obtain the coordinates of the initial position; when the device starts moving, use the device's inertial sensors to collect the corresponding acceleration and angular velocity data, and combine them with the coordinates of the initial position to calculate the moving distance and direction of the device using a dead reckoning algorithm.

[0052] In step 4, please refer to Figure 9 The initial position is set when the device begins to move, and its coordinates are obtained. As the device begins to move, its inertial sensors collect corresponding acceleration and angular velocity data. These data, combined with the initial position coordinates, are used in a dead reckoning algorithm to calculate the device's moving distance and direction. The relationships involved in this process are as follows: ; in, Indicates the first The x-coordinate of the next position. Indicates the first The distance to the next position, Represents the sine function. Indicates the first The dead recoil angle at the next position Indicates the first The ordinate of the next position This represents the cosine function.

[0053] Furthermore, the basic principle of dead reckoning algorithms is to calculate the direction and distance of movement using data obtained from sensors to obtain the relative position. The principle is as follows: Figure 9 As shown, Indicates the pedestrian's initial position. Indicates the initial position coordinates This represents the coordinates of the pedestrian's first position. This represents the coordinates of the pedestrian's second position. express and The azimuth angle of the satellite at the position. express and The azimuth angle of the position. When the pedestrian's initial position is known... At that time, dead reckoning algorithms can be used to calculate The position, similarly It can be derived from dead reckoning algorithms.

[0054] Step 5: Combining the unique sound source direction, the device's moving distance and direction, and the time difference between the sound source's arrival at the two microphones, the location of the sound source is finally calculated using a geometric positioning algorithm.

[0055] In step 5, please refer to Figure 10 By combining the unique direction of the sound source, the distance and direction of the device's movement, and the time difference between the sound source's arrival at the two microphones, the location of the sound source is finally calculated using a geometric positioning algorithm. The specific steps include: The initial position is marked as point A. The sound source signal is collected at point A, the arrival time difference of point A is calculated, and the arrival angle of the sound source relative to point A is calculated based on the arrival time difference of point A. Keeping the equipment attitude unchanged, move the equipment from point A to the east or west to the second position point, and mark the second position point as point B. Based on the sensor data collected during the equipment movement, calculate the moving distance from point A to point B using the dead reckoning algorithm, and calculate the angle of arrival of the sound source relative to point B at point B. The location of the sound source is denoted as point O. Based on the coordinates of points A and B, the arrival angle of the sound source relative to point A, and the arrival angle of the sound source relative to point B, the center coordinates and radius of the circle ABO formed by points A, B, and the sound source location point O are determined, and the corresponding central angle is calculated. The equipment is moved eastward from point B to the third position point, which is denoted as point C. The coordinates of point C are determined. During the movement of the equipment, the heading angle ABC is calculated using a weighted fusion algorithm, and the angle OBC is calculated using point B and the corresponding central angle. The length from point O to point B is calculated based on the heading angle ABC, angle OBC, and the distance traveled from point A to point B. Based on the length from point O to point B, and combined with the center coordinates and radius of circle ABO, the precise coordinates of the sound source location are finally determined.

[0056] Let the location of the sound source be denoted as point O. Based on the coordinates of points A and B, the angle of arrival of the sound source relative to point A, and the angle of arrival of the sound source relative to point B, determine the center coordinates and radius of the circle ABO formed by points A, B, and the sound source location point O, and solve for the corresponding central angle. The relationship in the corresponding process is as follows: ; in, Represents the x-coordinate of the circle's center. The x-coordinate representing the initial position. Represents the x and y coordinates of the center of the circle. The y-coordinate represents the initial position. This represents the x-coordinate of point B. This represents the ordinate of point B. Represents the radius of a circle The square of, This represents the distance from point A to point B. This represents the angle between the initial position and the center of the circle after the position has been moved; In the step of calculating the length from point O to point B based on the heading angle ABC, angle OBC, and the distance traveled from point A to point B, the corresponding relationship in the process is as follows: ; in, This represents the length from point O to point B. This represents the angle formed by the center of the circle to point B and the angle formed by point O to point B.

[0057] Furthermore, after moving two location points, the location of the sound source can be obtained. The location of the sound source can be verified by moving to a third location point D and using the same calculation method.

[0058] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0059] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0060] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A method for acoustic source localization using time delay and dead reckoning, characterized in that, The method includes the following steps: Step 1: Acquire the original sound signal and perform noise reduction preprocessing to obtain the effective sound signal after noise reduction; Step 2: For the effective sound signal after noise reduction, calculate the time delay of the sound source to reach the two array elements in the two microphone arrays using the improved generalized cross-correlation algorithm, and calculate the arrival angle of the sound source based on the time delay. Based on the symmetry of the device, obtain the directions of the two sound sources. Step 3: The two sound source directions are processed by virtual sound rotation. The effective sound signal after noise reduction is interpolated to simulate rotation. The angle difference before and after rotation is compared to eliminate false angles and determine the unique sound source direction. Step 4: Set the position when the device starts moving as the initial position and obtain the coordinates of the initial position; when the device starts moving, use the device's inertial sensors to collect the corresponding acceleration and angular velocity data, and combine them with the coordinates of the initial position to calculate the moving distance and direction of the device using a dead reckoning algorithm. Step 5: Combining the unique sound source direction, the device's moving distance and direction, and the time difference between the sound source's arrival at the two microphones, the location of the sound source is finally calculated using a geometric positioning algorithm.

2. The method of claim 1, wherein, In step 1, the original sound signal is acquired and subjected to noise reduction preprocessing to obtain the noise-reduced effective sound signal, specifically including the following steps: The original sound signal is acquired, and multi-level discrete wavelet decomposition is performed on the original sound signal. Low-pass and high-pass filters are used to extract the low-frequency approximation coefficients and high-frequency detail coefficients of the signal, respectively. The low-frequency approximation coefficients and high-frequency detail coefficients of the signal are processed using preset thresholds. Coefficients below the preset thresholds are set to zero to achieve noise suppression. At the same time, speech components above the thresholds are retained to obtain the processed approximation coefficients and processed detail coefficients. The processed approximation coefficients and processed detail coefficients are used to reconstruct the signal using a multi-resolution reconstruction algorithm to achieve noise reduction, resulting in a noise-reduced effective sound signal.

3. The method of claim 2, wherein, In step 2, for the noise-reduced effective sound signal, the time delay of the sound source arriving at the two array elements in the two microphone arrays is calculated using an improved generalized cross-correlation algorithm. The arrival angle of the sound source is then calculated based on the time delay, and the directions of the two sound sources are obtained according to the device symmetry. Specifically, the steps include the following: The device uses its microphones to receive the noise-reduced effective sound signals, thus obtaining a dual-channel microphone reception signal. The cross-correlation function of the received signals from the dual-channel microphones is calculated, and peak search is performed to obtain the estimated time delay of the sound source reaching the two microphones. The estimated time delay of the sound source to reach the two microphones, the known speed of sound, and the distance between the microphones are substituted into the formula for calculating the angle of arrival. At the same time, based on the symmetry of the equipment, the directions of the two sound sources are finally obtained.

4. The method of claim 3, wherein, The cross-correlation function of the received signals from the dual-channel microphones is calculated, and a peak search is performed to obtain the estimated time delay of the sound source arriving at the two microphones. Specifically, this includes: Perform a discrete Fourier transform on the received signals from the dual-channel microphones to obtain the frequency domain representation of the dual-channel signals; Based on the frequency domain representation of the dual-channel signal, the cross-power spectrum between the two signals is calculated and obtained; The cross-power spectrum between the two signals is weighted using the PHAT weighting function to obtain the first weighted cross-power spectrum. Perform an inverse Fourier transform on the first weighted cross-power spectrum to obtain the first generalized cross-correlation function; An improved weighting function is applied to the cross-power spectrum between the two signals to obtain a second weighted cross-power spectrum; wherein, the improved weighting function introduces a noise cross-power spectrum estimation term and an adjustable weighting factor into the denominator of the PHAT weighting function. Perform an inverse Fourier transform on the second weighted cross-power spectrum to obtain the second generalized cross-correlation function; A joint analysis of the first and second generalized cross-correlation functions is performed, and a peak search is conducted. By finding the time offset corresponding to the global maximum value of the peak, the estimated time delay of the sound source reaching the two microphones is obtained.

5. The method of claim 4, wherein, The device uses its microphones to receive the noise-reduced effective sound signals, resulting in dual-channel microphone received signals. The corresponding relationship in this process is as follows: ; wherein, represents a signal received by the first microphone, represents an independent variable time, represents a signal received by the second microphone, represents a first transfer gain of a signal to a microphone, represents a second transfer gain of a signal to a microphone, represents a sound source, represents a time of sound arrival at the first microphone, represents a time of sound arrival at the second microphone, represents a noise signal received by the first microphone, represents a noise signal received by the second microphone; In the process of substituting the estimated time delay of the sound source to the two microphones, the known speed of sound, and the distance between the microphones into the formula for calculating the angle of arrival, and considering the symmetry of the equipment, the relationship between the two sound sources is as follows: ; wherein, denotes the angle of arrival of the sound source, denotes the inverse trigonometric function, denotes the sound velocity, denotes the distance between the two microphones.

6. The sound source localization method using time delay and dead reckoning according to claim 5, characterized in that, Performing a Discrete Fourier Transform on the received signals from the dual-channel microphones yields the frequency domain representation of the dual-channel signals. The corresponding relationship in this process is as follows: ; in, This indicates the signal received by the first microphone. The signal received by the second microphone cross power spectral density, Represents a complex exponential function. This indicates the signal received by the second microphone. Representation in the frequency domain The complex conjugate, Represents angular frequency. This indicates the signal received by the first microphone. The representation in the frequency domain; In the step of applying the PHAT weighting function to the cross-power spectrum between the two signals to obtain the first weighted cross-power spectrum, the corresponding relationship in the process is as follows: ; in, This indicates the signal received by the first microphone. The signal received by the second microphone A weighted function for the numerical values ​​of the correlation functions between them; In the step of applying an improved weighting function to the cross-power spectrum between the two signals to obtain the second weighted cross-power spectrum, the corresponding relationship in the process is as follows: ; in, This represents the weighting function used in the first calculation of the relevant function values. This represents the weighting function used when calculating the correlation function value for the second time. This indicates the signal received by the first microphone. The self-power spectral density, This represents the cross-power spectral density calculated during the second cross-correlation of the two signals after weighting function processing. The cross-power spectral density represents the noise. Indicates the weighting factor; In the step of performing an inverse Fourier transform on the second weighted cross-power spectrum to obtain the second generalized cross-correlation function, the corresponding relationship in the process is as follows: ; in, This indicates the signal received by the first microphone. The autocorrelation function, This represents the autocorrelation function when calculating the autocorrelation function value for the second time.

7. The sound source localization method using time delay and dead reckoning according to claim 6, characterized in that, In step 3, the two sound source directions are processed by virtual sound rotation. The effective sound signal after noise reduction is interpolated to simulate rotation. The angle difference before and after rotation is compared to eliminate false angles and determine the unique sound source direction. The specific steps include the following: Based on the central symmetry of the device structure, the surrounding space of the two sound source directions is divided into four regions. The sound source arrival angles calculated for the two opposite regions are the same, thus obtaining two symmetrical candidate values ​​for the sound source directions. Interpolation processing is performed on the signals received by the dual-channel microphone. By changing the signal sample points, the change in the arrival time difference of the sound waves is simulated, which is equivalent to realizing the virtual rotation of the sound signal, and thus obtaining the interpolated sound signal. The arrival angle of the sound source is recalculated on the interpolated sound signal to determine the direction of the sound source after virtual rotation. The virtual rotated sound source direction is compared with two symmetrical sound source direction candidate values. Based on the difference in the change of angle in different regions, false sound source directions are identified and eliminated, and the unique true sound source direction is determined.

8. The sound source localization method using time delay and dead reckoning according to claim 7, characterized in that, In step 4, the initial position is set when the device begins to move, and the coordinates of the initial position are obtained. When the device begins to move, the corresponding acceleration and angular velocity data are collected using the device's inertial sensors. Combined with the coordinates of the initial position, the moving distance and direction of the device are calculated using a dead reckoning algorithm. The corresponding relationship in this process is as follows: ; in, Indicates the first The x-coordinate of the next position. Indicates the first The distance to the next position, Represents the sine function. Indicates the first The dead recoil angle at the next position Indicates the first The ordinate of the next position This represents the cosine function.

9. The sound source localization method using time delay and dead reckoning according to claim 8, characterized in that, In step 5, the location of the sound source is finally calculated using a geometric positioning algorithm, combining the unique direction of the sound source, the moving distance and direction of the device, and the time difference between the sound source's arrival at the two microphones. This process includes the following steps: The initial position is marked as point A. The sound source signal is collected at point A, the arrival time difference of point A is calculated, and the arrival angle of the sound source relative to point A is calculated based on the arrival time difference of point A. Keeping the equipment attitude unchanged, move the equipment from point A to the east or west to the second position point, and mark the second position point as point B. Based on the sensor data collected during the equipment movement, calculate the moving distance from point A to point B using the dead reckoning algorithm, and calculate the angle of arrival of the sound source relative to point B at point B. The location of the sound source is denoted as point O. Based on the coordinates of points A and B, the arrival angle of the sound source relative to point A, and the arrival angle of the sound source relative to point B, the center coordinates and radius of the circle ABO formed by points A, B, and the sound source location point O are determined, and the corresponding central angle is calculated. The equipment is moved eastward from point B to the third position point, which is denoted as point C. The coordinates of point C are determined. During the movement of the equipment, the heading angle ABC is calculated using a weighted fusion algorithm, and the angle OBC is calculated using point B and the corresponding central angle. The length from point O to point B is calculated based on the heading angle ABC, angle OBC, and the distance traveled from point A to point B. Based on the length from point O to point B, and combined with the center coordinates and radius of circle ABO, the precise coordinates of the sound source location are finally determined.

10. The sound source localization method using time delay and dead reckoning according to claim 9, characterized in that, Let the location of the sound source be denoted as point O. Based on the coordinates of points A and B, the angle of arrival of the sound source relative to point A, and the angle of arrival of the sound source relative to point B, determine the center coordinates and radius of the circle ABO formed by points A, B, and the sound source location point O, and solve for the corresponding central angle. The relationship in the corresponding process is as follows: ; in, Represents the x-coordinate of the circle's center. The x-coordinate representing the initial position. Represents the x and y coordinates of the center of the circle. The y-coordinate represents the initial position. This represents the x-coordinate of point B. This represents the ordinate of point B. Represents the radius of a circle The square of, This represents the distance from point A to point B. This represents the angle between the initial position and the center of the circle after the position has been moved; In the step of calculating the length from point O to point B based on the heading angle ABC, angle OBC, and the distance traveled from point A to point B, the corresponding relationship in the process is as follows: ; in, This represents the length from point O to point B. This represents the angle formed by the center of the circle to point B and the angle formed by point O to point B.

Citation Information

Patent Citations

  • Self-tracking robot sound source positioning system based on dead reckoning method

    CN114325582A

  • Directional estimation enhancement for parameterized spatial audio capture using broadband estimation

    CN114424588A

  • Improved generalized cross-correlation time delay estimation positioning method based on microphone array

    CN119126022A

  • Sound source localization method based on double-microphone rotating array

    CN119828076A

  • Moving sound image presenting device

    JP2006222801A

Cited By

  • Intelligent guide system for positioning sound source

    CN122218616A