A sound source positioning method based on a dual-microphone rotating array

Through the dual-microphone rotating array method, frequency domain signal processing and spatial grid search are used to solve the problems of high hardware cost, large computational complexity and poor real-time performance of existing sound source localization algorithms, and achieve efficient sound source localization.

CN119828076BActive Publication Date: 2025-10-21CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510034661.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-10-21
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The existing sound source localization algorithm based on arrival time delay difference has the problems of high hardware cost, large computational complexity, weak anti-noise interference ability and poor real-time performance.

Method used

A dual-microphone rotating array method is adopted. By making the two microphones rotate at a uniform speed with the rotating disk, they are equivalent to a circular array. The cross-correlation function of the frequency domain signal is calculated, and the Kalman filter and spatial grid search strategy are used to quickly locate the sound source.

Benefits of technology

It reduces hardware costs and computational load, improves anti-interference capabilities and real-time performance, and enables rapid sound source localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119828076B_ABST
    Figure CN119828076B_ABST
Patent Text Reader

Abstract

The application relates to a sound source positioning method based on a double-microphone array, and belongs to the technical field of acoustic positioning. The method comprises the following steps: collecting sound source signals received by two microphones when the two microphones rotate at a constant speed every fixed time interval, equivalently simulating the collected multiple sound source signals with time delays as a circular array composed of multiple microphones; converting the time-domain signals into frequency-domain signals, calculating the cross-correlation functions between different frequency-domain signals, and then performing weighting and filtering processing on the cross-correlation functions; performing inverse Fourier transform on the processed cross-correlation functions, so as to convert the frequency-domain information into the time domain and obtain the cross-correlation functions between the positions of virtual microphones at different time points; and adopting the strategy of coarse search first and then fine search to find the sound source position in a unit space which is subjected to grid division. The application can reduce the number of microphones used, reduce the hardware cost, greatly reduce the calculation amount, and improve the real-time performance of sound source positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of acoustic positioning and relates to a sound source positioning method based on a dual-microphone rotating array. Background Art

[0002] Microphone array sound source localization technology is maturing and playing a vital role in fields such as speech enhancement, automatic speech recognition, and intelligent robotics. Microphone arrays combine multiple microphones and perform weighted processing based on the spatial characteristics of the signal, achieving signal enhancement and interference suppression, thereby improving signal quality and performance. The accuracy of sound source localization is affected by various factors, particularly the sound source localization algorithm, which directly processes the sound data received by the microphones to accurately estimate the sound source's location. Currently, mainstream sound source localization algorithms fall into three categories: steerable beamforming, subspace sound source localization, and time delay difference of arrival.

[0003] Among them, steerable beamforming algorithms primarily use beamforming technology to process the sound signals received by the array. After removing useless frequencies through filtering, the signals are combined and weighted, and the array's beam direction is controlled to enhance signal energy in a specific direction, thereby achieving sound source localization. Commonly used maximum output power steerable beamforming algorithms include delay-and-sum beamforming and adaptive beamforming. The delay-and-sum beamforming algorithm offsets propagation delays through time delay processing, requiring minimal computation and producing low signal distortion. However, its noise immunity is poor, and more microphones are typically required for improved performance. The adaptive beamforming algorithm adapts to complex environments, but requires a high computational load and can result in distorted output signals. Reverberation must be avoided for optimal performance.

[0004] The subspace sound source localization algorithm is a sound source localization method based on array signal processing. The commonly used high-resolution spectral estimation method extends the time-domain Fourier method to the spatial domain, constructing a cross-spectral matrix between the sound source signal and the microphone array to estimate the direction of arrival (DOA). Conventional beamforming (CBF) algorithms are the earliest DOA algorithms. Although they simplify time-domain data processing, they lack resolution within the beamwidth. To improve resolution, increasing the array aperture is often necessary, but this is often not feasible in practical applications. Therefore, improving the algorithm has become a research priority.

[0005] Sound source localization algorithms based on arrival delay differences calculate the time delay differences between multiple array elements receiving the same sound signal and use these delay differences and the geometric relationship of the microphones to estimate the sound source location. The key to this method is to accurately calculate the arrival time differences between different microphones. Delay estimation methods based on cross-correlation functions are commonly used, such as the generalized cross-correlation method (GCC), the maximum likelihood weighted method (ML), and the cross-spectral phase method (CSP). Although the GCC algorithm is commonly used, discrepancies between theoretical and practical results can occur in the presence of interfering sound sources, and a large amount of data calculation is required to improve accuracy. By obtaining information about the time delay differences and the geometric position of the microphones, triangulation or least squares methods can be used to calculate the sound source location. These algorithms infer the distance between the sound source and the microphones based on time and sound speed, and determine the exact sound source location using a curve equation constructed by multiple microphones. However, errors may cause the curves determined by two microphones to not intersect at a single point. Therefore, the sound source location is often determined using a least squares fit of the intersection area of ​​multiple microphones. Generally speaking, delay difference estimation and sound source localization are two-step processes. The results reflect the past location of the sound source, so accuracy is affected by both the delay difference estimation and the microphone geometry. In the presence of multiple sound sources, delay difference estimation becomes even more complex, and the geometry of the microphone array must be optimized to accommodate these requirements. Furthermore, the algorithm's accuracy is closely related to the signal sampling rate, requiring continuous sampling and computational fitting to improve localization accuracy.

[0006] In summary, the existing sound source localization algorithm based on arrival time delay difference has the problems of using multi-microphone arrays, high hardware cost, large placement space, and inconvenient movement. At the same time, the calculation of the cross-correlation values ​​of all grid points is computationally intensive, the anti-noise interference ability is weak, and the real-time performance is poor. Summary of the Invention

[0007] In view of this, the purpose of the present invention is to provide a sound source localization method based on a dual-microphone rotating array, so as to reduce the number of microphones used, reduce the amount of calculation of the cross-correlation value, improve the anti-interference ability of the cross-correlation function, and at the same time accelerate the sound source position search speed and improve real-time performance.

[0008] In order to achieve the above object, the present invention provides the following technical solutions:

[0009] A sound source localization method based on a dual-microphone array, the method comprising:

[0010] The two microphones are made to rotate at a constant speed along with the rotating disk. The sound source signals received by the two microphones are collected at fixed intervals. The multiple collected sound source signals with time delays are regarded as a group, which is equivalent to simulating a circular array composed of multiple microphones.

[0011] Convert the collected multiple time domain signals into frequency domain signals, calculate the cross-correlation function between different frequency domain signals, and then perform weighted and filtering processing on the cross-correlation function;

[0012] Performing an inverse Fourier transform on the processed cross-correlation function to convert the frequency domain information into the time domain, and obtaining the cross-correlation function between the virtual microphone positions at different time points;

[0013] A unit space is gridded using a triangular recursive method, and a simulated microphone array is set at the origin of the unit space. The time delay from any grid point in the unit space to each pair of microphones is obtained, and the cross-correlation value of each microphone pair is calculated based on the time delay and the signal received by the microphone pair. The cross-correlation values ​​of each microphone pair are accumulated as the energy value of the grid point. After assigning energy values ​​to all grid points in the unit space, the grid point with the largest energy value among all grid points is searched, and the grid point with the largest energy value is the location of the sound source.

[0014] The fixed time is the time it takes for the microphone to rotate 30 degrees. The microphone collects a sound source signal every 30 degrees of rotation. During the 360-degree rotation of the rotating disk, 12 time-delayed sound source signals are collected, effectively simulating a circular array consisting of 12 microphones.

[0015] Furthermore, the multiple time domain signals collected are converted into frequency domain signals, and then the cross-correlation function between different frequency domain signals is calculated, including: for the multiple time domain signals collected, they are first divided into L frames, and then the L frame signals corresponding to each time domain signal are converted into the frequency domain; and then the cross-correlation function between different frequency domain signals is calculated by the following formula:

[0016]

[0017] Where, X i [k] and X j [k] are the kth frame time domain signals x of virtual microphone i i (k) and the k-th frame time domain signal x of virtual microphone j j (k) is the discrete Fourier transform, τ represents x i (k) and x j (k) the time delay between (X i [k]X j [k]) * is x i (k) and x j The cross spectrum of (k), (·) * represents the complex conjugate.

[0018] Furthermore, the cross-correlation function is weighted by the weighting function shown in the following formula:

[0019]

[0020] Where S(k) is the cross-correlation function between different frequency domain signals in the kth frame, k = 0, 1, ..., L-1, and L represents the number of frames for framing the acquired time domain signal.

[0021] A Kalman filter is added every time the rotating disk completes a 360° rotation, and the corresponding weight is calculated for each cross-correlation function, as shown in the following formula:

[0022]

[0023] Where, Indicates the weighted prediction value calculated at the last moment, S n (k) represents the cross-correlation function between different frequency domain signals of the currently predicted k-th frame; N n (·) indicates the weight assignment operation.

[0024] Furthermore, after assigning energy values ​​to all grid points in the unit space, the four grid points with the largest energy values ​​among all grid points are searched first to preliminarily define the area where the sound source is located; then, the grid points are further divided within the defined area to improve the resolution of the area;

[0025] For the area with improved resolution, an energy value is reassigned to each grid point in the area, and the grid point with the largest energy value is searched again in the limited area, which is the sound source position.

[0026] The beneficial effects of the present invention are as follows: by simulating the signals collected at different time points as equivalent to a circular array composed of multiple microphones and then calculating the cross-correlation values ​​of the virtual microphone elements, the present invention can effectively reduce hardware costs, improve practicality, and reduce the impact of sound source propagation barriers on the sound source localization effect. In addition, the present invention uses different resolutions to divide the space into grids, pre-calculates the time delay values ​​from the grid position to each pair of microphones, and adopts a strategy of first performing a coarse search and then a fine search when searching for the sound source location. This greatly reduces the amount of calculation, effectively speeds up the sound source localization, and improves the real-time performance of sound source localization.

[0027] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0029] Figure 1 is a flow chart of a sound source localization method shown in an exemplary embodiment of the present application;

[0030] Figure 2 is a schematic diagram of a handheld rotating microphone shown in an exemplary embodiment of the present application;

[0031] Figure 3 is a schematic diagram of a simulated circular microphone array shown in an exemplary embodiment of the present application;

[0032] Figure 4 is a spatial grid division diagram shown in an exemplary embodiment of the present application;

[0033] Figure 5 is a schematic diagram of microphone directivity shown in an exemplary embodiment of the present application;

[0034] Figure 6 is a diagram of a simulated gain function shown in an exemplary embodiment of the present application;

[0035] Figure 7 This is a schematic diagram of a unit sphere search strategy shown in an exemplary embodiment of the present application.

[0036] Reference numerals: 1-motor module; 2-acrylic plate; 3-MEMS microphone; 4-communication module; 5-handheld handle. DETAILED DESCRIPTION

[0037] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0038] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0039] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0040] like Figure 2 The figure shows a handheld rotating microphone device according to an embodiment of the present invention. The device comprises a motor module 1, a circular acrylic plate 2, a MEMS microphone 3, a communication module 4, and a handheld handle 5. The motor module 1 includes a motor and a battery, responsible for the motor's rotation. The acrylic plate 2 is used to secure the microphone. The MEMS microphone 3 receives sound signals and converts analog signals into digital signals. The communication module 4 is used to remotely transmit the sound signals output by the microphone to a computer.

[0041] The diameter of the circular acrylic plate 2 is 10 cm. Two MEMS microphones 3 are symmetrically fixed on one diameter of the circular acrylic plate 2 and are located at the edge of the circular acrylic plate 2. The sound collection point of the MEMS microphone 3 is located at the outermost side of the circular acrylic plate 2, thereby enhancing the spatial resolution.

[0042] Based on the above rotating microphone device, the sound source localization method based on the dual-microphone rotating array proposed in the present invention is as follows Figure 1 As shown, it includes the following steps:

[0043] N1: Two microphones receive sound signals at different angles as the rotating disk rotates. The collected analog signals are converted to digital signals by the sound card and transmitted to the computing device. The preprocessing module frames the time domain signals collected by each microphone at different time points. The signals collected at each rotation are divided into L = 1024 independent frames.

[0044] In the computing device, the signal is first processed by a power complementary window to reduce the impact of spectral leakage, and then the time domain signal is converted into a frequency domain signal by fast Fourier transform (FFT). Among them, the signals obtained at different time points can be regarded as receiving sound source signals from multiple virtual microphone positions at the same time. Specifically, according to the rotation speed of the circular acrylic plate, the time interval required for the microphone to rotate 30° is calculated, and a total of 12 time intervals are obtained. Therefore, the 12 delayed signals are regarded as a group, which is equivalent to simulating an array consisting of 12 microphones. Figure 3The figure shows a simulated circular array formed by two microphones after rotation. The simulated circular array is formed by twelve microphones, and the angle between each two adjacent microphones and the line connecting the center of the circle is 30°.

[0045] The output of the M-microphone delay-sum beamformer is defined as:

[0046]

[0047] Where M = 2, x m (n) is the time domain signal collected by the mth microphone, τ m is the time delay when the sound source arrives at the mth microphone relative to the first microphone. According to the output of the delay-sum beamformer, the output energy of the beamformer in the frame length of L = 1024 can be obtained:

[0048]

[0049] From formula (2), we can see that if the maximum value of E is required, then it is only necessary that the signals of the microphones are in the same direction, that is, the peak value of each microphone signal corresponds to the peak value. Expanding the energy formula using the square sum formula, we get:

[0050]

[0051] When calculating the maximum value of E in formula (3), the square term on the right side of the equal sign is approximately a constant and can be ignored.

[0052] In order to reduce the amount of calculation and facilitate the later whitening processing of the signal, the time domain signal is converted into a frequency domain signal, and the cross-correlation value between the frequency domain signals of the two microphones is calculated to obtain:

[0053]

[0054] Among them, X i [k] is x i The discrete Fourier transform of (k), (X i [k]X j [k]) * is x i (k) and x j The cross spectrum of (k), (·) * represents the complex conjugate, R ij(τ) represents the cross-correlation function between the frequency domain signals of virtual microphone i in the kth frame and the frequency domain signals of virtual microphone j in the kth frame. For signals sampled at a 16kHz sampling rate, overlapping windows of length L = 1024 sampling points (50% overlap) are used to calculate the power spectrum and cross-power spectrum for each window. By calculating the cross-correlation value between the frequency domain signals of two microphones, the degree of similarity between the two signals can be assessed. By finding the maximum value in the cross-correlation function, the time delay τ of the sound source signal can be determined. This time delay reflects the difference in the time delay between the sound source reaching each microphone, providing key data for subsequent sound source localization.

[0055] N2: After calculating the frequency domain signals of the microphones at different time points, the time series relationship of these signals is used to rewrite the frequency domain summation function of the signals to obtain the cross-correlation function based on the signals collected at different time points (as shown in Equation (4)). By whitening these cross-correlation functions to reduce the impact of environmental noise, and introducing a weighting function to further suppress the noise, an optimized cross-correlation function is obtained. Then, an inverse Fourier transform is performed on the optimized cross-correlation function to convert the frequency domain information back to the time domain. The cross-correlation function between the virtual microphone positions at different time points is obtained, providing accurate time delay estimation information for sound source localization.

[0056] The cross-correlation value R in formula (4) ij (τ) is the average cross power spectrum (X i [k]X j [k]) * Calculated. Pre-calculated R ij (τ) can be calculated using only M(M-1) / 2 lookup and accumulation operations, while calculating E in the time domain requires 2L(M+2) operations. For M=2 and 2562 directions, the complexity of the search itself is reduced from 2.4Gflops to only 3.5Mflops. After computing all time-frequency transforms, the complexity is only 96.8Mflops, which is 25 times lower than the time-domain search at the same resolution.

[0057] In order to facilitate the observation and comparison of the calculated results, equation (4) is normalized:

[0058]

[0059] Normalized cross-correlation produces sharper cross-correlation peaks, but it also has a significant drawback: each frequency domain of the spectrum has the same impact on the final cross-correlation result. This means that noise has a significant impact, making the system less resilient to noise and making it more difficult to detect sound sources. To address this issue, a weighting function is introduced based on the signal-to-noise ratio, as shown in the following formula:

[0060]

[0061] Wherein, S(k) is the cross-correlation function of the two microphones on the frequency domain signal of the kth frame.

[0062] At each update (i.e., every time the circular acrylic plate rotates 360° and the microphone starts collecting signals for the next cycle), a Kalman filter is added to calculate the corresponding weight for each signal, as shown in the following formula:

[0063]

[0064] in, is the weighted prediction value calculated at the previous moment, S n (k) is the cross-correlation function of the different frequency domain signals of the currently predicted k-th frame. N n (·) indicates the weight assignment operation.

[0065] N3: After obtaining the optimized cross-correlation function, the sound source location estimation strategy first searches for the most likely location of the sound source on the simulated circular microphone array through a spatial search algorithm based on the time delay information between these virtual microphone positions and the time delay differences collected by the microphones at various angles on the rotating disk.

[0066] This strategy makes full use of the spatial characteristics of virtual array coverage and can quickly converge to the optimal position of the sound source with relatively small computing resources, which is suitable for the needs of real-time sound source localization.

[0067] In order to facilitate the search for the sound source location and reduce the amount of calculation, a unit space is defined and divided into a uniform triangular grid to obtain a spatial grid for sound source localization. To create the spatial grid, we start with an initial 20-sided spherical grid. Each triangle in the initial 20-element grid is recursively subdivided into 4 smaller triangles. Each smaller triangle after recursive subdivision is recursively subdivided again, as shown in the following example: Figure 4 The final mesh consists of 5120 triangles and 2562 points. The accuracy of the mesh is determined by the number of recursions. To quickly search for the sound source location, the mesh is divided twice, corresponding to the coarse search and the fine search, respectively.

[0068] The microphone array is set at the origin of the unit space. Based on the coordinates of each grid point and the position parameters of the microphones, the time delay difference from any grid point to all microphone pairs can be calculated. The cross-correlation value of each pair of microphones can be calculated based on the time delay difference and the microphone signals. In microphone arrays, it is generally assumed that the microphones are omnidirectional, that is, they acquire signals from all directions with the same gain. However, in practice, the microphones are mounted on a rigid body, which may block the direct propagation path between the sound source and the microphones, causing attenuation of the sound source signal. This attenuation is mainly caused by diffraction and varies with frequency.

[0069] Since an exact diffraction model is unavailable, the proposed model relies on simpler assumptions: (1) a source with a direct propagation path has unity gain; and (2) the gain is zero when the path is blocked by an object. Since the signal-to-noise ratio of an obstructing microphone is usually unknown, it is safer to assume a lower signal-to-noise ratio, and setting the gain to zero prevents noise from being injected into the observation. In addition, the gain is constant at all frequencies, and a smooth transition band connects the unity gain region and the zero gain region. This transition band prevents abrupt changes in gain when the sound source position changes. Figure 5 θ(u,d) is shown, which represents the angle between the sound source at position u and the direction of the microphone modeled by the unit vector d. θ(u,d) is expressed as follows:

[0070]

[0071] Let the analog gain G(u,D) be a function of θ(u,d), as shown below:

[0072]

[0073] Where D is a set of parameters {d, α, β}, α represents the angle where the gain is 1, and β represents the angle where the gain is zero. The area between α and β can be regarded as a transition zone, such as Figure 6 shown.

[0074] To make sound source localization more reverberant, the scan space is restricted to a specific direction. For example, the scan space is restricted to a hemisphere pointing toward the ceiling to ignore reflections from the floor. This also reduces the number of search points and speeds up the search.

[0075] Based on the spatial grid constructed above and the assumptions, after calculating the cross-correlation relationship of each pair of microphones, we begin to search each grid point of the spatial grid in a loop to search for the optimal direction of the sound source. First, the energy value of each grid point is set to 0, and the microphone pair i, j is looped to obtain the pre-calculated time delay τ from each grid point to each pair of microphones, and the cross-correlation value of each pair of microphones is calculated. The cross-correlation value of each microphone team The energy value E accumulated to the grid point p p Therefore, the energy value of each grid point is the sum of the cross-correlation values ​​of all microphone pairs, and the direction of the sound source is the grid point with the largest energy value.

[0076] The search strategy is to find the grid point with the largest energy value as quickly as possible, such as Figure 7 As shown in the figure, the approximate sound source location is first searched within the lower-resolution sphere A. After the approximate sound source location is found, the refined grid points near the approximate sound source point are searched, as shown in sphere B. Specifically, the lower-resolution sphere is first divided for a coarse search, searching for the point with the maximum energy value (denoted as the first grid point) within the sphere's grid points. This is the approximate direction of the sound source. A high-resolution grid is then divided within a certain space around the first grid point, and the energy values ​​of each grid point within this space are recalculated. A fine search is then performed within this space, again searching for the grid point with the maximum energy value, which is the precise location of the sound source. This search strategy effectively reduces computation while improving positioning accuracy.

[0077] At the same time, in order to draw a visual heat map, four cycles are performed when searching for the sound source location in both the coarse search and fine search stages. The first cycle first traverses all grid points, finds the maximum value point, outputs the position coordinate information and energy value of this point, and then removes this point from the space. Then, the second cycle begins, which also traverses all grid points to find the maximum value. The maximum value found in the second cycle is equivalent to the second largest value point overall. The position coordinates and energy value of this point are also output, and the point is removed from the space. The third cycle performs the same operation, and in the fourth cycle, the grid point with the largest energy value in the space is found and the position coordinate information of this point is output. After four cycles of the coarse search and fine search stages, eight points are obtained, and a visual heat map is drawn based on the energy values ​​of these eight points.

[0078] As a practical application scenario of the present invention, the present invention can be used to detect the broken part of the motor. When performing the detection, the operator holds the Figure 2 The audio and video device shown in the figure is positioned at a certain distance from the motor under test. The audio and video device has a built-in camera that captures real-time image information from the motor surface and transmits the captured images and related data wirelessly or wiredly to a computer for processing and display. At the same time, the audio and video device uses a built-in sound source localization module to synchronously detect abnormal sound source signals emitted during motor operation. Combined with the camera image information, it accurately identifies and locates the motor's broken part, providing the operator with real-time, intuitive fault diagnosis assistance. This device is designed to be portable and suitable for rapid detection needs in complex industrial scenarios.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A sound source localization method based on a dual-microphone array, characterized in that: The method includes: The two microphones are made to rotate at a constant speed along with the rotating disk. The sound source signals received by the two microphones are collected at fixed intervals. The multiple collected sound source signals with time delays are regarded as a group, which is equivalent to simulating a circular array composed of multiple microphones. Convert the collected multiple time domain signals into frequency domain signals, calculate the cross-correlation function between different frequency domain signals, and then perform weighted and filtering processing on the cross-correlation function; Performing an inverse Fourier transform on the processed cross-correlation function to convert the frequency domain information into the time domain, and obtaining the cross-correlation function between the virtual microphone positions at different time points; A unit space is gridded using a triangular recursive method, and a simulated microphone array is set at the origin of the unit space. The time delay from any grid point in the unit space to each pair of microphones is obtained, and the cross-correlation value of each microphone pair is calculated based on the time delay and the signal received by the microphone pair. The cross-correlation values ​​of each microphone pair are accumulated as the energy value of the grid point. After assigning energy values ​​to all grid points in the unit space, the grid point with the largest energy value among all grid points is searched, and the grid point with the largest energy value is the location of the sound source.

2. The sound source localization method according to claim 1, wherein: The fixed time is the time required for the microphone to rotate 30°.

3. The sound source localization method according to claim 2, wherein: The microphone collects a sound source signal every time it rotates 30°. During the 360° rotation of the rotating disk, 12 sound source signals with time delays are collected, which is equivalent to simulating a circular array consisting of 12 microphones.

4. The sound source localization method according to claim 1, wherein: The multiple acquired time domain signals are converted into frequency domain signals, and then the cross-correlation function between different frequency domain signals is calculated, including: firstly dividing the multiple acquired time domain signals into L frames, and then converting the L frame signals corresponding to each time domain signal into the frequency domain; and then calculating the cross-correlation function between different frequency domain signals using the following formula: Where, X i [k] and X j [k] are the kth frame time domain signals x of virtual microphone i i (k) and the k-th frame time domain signal x of virtual microphone j j (k) is the discrete Fourier transform, τ represents x i (k) and x j (k) the time delay between (X i [k]X j [k]) * is x i (k) and x j The cross spectrum of (k), (·) * represents the complex conjugate.

5. The sound source localization method according to claim 1 or 4, characterized in that: The cross-correlation function of the frequency domain signal is weighted and filtered, including: weighting the cross-correlation function by a weighting function shown in the following formula: Where S(k) is the cross-correlation function between different frequency domain signals in the kth frame, k = 0, 1, ..., L-1, and L represents the number of frames for framing the acquired time domain signal. A Kalman filter is added every time the rotating disk completes a 360° rotation, and the corresponding weight is calculated for each cross-correlation function, as shown in the following formula: Where, Indicates the weighted prediction value calculated at the last moment, S n (k) represents the cross-correlation function between different frequency domain signals of the currently predicted k-th frame; N n (·) indicates an empowerment operation.

6. The sound source localization method according to claim 1, wherein: After assigning energy values ​​to all grid points in the unit space, first search for the grid point with the largest energy value among all grid points to preliminarily define the area where the sound source is located; then continue to divide the grid points in the defined area to improve the resolution of the area; For the area with improved resolution, an energy value is reassigned to each grid point in the area, and the grid point with the largest energy value is searched again in the limited area, which is the sound source position.

Citation Information

Patent Citations

  • Sound localization method and system

    CN101201399A

  • Sound source positioning device and method based on track moving double microphone arrays

    CN104459625A