Indoor multi-sound-source positioning method based on DOA estimation and DOA association
By combining speech activity detection and FSSZ correlation between arrays, a sound source position histogram is constructed, which solves the positioning accuracy and missed detection problems in indoor multi-sound source positioning, and achieves high-precision and low-complexity sound source positioning.
Patent Information
- Application Number
- CN202510238482.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-27
AI Technical Summary
The existing indoor multi-sound source positioning technology lacks positioning accuracy under reverberation and noise conditions, and the missed detection problem in the case of overlapping multiple sound sources and DOA cannot be effectively solved.
By using a method based on DOA estimation and DOA correlation, speech activity detection is carried out to construct elastic single-source time-frequency regions with variable lengths, combined with the correlation of FSSZ between arrays, sound source position estimation is used using circular integral cross spectrum and 2D-CFAR algorithm to construct sound source position histogram to improve positioning accuracy and robustness.
The centimeter-level positioning accuracy is achieved under indoor reverb and noise conditions, which reduces the computational complexity, and improves the robustness of the overlapping DOA of multiple sound sources and reduces the calculation amount.
Smart Images

Figure CN120214697A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microphone array data processing, and in particular to an indoor multi-source sound source localization method based on DOA estimation and DOA association. Background Technique
[0002] In the information society, sound source localization technology has a wide range of applications in multiple fields, such as hearing aids, video conferencing, and robot auditory systems. The position information can be used to align the camera with the speaker, enhance specific sound sources, and track wild animals, etc. Humans judge the sound source position through the "binaural effect" of the auditory system, judge the angle and distance of the sound source by the intensity and time difference of the signals heard by both ears, and judge the height of the sound source by the pinna effect caused by the reflection of the sound signal by the outer ear structure. A microphone array is composed of multiple microphone elements arranged in a specific geometric structure. In a sound source localization system, by the similarity and difference of the audio signals collected by each element, combined with the geometric distribution of the elements, and supplemented by a variety of algorithms, it can simulate the positioning effect of the human ear and detect information such as the number, angle, and position of the sound source.
[0003] Sound source localization methods can be mainly divided into the following three categories: The first category is the method based on beamforming. This method locates the sound source by searching for the peak of the spatial power spectrum. Among them, the SRP-PHAT algorithm based on phase weighting is widely used. The method based on beamforming has strong anti-reverberation performance, but is sensitive to noise and has insufficient detection performance for multiple active sound sources. The second category is the method based on deep learning. This method uses a large amount of data to train a neural network. Commonly used network architectures include convolutional neural network (CNN) and convolutional recurrent neural network (CRNN). The method based on deep learning can automatically extract various features in the sound signal by the network, but this method has a large amount of calculation and usually requires the number of sound sources to be known. The third category of methods is the method based on DOA estimation and DOA association. This method first measures the DOA estimation of each sound source relative to each array, then finds the corresponding measurements of each sound source in different arrays through the DOA association method, and finally obtains the sound source position through triangulation. In the DOA estimation method, the method based on the time delay difference (TDOA) of the sound signal arriving at the element is relatively commonly used. In the DOA association method, histograms and inter-channel phase differences (IPD) are usually used as association features. Summary of the Invention
[0004] The purpose of the present invention is to provide an indoor multi-source sound source localization method based on DOA estimation and DOA association.
[0005] The technical solution for realizing the purpose of the present invention is: An indoor multi-source sound source localization method based on DOA estimation and DOA association, including the following steps:
[0006] Step 1: Perform voice activity detection, and identify voiced frames based on the energy threshold and short-time zero-crossing rate detection method;
[0007] Step 2: Within the voiced frames identified in Step 1, taking one frame of data as a unit, judge the single-source time-frequency points through the correlation of the phase differences between different microphone array elements in the array;
[0008] Step 3: Within each frame, according to the SSP identified in Step 2, construct an elastic single-source region with variable length;
[0009] Step 4: According to the highly correlated characteristics of the FSSZ time-frequency distribution of the same sound source at different microphone arrays within one frame, apply the FSSZ distribution of this sound source at the reference array to all arrays;
[0010] Step 5: Adopt the circular integral cross-spectrum method to perform intermediate DOA estimation for each SSP, combine with the FSSZ aggregation degree in Step 3 for weighting, and construct a sound source position histogram;
[0011] Step 6: Use 2D-CFAR to perform peak detection on the sound source position histogram in Step 5 to obtain the 2D positions of multiple sound sources.
[0012] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above indoor multi-source sound source localization method based on DOA estimation and DOA association.
[0013] A computer-readable storage medium has a computer program stored thereon. When the program is executed by a processor, it implements the above indoor multi-source sound source localization method based on DOA estimation and DOA association.
[0014] A computer program product includes a computer program. When the computer program is executed by a processor, it implements the above indoor multi-source sound source localization method based on DOA estimation and DOA association.
[0015] Compared with the prior art, the significant advantages of the present invention are as follows: 1) The present invention can localize multiple sound sources with unknown quantities under indoor reverberation and noise conditions, and the localization accuracy can reach the centimeter level; 2) The present invention is robust to the "missed detection" situation caused by the DOA coincidence of multiple sound sources relative to a certain microphone array, making up for the deficiencies of the prior art; 3) The present invention makes full use of the correlation of FSSZ between arrays, applies the FSSZ distribution of the reference array to the entire positioning system, greatly reduces the calculation, and improves the positioning speed. Description of the Drawings
[0016] Figure 1 It is a flowchart of the indoor multi-source sound source localization method based on DOA estimation and DOA association of the present invention.
[0017] Figure 2 Schematic diagram of a six - element uniform circular microphone array in the embodiment.
[0018] Figure 3 Schematic diagram of the FSSZ distribution of arrays M1 and M2 in the embodiment.
[0019] Figure 4 Schematic diagram of the binary sub - array structure in the embodiment.
[0020] Figure 5 Histogram of sound source positions in the embodiment. Detailed implementation manners
[0021] The present invention will be further described below in conjunction with the accompanying drawings of the specification.
[0022] Sound source localization has always been an important branch in the field of microphone array signal processing. Indoor multi - sound source localization technology is used in many fields such as video conferencing, hearing aids, and robot auditory systems. Existing indoor multi - sound source localization technologies have problems such as insufficient anti - noise and anti - reverberation performance, requiring the number of sound sources to be known, and poor performance in the case of "missed detection". The present invention constructs a single - source time - frequency region with variable length, combines aggregation - degree weighting to form a sound source position histogram, enhances the anti - noise and anti - reverberation performance of the localization method, and at the same time increases the robustness to the "missed detection" situation. Utilizing the correlation of FSSZ between arrays, the sharing of FSSZ distribution conditions is realized, and the computational complexity is effectively reduced on the premise of ensuring the localization accuracy.
[0023] Combined with Figure 1 , an indoor multi - sound source localization method based on sound source direction - of - arrival (DOA) estimation and DOA association, includes the following steps:
[0024] Step 1: Perform voice activity detection (VAD), and identify the voiced frames based on the energy threshold and short - time zero - crossing rate detection (STZCR) method.
[0025] Step 2: Within the voiced frames identified in Step 1, taking one - frame data as a unit, judge the single - source time - frequency points (SSP) through the correlation of the phase difference (IPD) between different microphone elements in the array.
[0026] In an M - element circular microphone array, at a certain specific SSP, the following formula is satisfied:
[0027]
[0028] Where X i (n,k) represents the time - frequency domain representation of the audio signal received by the i - th element in the array, ω(n,k) represents the frequency corresponding to the time - frequency point (n,k), τ i,j(n,k) is denoted as the TDOA between microphones m on SSP(n,k). i and m j Let the TDOA between them be τ. F represents the number of FFT points in the STFT, angle(θ) represents the angle of the complex number θ, and "*" represents the conjugate. Thus, the time-delay vector on this SSP is obtained as follows:
[0029] τ(n,k) = [τ 0,1 (n,k), τ 1,2 (n,k), …, τ M-1,0 (n,k)] T
[0030] where [·] T denotes the transpose. In the n-th frame, define the time-delay vector correlation coefficients e n k-1,k and e n k,k+1 for three consecutive time-frequency points (n,k - 1), (n,k), and (n,k + 1):
[0031]
[0032] where corrcoef(x1,x2) is used to calculate the correlation coefficient between x1 and x2, n represents the time-domain index, k represents the frequency-domain index, τ(n,k) is the time-delay vector at the time-frequency point (n,k), and N represents the total number of time frames of the audio to be processed. When e n k-1,k and e n k,k+1 meet the following conditions, the three adjacent time-frequency points are judged as single-source time-frequency points:
[0033]
[0034] where δ is a small positive number.
[0035] Step 3: In each frame, according to the SSP identified in Step 2, construct a flexible single-source zone (FSSZ) with variable length.
[0036] In the n-th frame, the r-th (1 ≤ r ≤ B n ) FSSZ can be represented as:
[0037] Φ (n,r) = {(n,k r ),(n,k r + 1),...,(n,k r + D r - 1)}
[0038] where D ris the number of SSPs in the FSSZ, expressed as the aggregation degree of SSPs in this area, k r is the serial number of the first SSP in the current FSSZ, B n is expressed as the number of FSSZs in the nth frame. To increase the estimation accuracy of FSSZs and prevent non-single-source time-frequency points from being mixed in, it is required that the FSSZ contains at least two consecutive groups of adjacent-frequency SSPs estimated in step 2.
[0039] Step 4: According to the highly correlated characteristics of the FSSZ time-frequency distribution of the same sound source at different microphone arrays within one frame, apply the FSSZ distribution of the sound source at the reference array to all arrays.
[0040] Step 5: Adopt the Circular Integral Cross-Spectrum (CICS) method to perform intermediate DOA estimation for each SSP, and combine the FSSZ aggregation degree in step 3 for weighting to construct a sound source position histogram.
[0041] In a circular microphone array, given a pair of adjacent microphones m i and m j , the following relationship is satisfied:
[0042] τ i→0 (φ) = τ 0,1 (φ) - τ i,j (φ), i = 0, 1, … M - 1, j = mod(i + 1, M)
[0043] where φ ∈ [0, 2π) is the candidate angle of the DOA. Calculate the phase rotation factor:
[0044]
[0045] Calculate the CICS:
[0046]
[0047] where
[0048]
[0049] Calculate the angle estimation of the current SSP(n, k):
[0050]
[0051] where represents the narrowband DOA estimation at the time-frequency point (n, k), CICS (n,k) (φ) represents the circular integral cross-spectrum value in the direction φ, G i,j (n, k) represents the phase transformation value at the time-frequency point (n, k), represents the phase rotation factor in the direction φ, τi→0 Let \(\tau(\varphi)\) denote the time delay difference between the \(i\)-th array element and the 0-th array element at direction \(\varphi\), where \(\varphi\in[0, 2\pi)\) is the candidate angle of DOA.
[0052] To reduce the computational complexity of CICS, first, a narrowband far-field beamforming algorithm is used in combination with the geometric characteristics of the circular array to perform a preliminary estimation of DOA, and the estimated value is Then the search range of CICS can be reduced to
[0053] For an \(M\)-element uniform circular microphone array, every two adjacent array elements can be regarded as a binary sub-array, satisfying:
[0054] \(\tau\) i,j \(\tau(n,k)\cdot c = l\cdot\sin(\theta\) i,j ), \(i = 0,1,\cdots,M - 1\), \(j=\text{mod}(i + 1,M)\)
[0055] where \(c\) is the speed of sound, \(l\) is the distance between two array elements, and a rough DOA estimate can be obtained for each sub-array:
[0056]
[0057] Combining the rough DOA estimates of all \(M\) binary sub-arrays in an array, the estimated value of the sound source DOA can be obtained
[0058] \(i = 0,1,\cdots,M - 1\), \(j=\text{mod}(i + 1,M)\)
[0059] \(\text{angle}(X)\) represents the phase angle of \(X\), represents the rough DOA estimate of the binary sub-array composed of the \(i\)-th array element and its adjacent array element \(j=\text{mod}(i + 1,M)\) in the array, \(X\) i Let \(X(n,k)\) denote the time-frequency domain representation of the audio signal received by the \(i\)-th array element in the array, \(*\) represents conjugate, \(M\) is the number of array elements in an array, \(c\) is the speed of sound, \(l\) is the distance between array elements, \(\omega(n,k)\) represents the frequency corresponding to the time-frequency point \((n,k)\), and then the CICS method is used to obtain the intermediate DOA estimate:
[0060]
[0061] Let \(\Delta\theta\) denote the search range centered at
[0062] In step 4, the FSSZ distribution of the reference array is applied to all arrays. In one frame, the CICS method is used to obtain the estimated value of the sound source DOA of the SSP in the FSSZ of all arrays, and the sound source position histogram of the \(n\)-th frame is constructed accordingly:
[0063]
[0064] where y (n,r) (u, v) represents the contribution value of the r-th FSSZ in the n-th frame to the current frame-level histogram. (u, v) represents a small block interval in 3D space. In the entire 3D space, the total number of small block intervals is U × V, D r is the number of SSPs in the current FSSZ, M b (u, v) refers to the value of the interval (u, v) in the 3D histogram:
[0065]
[0066] where refers to the position estimation of the sound source by the current SSP. Considering all frames, the number D of SSPs in the current FSSZ is used r for weighting to construct a frame-level sound source position histogram:
[0067]
[0068] where y n (u, v) represents the value of the frame-level sound source histogram at the coordinate (u, v) in the n-th frame. U and V respectively represent the number of divisions of the x-axis and y-axis, y (n,r) (u, v) represents the contribution value of the r-th FSSZ in the n-th frame to the current frame-level histogram. (u, v) represents a small block interval in 3D space, D r is the number of SSPs in the current FSSZ, M b (u, v) refers to the value of the interval (u, v) in the 3D histogram.
[0069] Step 6: Use 2D-CFAR to perform peak detection on the sound source position histogram in Step 5 to obtain the 2D positions of multiple sound sources.
[0070] The sliding window size is G × G, the guard cell size is Q × Q, and the reference cells are represented as X = (x1, x2,... x U ) and U = G 2 - Q 2 . The threshold factor P fa is the false alarm probability, is the average clutter intensity, where x j refers to the clutter intensity at the reference cell j. η CUT is the signal intensity of the cell to be detected, and the decision strategy is where H1 represents that there is a sound source in the cell to be detected, and H0 represents that there is no sound source in the cell to be detected.
[0071] The present invention will be further described in conjunction with the embodiments as follows:
[0072] Embodiment
[0073] In conjunction with Figure 1 , an indoor multi-source sound localization method based on DOA estimation and DOA association, comprising the following steps:
[0074] Step 1: Perform voice activity detection (VAD), and identify the voiced frames based on the energy threshold and short-time zero-crossing rate detection (STZCR) method. In this embodiment, two uniform circular six-element microphone arrays M1 and M2 are used, and the single microphone array model is as Figure 2 shown, and are respectively placed at (0m, 0m) and (0m, 1.5m). The column radius is 0.0475m, the sampling frequency is 44.1kHz, the number of FFT points is 2048, the time overlapping frame rate is 50%, the frame length is 1024, the frequency range is 0.3kHz - 3.6kHz, and the cumulative number of frames is 86 frames (1s). Three simultaneously active speaker sound sources S1, S2, and S3 are placed at (0.3m, 1.5m), (1.7m, 0.4m), and (0.5m, -1.6m), and each audio is randomly selected and played in the NTT database.
[0075] Step 2: Within the voiced frames identified in Step 1, taking one frame of data as a unit, judge the single-source time-frequency points (SSP) through the correlation of the phase difference (IPD) between different microphone elements within the array.
[0076] In an M-element circular microphone array, at a certain specific SSP, the following formula is satisfied:
[0077]
[0078] where X i (n,k) represents the time-frequency domain representation of the audio signal received by the i-th element in the array, ω(n,k) represents the frequency corresponding to the time-frequency point (n,k), τ i,j (n,k) represents the TDOA between microphone m i and m j at the SSP (n,k), F represents the number of FFT points in the STFT, angle(θ) represents the angle of the complex number θ, and "*" represents the conjugate. Thus, the delay vector at this SSP is obtained:
[0079] τ(n,k) = [τ 0,1 (n,k), τ 1,2 (n,k), …, τ M-1,0 (n,k)] T
[0080] where [·] TDenotes transpose. In the n-th frame, define the time-delay vector correlation coefficient e of three consecutive frequency-time points (n, k-1), (n, k), and (n, k+1). n k-1,k and e n k,k+1 :
[0081]
[0082] where corrcoef(x1, x2) is used to calculate the correlation coefficient of x1 and x2. When e n k-1,k and e n k,k+1 meet the following conditions, the three adjacent frequency-time points are judged as single-source frequency-time points:
[0083]
[0084] In this example, set δ to 0.008.
[0085] Step 3: Within each frame, according to the SSPs identified in Step 2, construct a flexible single-source zone (FSSZ) with variable length.
[0086] In the n-th frame, the r-th (1 ≤ r ≤ B n ) FSSZ can be expressed as:
[0087] Φ (n,r) = {(n, k r ), (n, k r +1),..., (n, k r +D r -1)}
[0088] where D r is the number of SSPs in this FSSZ, expressed as the aggregation degree of SSPs in this area, k r is the serial number of the first SSP in the current FSSZ, and B n represents the number of FSSZs in the n-th frame. To increase the estimation accuracy of the FSSZ and prevent non-single-source frequency-time points from mixing in, it is required that the FSSZ contains at least two consecutive groups of adjacent frequency SSPs estimated in Step 2. As Figure 3 shown, the schematic diagram of the FSSZs constructed within 39 frames when measuring three sound sources S1, S2, and S3 by two arrays M1 and M2. The three colors respectively represent the FSSZs dominated by the three sound sources.
[0089] Step 4: According to the high correlation characteristics of the FSSZ time-frequency distribution of the same sound source at different microphone arrays within one frame, apply the FSSZ distribution of this sound source at the reference array to all arrays.
[0090] As Figure 3 shown, the time-frequency distributions of each sound source at the FSSZ of arrays M1 and M2 are highly correlated. Taking M1 as the reference array, the FSSZ distribution of M1 is applied to M2.
[0091] Step 5: Adopt the Circular Integral Cross-Spectrum (CICS) method to perform intermediate DOA estimation for each SSP, and combine the FSSZ aggregation degree in Step 3 for weighting to construct a sound source position histogram.
[0092] In a circular microphone array, for a given pair of adjacent microphones m i and m j , the following relationship is satisfied:
[0093] τ i→0 (φ) = τ 0,1 (φ) - τ i,j (φ), i = 0, 1,... M - 1, j = mod(i + 1, M)
[0094] where φ ∈ [0, 2π) is the candidate angle of the DOA. Calculate the phase rotation factor:
[0095]
[0096] Calculate the CICS:
[0097]
[0098] where
[0099]
[0100] Calculate the angle estimation of the current SSP(n, k):
[0101]
[0102] To reduce the computational complexity of the CICS, first use the narrowband far-field beamforming algorithm combined with the geometric characteristics of the circular array to perform a preliminary estimation of the DOA, and the estimated value is Then the search range of the CICS can be reduced to
[0103] For an M-element uniform circular microphone array, every two adjacent array elements can be regarded as a binary sub-array, as Figure 4 shown, satisfying:
[0104] τ i,j (n, k)·c = l·sin(θ i,j ), i = 0, 1,..., M - 1, j = mod(i + 1, M)
[0105] Among them, c is the speed of sound, which is 342 m / s, and l is the distance between two array elements, which is 0.0475 m in this example. A rough DOA estimate can be obtained for each subarray:
[0106]
[0107] Combining the rough DOA estimates of all M binary subarrays in an array, the DOA estimate value of the sound source can be obtained
[0108] i = 0, 1,..., M - 1, j = mod(i + 1, M)
[0109] Under the current SSP, the final CICS estimate of DOA is expressed as:
[0110]
[0111] In step 4, the FSSZ distribution of the reference array is applied to all arrays. In one frame, the DOA estimate value of the sound source of the SSP in all arrays' FSSZs is obtained using the CICS method, and the sound source position histogram of the nth frame is constructed accordingly:
[0112]
[0113] Among them, (u, v) represents the small block interval in 3D space. In the entire 3D space, the total number of small block intervals is U × V, and D r is the number of SSPs in the current FSSZ, and M b (u, v) refers to the value of the interval (u, v) in the 3D histogram:
[0114]
[0115] Among them refers to the position estimate of the sound source by the current SSP. Considering all frames, the number D of SSPs in the current FSSZ is r used for weighting to construct the frame-level sound source position histogram:
[0116]
[0117] Figure 5 is the schematic diagram of the sound source position histogram after an 8 * 8 averaging window. The three sound sources correspond to the three peaks in the sound source position histogram, and their positions are the projection positions of the three peaks on the plane.
[0118] Step 6: Use 2D-CFAR to perform peak detection on the sound source position histogram in step 5 to obtain the 2D positions of multiple sound sources.
[0119] The sliding window size is G×G, and the protection cell size is Q×Q. In this example, G is taken as 8 and Q is taken as 3. The reference cell is represented as X = (x1, x2, … x U ), U = G 2 -Q 2 . The threshold factor P fa is the false alarm probability, is the average clutter intensity. η CUT is the signal intensity of the cell to be detected, and the decision strategy is where H1 indicates that there is a sound source in the cell to be detected, and H0 indicates that there is no sound source in the cell to be detected.
[0120] 2D-CFAR is used to detect the peaks of the sound source position histogram, and the detected sound source positions are S1: (0.3 m, 1.5 m), S2: (1.7 m, 0.4 m) and S3: (0.5 m, -1.6 m).
[0121] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for indoor multi-sound source localization based on DOA estimation and DOA association, characterized in that: The following steps are involved: Step 1: Perform voice activity detection and identify voiced frames based on energy threshold and short-time zero-crossing rate detection method; Step 2: In the voiced frame identified in step 1, taking one frame of data as a unit, determine the single source time-frequency point by the correlation of the phase differences between different microphone array elements in the array; Step 3: In each frame, construct an elastic single source region with variable length based on the SSP identified in step 2; Step 4: Based on the high correlation of the FSSZ time-frequency distribution of the same sound source at different microphone arrays within a frame, the FSSZ distribution of the sound source at the reference array is applied to all arrays; Step 5: Use the circular integral cross-spectrum method to estimate the intermediate DOA of each SSP, and weight it with the FSSZ concentration in step 3 to construct a sound source position histogram; Step 6: Use 2D-CFAR to perform peak detection on the sound source position histogram in step 5 to obtain the 2D positions of multiple sound sources.
2. The indoor multiple sound source localization method based on DOA estimation and DOA association according to claim 1, characterized in that: Use the following method to determine that three adjacent time-frequency points (n, k-1), (n, k) and (n, k+1) are single-source time-frequency points, where n represents the time domain index and k represents the frequency domain index. n k-1,k and e n k,k+1 When the following conditions are met, three adjacent time-frequency points are judged as single-source time-frequency points: Where δ is a small positive number, e n k-1,k and e n k,k-1 is the delay vector correlation coefficient: Where corrcoef(x1,x2) is used to calculate the correlation coefficient of x1 and x2, τ(n,k) is the delay vector of the time-frequency point (n,k), F represents the number of FFT points in STFT, and N represents the total number of time frames of the audio to be processed.
3. The indoor multiple sound source localization method based on DOA estimation and DOA association according to claim 1, characterized in that: In the nth frame, the rth FSSZ is constructed using the following method: Φ (n,r) {(n,k r ),(n,k r +1),…,(n,k r +D r -1)} Where D r is the number of SSPs in the FSSZ, expressed as the aggregation degree of SSPs in the region, k r is the serial number of the first SSP in the current FSSZ, 1≤r≤B n , B n Expressed as the number of FSSZs in the nth frame.
4. The indoor multiple sound source localization method based on DOA estimation and DOA association according to claim 1, characterized in that: When estimating the intermediate DOA, we first use the narrowband far-field beamforming algorithm combined with the geometric characteristics of the circular array to make a preliminary estimate of the DOA: i=0,1,...,M-1,j=mod(i+1,M) in, represents the preliminary DOA estimation before CICS method, angle(X) represents the phase angle of X, represents the rough DOA estimation of the binary subarray composed of the i-th array element and its adjacent array element j=mod(i+1,M), X i (n,k) represents the time-frequency domain representation of the audio signal received by the i-th array element in the array, "*" represents conjugation, M is the number of array elements in an array, c is the speed of sound, l is the distance between array elements, ω(n,k) represents the frequency corresponding to the time-frequency point (n,k), and then the CICS method is used to obtain the intermediate DOA estimate: in, in represents the narrowband DOA estimation at the time-frequency point (n, k), Δθ represents the The search scope is centered on CICS (n ,k) (φ) represents the circular integrated cross spectrum value in direction φ, G i,j (n,k) represents the phase change value at the time-frequency point (n,k), represents the phase rotation factor in direction φ, τ i→0 (φ) represents the time delay difference between the ith array element and the 0th array element in the direction φ, and φ∈[0,2π) is the candidate angle of DOA.
5. The indoor multiple sound source localization method based on DOA estimation and DOA association according to claim 1, characterized in that: The frame-level sound source position histogram is constructed according to the following rules: where y n (u,v) represents the value of the frame-level sound source histogram at the coordinate (u,v) in the nth frame. U and V represent the number of segmentations on the x-axis and y-axis respectively. y (n,r) (u, v) represents the contribution value of the rth FSSZ in the nth frame to the current frame-level histogram, (u, v) represents a small block interval in the 3D space, D r is the number of SSPs in the current FSSZ, M b (u,v) refers to the value of the interval (u,v) in the 3D histogram: in Refers to the current SSP's estimate of the sound source's position.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 5 are implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 5 are implemented.
Citation Information
Cited By
Method and device for simultaneously counting and positioning multiple sound sources
CN120669197A
A method and apparatus for simultaneous counting and localization of multiple sound sources
CN120669197B