A robust sound source localization method fusing inter-frame correlation of sound intensity amplitude
By using an equilateral triangle array and a time-frequency point screening method based on inter-frame correlation of sound intensity amplitude in sound source localization, the problems of high computational complexity and low precision in sound source localization are solved, and higher positioning accuracy and anti-interference capability are achieved.
Patent Information
- Application Number
- CN202511045318.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-07-29
AI Technical Summary
现有声源定位方法在计算量大、声强估计不准确和抗混响、抗噪声能力不足的问题,尤其是在室内声源定位场景中效果不佳。
An equilateral triangle array is used to estimate the sound intensity, and the time-frequency point screening method based on the inter-frame correlation of the sound intensity amplitude is used. The comprehensive weight is calculated using the sound intensity estimation consistency method for weighted selection, and reliable time-frequency points are screened out. Finally, the sound source azimuth is estimated by the sound intensity localization method.
The computational complexity is reduced, and the accuracy and robustness of sound source localization are improved, especially in indoor environments, where the system has a strong anti-interference ability against reverberation and noise.
Smart Images

Figure CN120539672B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a robust sound source positioning method fusing sound intensity amplitude inter-frame correlation, and belongs to the technical field of indoor sound source positioning. BACKGROUND
[0002] Sound intensity refers to the average sound energy flow per unit area perpendicular to the sound propagation direction, which is a vector closely related to the direction of arrival (DOA) of the sound source. By estimating the sound intensity vector through an algorithm, the azimuth angle of the sound source can be estimated. The equilateral triangle array is a non-orthogonal first-order difference array. If the sound intensity is estimated by using the difference approximation method in the orthogonal first-order difference array of the reference four microphones. First, the sound pressure at the array midpoint is estimated based on the average sound pressure of the three microphones, and then the velocity components estimated by the three pairs of dipoles sharing the microphone are projected onto the coordinate axes and superimposed. This method has the problems of large calculation amount and inaccurate sound intensity estimation, which leads to poor sound source positioning effect.
[0003] After estimating the sound intensity, not all time-frequency points corresponding to the sound intensity are useful for sound source positioning. In fact, only a small part of the time-frequency points are reliable time-frequency points useful for sound source positioning. Therefore, screening out reliable time-frequency points useful for sound source positioning is the key to improving the algorithm's ability to resist reverberation and noise. The pseudo sound intensity estimation consistency (EC) algorithm uses the high consistency between pseudo sound intensities (PIV) in the time-frequency region dominated by the direct path of a single sound source to select time-frequency points. It has fast operation speed, good reverberation resistance, and strong algorithm versatility, and can also be applied to sound intensity estimation algorithms. However, it also has some shortcomings: when calculating the time frame weight, the normalized sound intensity is used, which causes sound intensity vectors with different amplitudes to contribute the same to the time frame weight, ignoring the amplitude information of the single time-frequency point sound intensity. And when calculating the weight of a single frequency point, the EC algorithm only considers the current frame information, without considering the information before and after the frame, which can improve the detection accuracy of the direct sound. SUMMARY
[0004] The technical problem to be solved by the present application is to provide a robust sound source positioning method fusing sound intensity amplitude inter-frame correlation, which effectively screens out time-frequency points less affected by reverberation and noise based on the estimated sound intensity, thereby improving the anti-reverberation and anti-noise ability of the method in the indoor sound source positioning scene.
[0005] The present application solves the above technical problems by adopting the following technical solutions:
[0006] A robust sound source positioning method fusing sound intensity amplitude inter-frame correlation, comprising the following steps:
[0007] Step 1, a sound source signal is received by using an equilateral triangle array composed of three omnidirectional microphones, time-domain sound pressure signals received by the microphones are converted into a time-frequency domain by a short-time Fourier transform, and sound intensity components of each time-frequency point at the center of the equilateral triangle array are calculated by using a sound pressure Taylor series expansion sound intensity estimation method in the time-frequency domain;
[0008] Step 2, sound intensity amplitudes are calculated according to the sound intensity components calculated in step 1, comprehensive weights corresponding to each time-frequency point are calculated by using a sound intensity estimation consistency method, the sound intensity amplitudes are weighted by using the comprehensive weights, and weighted sound intensity amplitudes are obtained;
[0009] Step 3, time-frequency point screening factors based on sound intensity amplitude interframe correlation are constructed according to the weighted sound intensity amplitudes and unweighted sound intensity amplitudes, and reliable time-frequency points are screened from all time-frequency points according to the time-frequency point screening factors;
[0010] Step 4, sound intensity components corresponding to the reliable time-frequency points screened in step 3 are superimposed, and a sound source azimuth angle estimation value is obtained by using a sound intensity positioning method.
[0011] Compared with the prior art, the above technical scheme has the following technical effects:
[0012] 1, the sound intensity is estimated by using the equilateral triangle array, and then the sound source is positioned, so that the calculation complexity is lower than that of a four-microphone uniform circular array, and the azimuth angle omnidirectional spatial positioning capability is equivalent to that of the four-microphone uniform circular array.
[0013] 2, the Taylor series expansion sound intensity estimation method in the three-dimensional regular tetrahedral array is applied to a two-dimensional planar array, and the method is more convenient to calculate than sound pressure difference orthogonal decomposition and cross power spectrum, and has equivalent positioning performance.
[0014] 3, on the basis of the sound intensity estimation consistency detection method, the sound intensity amplitude information of a single time-frequency point is fully utilized, the comprehensive weights are calculated by using the consistency detection method, the sound intensity amplitudes are weighted, and the interframe correlation of the two is utilized. The time-frequency point screening method fusing the interframe correlation of the sound intensity amplitude can more accurately detect time-frequency points less disturbed by reverberation and noise, so that the positioning accuracy of the method can be improved.
[0015] 4, in the indoor sound source positioning of a simulation scene, the method has higher positioning accuracy, is more robust to noise and reverberation interference, and has certain practical value. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is an equilateral triangle array of three microphones designed by the application;
[0017] Figure 2It is a schematic diagram of an indoor sound source positioning simulation scene.
[0018] Fig. 3 (a) and Fig. 3 (b) are root mean square error comparisons of the present application method and SI, SI-EC, DPD-MUSIC algorithm under different reverberation time conditions when the signal-to-noise ratio is fixed at 15 dB and 5 dB respectively;
[0019] Fig. 4 (a) and Fig. 4 (b) are root mean square error comparisons of the present application method and SI, SI-EC, DPD-MUSIC algorithm under different signal-to-noise ratio conditions when the reverberation time is fixed at 300 ms and 800 ms respectively;
[0020] Figure 5 Fig. 5 is a root mean square error comparison of the present application method and SI, SI-EC, DPD-MUSIC algorithm under different array radius conditions when the reverberation time is 0.5 s and the signal-to-noise ratio is 10 dB.
[0021] Figure 6 It is a flowchart of a robust sound source positioning method of the present application fusing sound intensity amplitude inter-frame correlation. DETAILED DESCRIPTION
[0022] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be explained as a limitation of the present application.
[0023] The present application proposes a robust sound source positioning method fusing sound intensity amplitude inter-frame correlation, Figure 6 It is a flowchart of the method of the present application, and the specific implementation steps are as follows:
[0024] Step 1, considering three omnidirectional microphones , , forming an equilateral triangle array as shown in with a radius and an equi-azimuthal interval Figure 1 . Then the time-domain signals received by the microphones are subjected to short-time Fourier transform, and the sound intensity components and are estimated in the time-frequency domain by the sound pressure Taylor series expansion method.
[0025] 1.1, for an indoor single sound source positioning scene, the incident direction of the far-field plane wave sound source is defined as the azimuth angle with the positive half-axis of the coordinate system , and the value range is , and the time-domain sound pressure signals received by the three microphones can be represented as:
[0026] (1)
[0027] where, , is the time-domain clean speech signal, and the symbol represents the convolution operation, represents the room impulse response from the sound source to the th microphone, is the additive noise signal of the th microphone, and is assumed to be mutually independent with all microphones and uncorrelated with the sound source signal.
[0028] 1.2, since the speech signal has the short-time stationary characteristic, the time-domain observation signal is generally subjected to the short-time Fourier transform (STFT), and the above formula is expressed as:
[0029] (2)
[0030] where, and represent the th time unit and the th frequency unit, respectively, , , and are the signals after STFT. , , and
[0031] 1.3, in the time-frequency domain, the sound intensity estimation algorithm of the sound pressure Taylor series expansion in the three-dimensional cubic tetrahedral array is extended to the two-dimensional equilateral triangular array, the input sound pressure and the absolute coordinates of the microphone are subjected to matrix transformation, and the sound pressure and the vibration velocity are calculated. In order to facilitate derivation, the frequency is omitted, and the specific derivation process is as follows:
[0032] In an ideal fluid medium, the relationship between the particle vibration velocity and the sound pressure gradient is given by the linear Euler equation:
[0033] (3)
[0034] where, is the air density, and are the vibration velocity and the sound pressure of the particle at time, respectively, represents the gradient operation.
[0035] Then the particle vibration velocity is expressed by the sound pressure gradient as:
[0036] (4)
[0037] where is the angular frequency, finite-difference approximation is usually used to approximate one component of the pressure gradient, then the particle velocity is calculated using equation (4).
[0038] The sound pressure at the midpoint of two microphones very close to each other can be obtained using the zero-order term and the first-order term of the Taylor series expansion, and the and , then the sound pressure of the th microphone located at a distance of from the origin of the coordinate system is expressed as:
[0039] (5)
[0040] where is the sound pressure at the origin, , is the sound pressure gradient of the th microphone relative to the origin.
[0041] Taking a regular triangular array as an example, the position coordinates of the three microphones are , , Substituted into equation (5), written in the form of the following 3 linear equations:
[0042] (6)
[0043] where , , for ease of subsequent derivation, the above equation is written in matrix form.
[0044] Let , , Figure 1 the relative position coordinate matrix of the microphones in the middle is expressed as:
[0045] (7)
[0046] The inverse of the matrix is multiplied by , and the vector related to the sound pressure and velocity is obtained:
[0047] (8)
[0048] where , , is omitted , corresponding to formula (2) ; by vector The sound pressure at the center point can be approximated as: , The directional sound pressure gradient is , The directional sound pressure gradient is .
[0049] 1.4, according to the corresponding 、 、 , estimate the array center point Direction and The sound intensity components in the direction are:
[0050] (9)
[0051] (10)
[0052] in, is the imaginary unit, This is the real part operation.
[0053] Step 2: Calculate the sound intensity amplitude based on the sound intensity component estimated in step 1 , and then use the sound intensity estimation consistency method to calculate the comprehensive weight , use the comprehensive weight to weight the sound intensity amplitude, and get the weighted sound intensity amplitude .
[0054] The traditional DOA estimation algorithm based on sound intensity is to calculate the and The components are superimposed separately to achieve the effect of anti-noise and anti-reverberation. Therefore, according to equations (9) and (10), the closed-form solution based on the azimuth angle estimation value of the regular triangle array is:
[0055] (11)
[0056] As shown in formula (11), most positioning methods based on sound intensity estimation obtain a positioning result by superimposing the sound intensity components of all time-frequency points. This is a positioning method based on sound intensity estimation without time-frequency point screening. The following is a positioning method based on time-frequency point screening based on sound intensity inter-frame correlation derived and proposed by the present invention.
[0057] When a time-frequency point is dominated by a single sound source direct path, its sound intensity vector can be expressed as:
[0058] (12)
[0059] in, for the clean speech signal after STFT transform, for the direct sound azimuth, subscript denotes the direct path, is the transpose. direction and the direct sound direction and , .
[0060] the direct sound dominant frame, the sound intensity amplitude is:
[0061] (13)
[0062] where, is the vector two-norm, when a frame is direct sound dominant, the sound intensity of most time-frequency points in this frame is concentrated in the direction of the sound source, which is called a clustering frame.
[0063] 2.1, according to the sound intensity of a single time-frequency point and the correlation degree of the sound intensity of all time-frequency points in a frame, the normalized weight of each frame is calculated , specifically:
[0064] (14)
[0065] where, is the number of time-frequency points in a frame, denotes the sound intensity vector of a single time-frequency point.
[0066] When a frame is a clustering frame, assuming that all time-frequency points in this frame are direct sound dominant, substituting equations (12) and (13) into equation (14) can be known that when the frame is direct sound dominant, , indicating that most sound intensity vectors in this frame are concentrated in a specific direction. When , it indicates that due to the existence of reverberation and noise, the sound intensity vectors in this frame are random.
[0067] 2.2, in fact, in each frame of signal, only part of the time-frequency points may be direct sound dominant, the weight of each time-frequency point can be calculated by the spatial distance between the sound intensity vector of each time-frequency point and the average sound intensity vector in a frame , specifically as follows:
[0068] (15)
[0069] where, denotes the average sound intensity of the frame signal, according to equation (15), if the sound intensity vector of a single time-frequency point is close to the average sound intensity vector in a frame When equal, it indicates that the spatial distance of the two is small, and according to the property of the inverse cosine function, at this time .
[0070] 2.3, multiply the current frame weight with the weight of each time-frequency point in a frame to obtain the comprehensive weight of each time-frequency point:
[0071] (16)
[0072] When a certain frame is an aggregation frame, the time frame weight , and when a certain time-frequency point in this frame is dominated by direct sound, the weight of a single time-frequency point , then according to formula (16), the comprehensive weight at this time is also close to 1.
[0073] The above sound intensity estimation consistency method can calculate the comprehensive weight of each time-frequency point.
[0074] 2.4, in the actual positioning scene, according to the co-point sound intensity estimation method, the sound intensity vector of each time-frequency point is calculated, and then the sound intensity amplitude is calculated:
[0075] (17)
[0076] 2.5, according to the use of the estimation consistency algorithm to calculate the comprehensive weight corresponding to each time-frequency point, and then the sound intensity amplitude is weighted to obtain the weighted sound intensity amplitude , which is:
[0077] (18)
[0078] According to (16) and (18), when a certain time-frequency point is dominated by direct sound, at this time, the two have strong correlation, and the weighted sound intensity amplitude is almost equal to the unweighted sound intensity amplitude, when reverberation and noise exist, the correlation between the two is weakened, and therefore the correlation between the two is considered to be used for reliable time-frequency point screening.
[0079] Step 3, according to the sound intensity amplitude inter-frame correlation, a time-frequency point screening factor is constructed, the weighted sound intensity amplitude and the unweighted sound intensity amplitude are input to the screening factor, and after comparing with the threshold , the reliable time-frequency points are screened out.
[0080] 3.1, the direct sound arrives earlier than reverberation in the speech signal, so the inter-frame correlation of the weighted sound intensity amplitude and the unweighted sound intensity amplitude of the continuous frame can be considered. The single-source area detection method of the present application considers the correlation degree of several continuous time-frequency points in a frame of signal, and the screening factor construction method of the correlation in the single-source area detection, and the continuous several time-frequency points in a frame are rewritten as the continuous several frames of the same frequency index, so that the time-frequency point screening factor is expressed as:
[0081] (19)
[0082] wherein, the screening factor is obtained by considering the trade-off between the calculation amount and the screening accuracy, the set of the continuous 3 frames, represents the inter-frame summation operation, represents the correlation degree of the sound intensity amplitude and the weighted sound intensity amplitude in a time-frequency analysis area.
[0083] 3.2, in each simulation, the values of the screening factors calculated by all time-frequency points are sorted in descending order, and the Mth largest value is taken as the dynamic threshold , then the screening factor is compared with the dynamic threshold one by one, and if it is greater than the threshold, the time-frequency point is regarded as a reliable time-frequency point, and the sound intensity components corresponding to the reliable time-frequency points are marked as and .
[0084] Step 4, extracting the sound intensity components corresponding to the reliable time-frequency points screened in step 3 and , then superimposing them, and then estimating the closed-form solution of the sound source azimuth angle estimation value by the sound intensity positioning method, which is:
[0085] (20)
[0086] The present application will be further described in combination with some specific examples.
[0087] The present application has good reverberation and noise resistance because the reliable time-frequency points are screened according to the inter-frame correlation of the sound intensity amplitude, and the positioning performance of the present application in the indoor reverberation and noise environment is verified by simulation experiments. The scene and process of the specific implementation simulation experiment are as follows.
[0088] The simulation scene is shown in Figure 2 , the size of the simulated rectangular room is set to , the array is located at the geometric center point of the room, and the array is a regular triangle array of 3 microphones with a radius of 2.6 cm, as shown in Figure 1 .
[0089] In the simulation experiment, the loudspeaker is placed at the same height as the microphone array, with a distance of 1.5 m from the center of the array, and the signal sampling rate is set to 16 kHz. The simulated room reverberation is generated by the mirror method. The microphone noise is zero-mean additive Gaussian white noise and is not correlated with the speech signal. The sound speed c is 340 m / s. The frame length is 512 and the frame shift is 256 in the short-time Fourier transform. As shown in the middle gray circle, the azimuth angle of the sound source is placed on the interval of 10 degrees, so there are a total of 36 sound source placement test positions. In each sound source placement position, 5 male and 5 female voices are randomly selected from the TIMIT database as the test database. After removing the silent segment, 2s of each voice is extracted as the experimental sound source signal. The frequency band range of the signal is set to 300-5000Hz. Each selected voice is subjected to 10 Monte Carlo experiments. Therefore, each sound source position is subjected to a total of 100 Monte Carlo experiments, and the total number of experiments for the 36 candidate positions is 3600. In the DPD-MUSIC algorithm, the time smoothing factor and the frequency smoothing factor are set to 3 and 4 respectively. In the present method and the SI-EC algorithm, the frequency points with the top 35% values are extracted. Figure 2
[0090] To evaluate the positioning performance of each algorithm, the root mean square error is used as an evaluation index, which is defined as:
[0091] (21)
[0092] In formula (21), represents the sound source placement position index, represents the azimuth angle estimated by the th simulation experiment algorithm at the th sound source position, represents the real azimuth angle of the position sound source. To fully demonstrate the positioning performance of various algorithms in the interval of 0 to 360 degrees of the two-dimensional plane azimuth angle for different sound source positions, the sound source position starts from 0 degrees with an increment of 10 degrees, and there are a total of 36 candidate sound source positions. Therefore, the value of is 36. represents the number of positioning times in the simulation at the current sound source position. In a sound source position, 10 voices are selected for experiment, and each is subjected to 10 Monte Carlo times. Therefore, the value of
[0093] To illustrate the advantages of the method of the present invention over other indoor sound source localization methods in terms of anti-reverberation and anti-noise, the following positioning performance comparison is performed between the method of the present invention and the sound intensity estimation localization algorithm (SI) without time-frequency point screening, the consistency detection algorithm based on SI (SI-EC), and the direct path detection in the modal domain (DPD-MUSIC) algorithm.
[0094] Figure 3(a) and Figure 3(b) show the reverberation time when the signal-to-noise ratio is fixed at 15dB and 5dB respectively. The root mean square error (RMSE) changes of each sound source localization algorithm at .
[0095] Analysis shows that the SI algorithm has the worst positioning performance, because it resists the influence of reverberation and noise by superimposing all time-frequency points. Many time-frequency points dominated by reverberation will affect its positioning performance. The DPD-MUSIC algorithm in the modal domain has modal aliasing when converting the speech signal from the time-frequency domain to the modal domain. The fewer the number of microphones, the greater the impact. Therefore, in the case of a three-microphone array, the modal domain algorithm has poor positioning performance. Since the SI-EC algorithm only uses the current frame information, it is easy to miss the starting end when detecting the starting end of the speech, which makes its anti-reverberation ability limited. The algorithm of the present invention makes full use of the inter-frame information of the sound intensity amplitude, and detects the starting end of the speech more accurately than the SI-EC algorithm, and retains the speech-dominated segment more accurately than the EC algorithm. Therefore, the accurate detection of the initial direct sound makes the algorithm of the present invention have a strong anti-reverberation ability. As shown in Figure 3(b), under strong noise conditions with a signal-to-noise ratio of 5 dB, the proposed algorithm reduces the RMSE by approximately 3 degrees compared to the SI algorithm at a reverberation time of 0.7 seconds, and improves positioning accuracy by approximately 1.3 degrees compared to the SI-EC algorithm. Overall, the proposed algorithm demonstrates very robust anti-reverberation capabilities.
[0096] Figure 4(a) and Figure 4(b) show that when the reverberation time is fixed at 300ms and 800ms respectively, the signal-to-noise ratio is set to The RMSE changes of each sound source localization algorithm at different times.
[0097] As can be seen from FIG. 4(a), when the reverberation time is 0.3 s and the signal-to-noise ratio changes from 0 dB to 25 dB, the root mean square errors of the four algorithms are all between 2 degrees and 3 degrees, which indicates that the algorithms are all relatively robust to noise under low reverberation conditions. As can be seen from FIG. 4(b), the SI algorithm still has the worst positioning performance, because it does not perform any time-frequency point screening and only estimates an angle by superimposing all frequency point sound intensity components. The DPD-MUSIC algorithm still has poor noise resistance under this array due to modal aliasing. The SI-EC has good noise resistance because it is a pseudo sound intensity screening algorithm. The method fully utilizes the sound intensity amplitude inter-frame correlation to perform time-frequency point screening, so that the positioning performance of the algorithm has the best noise resistance under two reverberation intensities. As can be seen from FIG. 4(a), when the low reverberation intensity is 300 ms, the method has a positioning precision improvement of nearly 1 degree compared with the EC algorithm. As can be seen from FIG. 4(b), the greater the reverberation intensity and the lower the signal-to-noise ratio, the more obvious the positioning precision improvement of the method. Compared with the basic sound intensity estimation algorithm, the average RMSE is reduced by about 4 degrees under strong reverberation conditions.
[0098] Figure 5 FIG. 4(b) shows the positioning performance of the four algorithms under the condition that the reverberation time is 0.5 s and the signal-to-noise ratio is 10 dB.
[0099] As can be seen from FIG. 4(a), when the reverberation time is 0.3 s and the signal-to-noise ratio changes from 0 dB to 25 dB, the root mean square errors of the four algorithms are all between 2 degrees and 3 degrees, which indicates that the algorithms are all relatively robust to noise under low reverberation conditions. As can be seen from FIG. 4(b), the SI algorithm still has the worst positioning performance, because it does not perform any time-frequency point screening and only estimates an angle by superimposing all frequency point sound intensity components. The DPD-MUSIC algorithm still has poor noise resistance under this array due to modal aliasing. The SI-EC has good noise resistance because it is a pseudo sound intensity screening algorithm. The method fully utilizes the sound intensity amplitude inter-frame correlation to perform time-frequency point screening, so that the positioning performance of the algorithm has the best noise resistance under two reverberation intensities. As can be seen from FIG. 4(a), when the low reverberation intensity is 300 ms, the method has a positioning precision improvement of nearly 1 degree compared with the EC algorithm. As can be seen from FIG. 4(b), the greater the reverberation intensity and the lower the signal-to-noise ratio, the more obvious the positioning precision improvement of the method. Compared with the basic sound intensity estimation algorithm, the average RMSE is reduced by about 4 degrees under strong reverberation conditions. Figure 5 It can be seen that the radius of 1 cm is a critical value affecting the positioning performance of the sound intensity method, so when the radius increases from 0.5 cm to 5 cm, the RMSE of each algorithm will have a downward trend first and then an upward trend. When the radius exceeds 1 cm, the non-orthogonal first-order difference sound intensity estimation method will have an overall upward trend in the root mean square error with the increase of the array radius. The reason is that with the increase of the array radius, the error of the difference approximation will increase, which will lead to inaccurate sound intensity estimation, and finally lead to a slight decline in the positioning performance. Therefore, the RMSE of the SI method, the SI-EC method and the method has a certain upward trend. Because the method can accurately detect the starting end of the speech through time-frequency point screening, it has good anti-reverberation and anti-noise ability, so it also shows a stable performance improvement trend under the condition of increasing array size. Compared with the basic EC algorithm, the root mean square error always has a downward trend of 2 degrees. It indicates that the method has good robustness to the increase of the array size in the indoor sound source positioning scene.
[0100] Based on the same inventive concept, the embodiment of the present application provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the fusion sound intensity amplitude inter-frame correlation robust sound source positioning method when executing the computer program.
[0101] Based on the same inventive concept, the embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the robust sound source positioning method based on fusion of inter-frame correlation of sound intensity amplitude.
[0102] Those skilled in the art will understand that embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) containing computer usable program code.
[0103] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 means for performing the functions specified in the flowchart
[0104] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including instruction means, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 one or more blocks or means for performing the functions specified in the flowchart
[0105] These computer program instructions can also be loaded into a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable data processing apparatus to produce a computer implemented process, so that the instructions executed on the computer or other programmable data processing apparatus provide a process for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 one or more blocks or means for performing the functions specified in the flowchart
[0106] The above examples only illustrate the technical idea of the present application, and cannot be used to limit the protection scope of the present application. Any modification made according to the technical idea of the present application on the basis of the technical scheme falls within the protection scope of the present application.
Claims
1. A robust sound source localization method fusing inter-frame correlation of sound intensity amplitude, characterized in that, The method comprises the following steps: Step 1, receiving a sound source signal using an equilateral triangle array composed of three omnidirectional microphones, and converting the time-domain sound pressure signals received by the three omnidirectional microphones into a time-frequency domain through a short-time Fourier transform, and calculating the sound intensity components of each time-frequency point at the center of the equilateral triangle array using a sound pressure Taylor series expansion sound intensity estimation method in the time-frequency domain; the specific process is as follows: Step 1.1, using three omnidirectional microphones, a pure speech signal emitted by a sound source is sampled to obtain the time-domain sound pressure signal p m (t) is: p m (t) = s(t) * h m (t) + n m (t) wherein m∈{1,2,3}, s(t) is a time-domain clean speech signal, the symbol * represents convolution operation, h m (t) is a room impulse response from the sound source to the mth microphone, n m (t) is an additive noise signal of the mth microphone; Step 1.2, for p m (t) performs short-time Fourier transform to obtain the mth microphone at each time-frequency point (t k ,f b ) of the time-frequency domain sound pressure signal P m (t k ,f b ): P m (t k ,f b ) = S(t k ,f b )H m (t k ,f b ) + N m (t k ,f b ) where t k represents the kth time unit, i.e., the t k th frame, f b represents the bth frequency unit, S(t k ,f b ), H m (t k ,f b ) and N m (t k ,f b ) are the signals after the short-time Fourier transform of s(t), h m (t) and n m (t), respectively. Step 1.3, in the time-frequency domain, the sound pressure Taylor series expansion sound intensity estimation algorithm in the three-dimensional regular tetrahedral array is extended to the two-dimensional equilateral triangle array, and an orthogonal coordinate system is established with the center of the equilateral triangle array as the origin, and one microphone is located on the positive y-axis of the coordinate system, and the sound pressure and velocity related vector G is calculated as: The sound pressure at the center of the equilateral triangle array at each time-frequency point (t k ,f b ) is P o (t k ,f b )≈(P1(t k ,f b )+P2(t k ,f b )+P3(t k ,f b )) / 3, the x-direction sound pressure gradient is D y (t k ,f b )≈(2P2(t k ,f b )-P1(t k ,f b )-P3(t k ,f b )) / 3r, P o is the sound pressure at the origin, and (x1, y1), (x2, y2), and (x3, y3) are the position coordinates of the three omnidirectional microphones, P represents sound pressure; Step 1.4, according to P o (t k ,f b ), D x (t k ,f b ) and D y (t k ,f b ), the sound intensity components of the center of the equilateral triangle array in the x direction and the y direction are estimated, respectively: where I x (t k ,f b ), I y (t k ,f b ) are the sound intensity components at the center of the equilateral triangle array in the x and y directions, respectively, Re{·} is the real part operation, j is the imaginary unit, and p0is the air density. Step 2, calculating the sound intensity amplitude according to the sound intensity components calculated in step 1, calculating the comprehensive weight corresponding to each time-frequency point using the sound intensity estimation consistency method, and weighting the sound intensity amplitude using the comprehensive weight to obtain the weighted sound intensity amplitude; Step 3, constructing a time-frequency point screening factor based on the inter-frame correlation of the sound intensity amplitude according to the weighted sound intensity amplitude and the unweighted sound intensity amplitude, and screening reliable time-frequency points from all time-frequency points according to the time-frequency point screening factor; Step 4, superimposing the sound intensity components corresponding to the reliable time-frequency points screened in step 3 to obtain the sound source azimuth estimation value through the sound intensity positioning method.
2. The robust sound source localization method fusing inter-frame correlation of sound intensity amplitude according to claim 1, characterized in that, The specific process of step 2 is as follows: Step 2.1, calculating the normalized weight of each frame according to the correlation degree of the sound intensity vector of a single time-frequency point and the sound intensity vectors of all time-frequency points in a frame: where l(t k ) is the normalized weight of the t k th frame, F is the number of time-frequency points in a frame, I(t k ,f b ) = [I x (t k ,f b )I y (t k ,f b )] T represents the single time-frequency point sound intensity vector, and ||·|| is the vector two-norm. Step 2.
2. Calculate the weight value μ(t,f) of each time-frequency point according to the spatial distance between the sound intensity vector of each time-frequency point and the average sound intensity vector of a frame k ,f b ) : wherein represents the t k the sound intensity mean value of the frame signal; Step 2.3, t k he normalization weight l(t k ) of the frame and the weight μ(t k ,f b ) of each time-frequency point in this frame are multiplied to obtain the comprehensive weight λ(t k ,f b ) corresponding to each time-frequency point: λ(t k ,f b ) = l(t k )μ(t k ,f b ) Step 2.
4. Estimate the sound intensity components I from the estimates of step 1.4 x (t k ,f b ) and I y (t k ,f b ) to compute the sound intensity amplitude I(t k ,f b ): Step 2.5, the sound intensity amplitude I(t k ,f b ) is weighted by the comprehensive weight λ(t k ,f b ) to obtain the weighted sound intensity amplitude I w (t k ,f b ): I w (t k ,f b )=λ(t k ,f b )I(t k ,f b )。 3. The robust sound source localization method fusing inter-frame correlation of sound intensity amplitude according to claim 2, characterized in that, The specific process of step 3 is as follows: Step 3.
1. Construct the time-frequency point screening factor C(t,f) based on the inter-frame correlation of the sound intensity amplitude, using the sound intensity amplitude I(t,f) calculated in step 2.4 and the weighted sound intensity amplitude I(t,f) obtained in step 2.
5. k b w k b k b where H is the t k-1 , t k and t k+1 consecutive frames; Step 3.2, sort the time-frequency point screening factors corresponding to all time-frequency points in descending order, select the Mth largest time-frequency point screening factor as the threshold value η, M is a preset value, compare the time-frequency point screening factor corresponding to each time-frequency point with the threshold value η, the time-frequency point whose time-frequency point screening factor is greater than η is a reliable time-frequency point, mark the sound intensity component corresponding to the reliable time-frequency point as and 4. The robust sound source localization method fusing inter-frame correlation of sound intensity amplitude according to claim 3, characterized in that, The specific process of step 4 is as follows: The sound intensity components corresponding to reliable time-frequency points are superimposed, and a closed-form solution of the sound source azimuth estimation value is obtained by a sound intensity positioning method 5. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, The processor executes the computer program to realize the steps of the robust sound source positioning method fusing the inter-frame correlation of the sound intensity amplitude according to any one of claims 1 to 4.
6. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 5. The computer program is executed by the processor to realize the steps of the robust sound source positioning method fusing the inter-frame correlation of the sound intensity amplitude according to any one of claims 1 to 4.
Citation Information
Patent Citations
Sound source direction identification method based on mini microphone array
CN106886010A
Time-frequency-space domain joint weighted circular harmonic domain pseudo-sound intensity sound source positioning method
CN108549052A