Sound source localization methods, devices and electronic equipment

By filtering mid-to-high frequency bands using a microphone array and calculating their weights, the problem of insufficient accuracy in sound source localization in existing technologies is solved, achieving more efficient and accurate sound source angle estimation.

CN115932728BActive Publication Date: 2026-04-03HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing microphone array-based sound source localization technologies lack accuracy and robustness in complex environments, affecting the estimation of sound source angles in voice interaction scenarios.

Method used

Sound signals are collected by a microphone array, the energy growth rate of each frequency point is determined, mid-frequency and high-frequency candidate frequency points are selected, the relative time delay of the mid-frequency points is calculated and the reference time delay is determined, the weight is determined based on the time delay difference of the high-frequency points, and the angle is estimated using a sound source localization algorithm.

Benefits of technology

It improves the accuracy and efficiency of sound source angle estimation, reduces the amount of calculation for low-frequency points, reduces the impact of noise, and enhances the accuracy of sound source localization in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115932728B_ABST
    Figure CN115932728B_ABST
Patent Text Reader

Abstract

This disclosure relates to a sound source localization method, apparatus, and electronic device. The method includes: acquiring sound signals through a microphone array; determining the energy growth amplitude of each frequency point of the sound signal; determining multiple candidate frequency points from each frequency point based on the energy growth amplitude of each frequency point; determining multiple first time delays for the transmission of each mid-frequency point to any two adjacent microphones in the microphone array; determining multiple time delay intervals based on each first time delay; determining a reference time delay from the multiple first time delays; determining a second time delay between any two adjacent microphones in the microphone array corresponding to each high-frequency point, obtaining multiple second time delays; determining the frequency weight of each high-frequency point based on the multiple second time delays and the reference time delay; determining an objective function based on a sound source localization algorithm and the frequency weight of each high-frequency point; and determining the angle that maximizes the function value of the objective function as the sound source angle. This solution can effectively improve the accuracy of sound source localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to speech signal processing technology. More specifically, it relates to a sound source localization method, apparatus, and electronic device. Background Technology

[0002] In voice interaction scenarios, sound source localization is a crucial step. By estimating the direction-of-arrival (DOA) of the voice signal as it reaches the microphone array, sound source localization can be achieved.

[0003] Currently, microphone array-based sound source localization methods can be broadly classified into three categories: controllable beamforming technology based on maximum output power, high-resolution spectral estimation technology, and sound source localization technology based on time-delay estimation (TDE).

[0004] However, due to the complexity of the external environment, the accuracy and robustness of the sound source angle obtained by using existing sound source localization technology need to be improved. Summary of the Invention

[0005] To solve the above-mentioned technical problems, or at least partially solve them, this application provides a sound source localization method, apparatus, and electronic device, which can effectively improve the accuracy of sound source localization.

[0006] In a first aspect, embodiments of this application provide a sound source localization method, the method comprising: acquiring sound signals through a microphone array; determining the energy growth amplitude of each frequency point of the sound signal; determining multiple candidate frequency points from the frequency points based on the energy growth amplitude of each frequency point, wherein the energy growth amplitude of each candidate frequency point is greater than or equal to an amplitude threshold, the multiple candidate frequency points including multiple mid-frequency points and multiple high-frequency points, wherein the frequency of each mid-frequency point is greater than a low-frequency threshold and less than or equal to a mid-frequency threshold, and the frequency of each high-frequency point is greater than a mid-frequency threshold; determining the relative time delay between any two adjacent microphones in the microphone array for transmitting each mid-frequency point as multiple first time delays; determining multiple time delay intervals based on each first time delay, wherein the midpoint of each time delay interval is the corresponding first time delay interval. Delay; determine a reference delay from multiple first delays, the reference delay corresponding to the delay interval that covers the most of the multiple first delays; determine the second delay between any two adjacent microphones in the microphone array corresponding to each high-frequency point, obtaining multiple second delays; based on the multiple second delays and the reference delay, determine the frequency weight of each high-frequency point, the smaller the absolute value of the difference between a second delay and the reference delay, the greater the frequency weight of the high-frequency point corresponding to the second delay; based on the sound source localization algorithm and the frequency weight of each high-frequency point, determine the objective function, the objective function is used to characterize the weighted angle function of each high-frequency point, the angle function is determined based on the sound source localization algorithm; the angle that maximizes the function value of the objective function is determined as the sound source angle.

[0007] In some embodiments of this application, the amplitude threshold is the energy growth amplitude at the Nth position when the energy growth amplitudes of each frequency point are arranged from largest to smallest, where N is an integer greater than 1.

[0008] In some embodiments of this application, multiple candidate frequencies are determined from each frequency point based on the energy growth amplitude of each frequency point, including: determining multiple candidate frequencies from each frequency point based on the energy growth amplitude of each frequency point, wherein the energy growth amplitude of each candidate frequency point is greater than or equal to an amplitude threshold; deleting the frequency points corresponding to transient noise from the multiple candidate frequency points to obtain multiple candidate frequencies.

[0009] In some embodiments of this application, multiple delay intervals are determined based on each first delay, including: determining multiple delay intervals based on each intermediate frequency point and the corresponding first delay; wherein, the higher the frequency of an intermediate frequency point, the smaller the interval length of the delay interval corresponding to the intermediate frequency point.

[0010] In some embodiments of this application, the length of the delay interval is inversely proportional to the frequency of the intermediate frequency point.

[0011] In some embodiments of this application, the interval length is determined by the following formula:

[0012]

[0013] Where τ represents the first time delay corresponding to the intermediate frequency point k, represents a positive integer less than 10, c represents the speed of sound, and μ0 represents the distance between any two adjacent microphones in the microphone array.

[0014] In some embodiments of this application, determining the frequency weight of each high-frequency point based on multiple second delays and reference delays includes: determining the frequency weight of each high-frequency point using a target formula based on multiple second delays and reference delays; the target formula is:

[0015]

[0016] Where π represents the mathematical constant pi, and exp represents the exponent with the natural logarithm e as the base. This indicates the second time delay corresponding to the high-frequency point k. σ represents the reference time delay, c represents the standard deviation of W(n,k), μ0 represents the distance between any two adjacent microphones in the microphone array, and |.| represents taking the absolute value.

[0017] In some embodiments of this application, the microphone array is a linear array or a circular array.

[0018] Secondly, this application provides a sound source localization device, comprising: a acquisition module and a determination module; the acquisition module is used to acquire sound signals through a microphone array; the determination module is used to determine the energy growth amplitude of each frequency point of the sound signal; the determination module is further used to determine multiple candidate frequency points from the multiple frequency points based on the energy growth amplitude of each frequency point, wherein the energy growth amplitude of each candidate frequency point is greater than or equal to an amplitude threshold, the multiple candidate frequency points include multiple mid-frequency points and multiple high-frequency points, the frequency of each mid-frequency point is greater than a low-frequency threshold and less than or equal to the mid-frequency threshold, and the frequency of each high-frequency point is greater than the mid-frequency threshold; the determination module is further used to determine the relative time delay between any two adjacent microphones in the microphone array for each mid-frequency point as multiple first time delays; the determination module is further used to determine multiple time delay intervals based on each first time delay, wherein the midpoint of each time delay interval is a corresponding... The determining module is further configured to determine a reference delay from multiple first delays, the reference delay corresponding to the delay interval that covers the most of the multiple first delays among multiple delay intervals; the determining module is further configured to determine the second delay between any two adjacent microphones in the microphone array corresponding to each high-frequency point, thereby obtaining multiple second delays; the determining module is further configured to determine the frequency weight of each high-frequency point based on the multiple second delays and the reference delay, the smaller the absolute value of the difference between a second delay and the reference delay, the greater the frequency weight of the high-frequency point corresponding to the second delay; the determining module is further configured to determine an objective function based on the sound source localization algorithm and the frequency weight of each high-frequency point, the objective function being used to characterize the weighted angle function of each high-frequency point, the angle function being determined based on the sound source localization algorithm; the determining module is further configured to determine the angle that maximizes the function value of the objective function as the sound source angle.

[0019] In some embodiments of this application, the amplitude threshold is the energy growth amplitude at the Nth position when the energy growth amplitudes of each frequency point are arranged from largest to smallest, where N is an integer greater than 1.

[0020] In some embodiments of this application, the sound source localization device further includes a deletion module; the determination module is specifically used to determine multiple candidate frequency points from each frequency point based on the energy growth amplitude of each frequency point, wherein the energy growth amplitude of each candidate frequency point is greater than or equal to an amplitude threshold; the deletion module is used to delete the frequency points corresponding to transient noise from the multiple candidate frequency points to obtain multiple candidate frequency points.

[0021] In some embodiments of this application, the determining module is specifically used to determine multiple delay intervals based on each intermediate frequency point and the corresponding first delay; wherein, the higher the frequency of an intermediate frequency point, the smaller the interval length of the delay interval corresponding to the intermediate frequency point.

[0022] In some embodiments of this application, the length of the delay interval is inversely proportional to the frequency of the intermediate frequency point.

[0023] In some embodiments of this application, the interval length is determined by the following formula:

[0024]

[0025] Where τ represents the first time delay corresponding to the intermediate frequency point k, represents a positive integer less than 10, c represents the speed of sound, and μ0 represents the distance between any two adjacent microphones in the microphone array.

[0026] In some embodiments of this application, the determining module is specifically used to determine the frequency weight of each high-frequency point based on a target formula using multiple second delays and reference delays.

[0027] The target formula is:

[0028]

[0029] Where π represents the mathematical constant pi, and exp represents the exponent with the natural logarithm e as the base. This indicates the second time delay corresponding to the high-frequency point k. σ represents the reference time delay, c represents the standard deviation of W(n,k), μ0 represents the distance between any two adjacent microphones in the microphone array, and |.| represents taking the absolute value.

[0030] In some embodiments of this application, the microphone array is a linear array or a circular array.

[0031] Thirdly, this application provides a computer-readable storage medium, comprising: storing a computer program on the computer-readable storage medium, wherein when the computer program is executed by a processor, it implements the sound source localization method as shown in the first aspect.

[0032] Fourthly, this application provides a computer program product, including: when the computer program product is run on a computer, causing the computer to implement the sound source localization method as shown in the first aspect.

[0033] Compared with the prior art, the technical solution provided in this application has the following advantages: In this application embodiment, sound signals are acquired through a microphone array; the energy growth amplitude of each frequency point of the sound signal is determined; based on the energy growth amplitude of each frequency point, multiple candidate frequency points are determined from each frequency point, the energy growth amplitude of each candidate frequency point is greater than or equal to an amplitude threshold, the multiple candidate frequency points include multiple mid-frequency points and multiple high-frequency points, the frequency of each mid-frequency point is greater than a low-frequency threshold and less than or equal to a mid-frequency threshold, and the frequency of each high-frequency point is greater than a mid-frequency threshold; the relative time delay between each mid-frequency point and any two adjacent microphones in the microphone array is determined as multiple first time delays; multiple time delay intervals are determined based on each first time delay, and the midpoint of each time delay interval is the corresponding... The process involves: determining the first time delay; identifying a reference time delay from multiple first time delays, where the reference time delay corresponds to the time delay interval that covers the most of the multiple first time delays; determining the second time delay between any two adjacent microphones in the microphone array corresponding to each high-frequency point, resulting in multiple second time delays; determining the frequency weight of each high-frequency point based on the multiple second time delays and the reference time delay, where the smaller the absolute value of the difference between a second time delay and the reference time delay, the greater the frequency weight of the high-frequency point corresponding to that second time delay; determining the objective function based on the sound source localization algorithm and the frequency weight of each high-frequency point, where the objective function is used to characterize the weighted angle function of each high-frequency point, and the angle function is determined based on the sound source localization algorithm; and determining the angle that maximizes the function value of the objective function as the sound source angle. Thus, by focusing only on the mid-frequency and high-frequency bands for each frequency point corresponding to the sound signal and eliminating some low-frequency band frequencies to reduce the amount of computation, the efficiency of sound source localization is improved. By performing a rough estimation in the mid-frequency band to obtain the most likely relative time delay between the two microphones as a reference time delay, and performing a precise estimation in the high-frequency band, smaller weights are assigned to the frequency points corresponding to time delays that differ significantly from the reference time delay, reducing the impact of these frequency points on the results of sound source localization, thereby improving the accuracy of sound source angle estimation. Attached Figure Description

[0034] To more clearly illustrate the implementation methods in the embodiments of this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0035] Figure 1 This illustrates an operational scenario between a microphone array and an electronic device according to some embodiments;

[0036] Figure 2 One of the flowcharts of a sound source localization method according to some embodiments is shown;

[0037] Figure 3 A schematic diagram of a uniform linear array according to some embodiments is shown;

[0038] Figure 4 A schematic diagram of a planar array according to some embodiments is shown;

[0039] Figure 5 A schematic diagram of a three-dimensional array according to some embodiments is shown;

[0040] Figure 6 A second schematic flowchart of a sound source localization method according to some embodiments is shown;

[0041] Figure 7 A third schematic flowchart of a sound source localization method according to some embodiments is shown;

[0042] Figure 8 A fourth schematic flowchart of a sound source localization method according to some embodiments is shown;

[0043] Figure 9 A schematic diagram of a sound source localization device according to some embodiments is shown;

[0044] Figure 10 A schematic diagram of the structure of an electronic device according to some embodiments is shown. Detailed Implementation

[0045] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.

[0046] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0047] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0048] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0049] The display device provided in this application can have various implementation forms, such as a television, a smart television, a laser projection device, a monitor, an electronic bulletin board, an electronic table, a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, etc.

[0050] For ease of understanding, a brief introduction to the technical aspects involved in this application is provided below.

[0051] Sound source localization technology refers to measuring sound signals using multiple microphones at different locations in the environment. Since the sound signals arrive at each microphone with varying degrees of delay, algorithms process the measured sound signals to obtain the direction of arrival (sound source angle) of the sound source relative to the microphones. Direction-of-arrival (DOA) based sound source localization methods estimate the direction in which the microphone array receives the sound source signal at each node to obtain the angle of the sound source relative to the microphone array. Multiple nodes can then determine a series of radial lines from the sound source to different microphone nodes; the intersection of these lines is the estimated sound source location. Existing methods for sound source localization mainly include methods based on relative time delay estimation, methods based on beamforming, and methods based on signal subspace.

[0052] The relative delay estimation method utilizes the geometry of the microphone array, where the signals received by each array element have different degrees of delay. The relative delay estimation method estimates the time delay difference between each sound arriving at each array element through cross-correlation, generalized cross-correlation (GCC), or phase difference, and then combines the array geometry to estimate the sound source angle.

[0053] Beamforming-based methods typically use phase compensation for all array elements to scan the target area, then perform a weighted sum of the signals, and take the direction with the highest beam output power as the direction of the target sound source. Common beamforming-based sound source azimuth estimation algorithms include Delay and Sum (DS), Minimum Variance Distortionless Response (MVDR), and Steered Response Power-Phase Transform (SRP-PHAT).

[0054] Signal subspace-based methods: These algorithms can generally be divided into coherent subspace methods and incoherent subspace methods. Among incoherent subspace algorithms, the most classic is the Multiple Signal Classification (MUSIC) algorithm. Its idea is to extract features from the signal's covariance, construct signal and noise subspaces using the feature vectors, and then construct a high-resolution spatial spectrum from the noise subspace. Since the sound source signal is a broadband signal, it can be decomposed into multiple narrowband signals using Fourier transform, and then the MUSIC algorithm is used to locate each narrowband signal. The results of the narrowband estimations are then weighted and combined to obtain a broadband azimuth estimate. Coherent subspace methods, on the other hand, converge the narrowband signals to a reference frequency, and then use narrowband subspace processing methods for azimuth estimation.

[0055] The traditional methods mentioned above are all based on the assumption that the energy of the speech signal is the greatest and is higher than that of noise and interference signals. However, in voice interaction scenarios, the human voice is not necessarily the strongest signal. In low signal-to-noise ratio environments, the accuracy of sound source localization of traditional methods will be seriously affected, which in turn will affect the effect of speech enhancement. Furthermore, the estimation of the sound source angle is based on a one-time estimation of all received sound signals, and the accuracy of sound source localization needs to be improved.

[0056] like Figure 1 The diagram shown is a structural diagram of a sound source localization system provided in an embodiment of this application, including a microphone array 100 and an electronic device 200. The microphone array is used to collect sound signals emitted by a sound source, and the electronic device is used to process the sound signals collected by the microphone array to determine the angle of the sound source and realize sound source localization.

[0057] In the embodiments of this application, such as Figure 2 As shown, a sound source localization method is provided, which includes the following steps 101 to 110.

[0058] 101. Acquire sound signals through a microphone array.

[0059] A microphone array is essentially a sound acquisition system composed of a certain number of microphones. It samples and filters the spatial characteristics of a sound field. Multiple microphones collect sound from different spatial directions. Based on the pre-arranged order of the microphones and corresponding algorithms (arrangement + algorithm), many acoustic problems can be solved, such as sound source localization, dereverberation, speech enhancement, and blind source separation. Commonly used microphone arrays can be classified by their layout shape into: linear arrays, planar arrays (circular arrays), and stereo arrays. All microphones have the same frequency response, and their sampling clocks are synchronized. Figure 3The aforementioned array is a uniform linear array (linear arrays also include nested arrays, i.e., non-uniform linear arrays), with equal distances between adjacent microphones; for example... Figure 4 The image shows a planar array. Common planar arrays include the "4+1" type, with four microphones evenly distributed around the circumference and one microphone at the center; the "6+0" type, with six microphones evenly distributed around the circumference and no microphone at the center; and the "7+1" type, with seven microphones evenly distributed around the circumference and one microphone at the center. Figure 5 As shown, this is a three-dimensional array. Common three-dimensional arrays include cylindrical arrays and spherical arrays.

[0060] Optionally, the microphone array used in the embodiments of this application can be a linear array or a circular array.

[0061] 102. Determine the energy increase amplitude at each frequency point of the sound signal.

[0062] It is understandable that the received signal is in the time domain. In order to facilitate calculation, the received sound signal needs to be processed. The specific processing includes: performing a short-time Fourier transform on the sound signal, calculating the amplitude mean of the cross power spectrum based on the transformed sound signal, determining the power envelope of each frequency point of the sound signal based on the amplitude mean of the cross power spectrum, and determining the energy growth amplitude of each frequency point based on the power envelope.

[0063] Specifically, the mean cross-spectral amplitude C(n,k) of the collected sound signal is calculated as shown in formula (1):

[0064]

[0065] Logarithmic transformation of C(n,k) yields the power envelope p(n,k) of the signal within the frequency range, as shown in equation (2):

[0066] p(n,k)=log 10 (C(n,k)+ε) (2)

[0067] Where, x i (n,k) is the short-time Fourier transform of the audio signal, I represents the number of microphones, n represents the frame number, ε is a regularization term used to reduce the influence of background noise, |.| represents the absolute value of a complex number, and the superscript * represents conjugate.

[0068] The increase in frequency energy is shown in formula (3):

[0069]

[0070] Where ΔP(n,k) represents the energy growth rate at a frequency point, i.e., the "rate of change" of the logarithmically transformed mean cross-spectral amplitude, N tP(nt,k) represents the frame number range, and is the power envelope of the frame number nt with frequency k that is t frames earlier than P(n,k).

[0071] It is understandable that the energy increase of each frequency point corresponding to the sound signal can be calculated using the above formula.

[0072] 103. Based on the energy growth rate of each frequency point, determine multiple candidate frequency points from each frequency point.

[0073] Among them, the energy growth amplitude of each candidate frequency point is greater than or equal to the amplitude threshold, and the multiple candidate frequency points include multiple intermediate frequency points and multiple high frequency points. The frequency of each intermediate frequency point is greater than the low frequency threshold and less than or equal to the intermediate frequency threshold, and the frequency of each high frequency point is greater than the intermediate frequency threshold.

[0074] It can be understood that the energy growth rate of each candidate frequency point is greater than or equal to the amplitude threshold, that is, the energy growth rate of each candidate frequency point is greater than the energy growth rate of any frequency point other than the multiple candidate frequency points.

[0075] It is understood that the frequency points obtained based on the sound signal include low-frequency, mid-frequency, and high-frequency points. However, since there are relatively few low-frequency points of the sound signal that are of interest in sound source localization, in order to reduce the subsequent computational load, this application performs sound source localization based on mid-frequency and high-frequency points and removes low-frequency points. At the same time, it can also reduce white noise to a certain extent, reduce the influence of noise, and increase the accuracy of sound source localization.

[0076] It is understood that in this embodiment of the application, the low-frequency threshold and the mid-frequency threshold are not specifically limited, but are limited according to the frequency of the sound signal of interest in the actual sound source localization.

[0077] For example, if the low-frequency threshold is set to 1000 Hz and the mid-frequency threshold is set to 2000 Hz, then the frequency points with a frequency value less than or equal to 1000 Hz are low-frequency points, the frequency points with a frequency value greater than 1000 Hz and less than or equal to 2000 Hz are mid-frequency points, and the frequency points with a frequency value greater than 2000 Hz are high-frequency points.

[0078] Optionally, the amplitude threshold can be a preset value (determined based on historical data or experience), or the amplitude threshold can be the energy growth amplitude of each frequency point arranged from largest to smallest, where N is an integer greater than 1.

[0079] It is understood that N can be determined according to actual needs, and this application embodiment does not limit it. For example, N can be set to 100, 1000 or other numbers. The more numbers there are, the more accurate the calculation results will be.

[0080] It can be understood that the amplitude threshold is the energy growth amplitude of each frequency point arranged from largest to smallest, and the energy growth amplitude at the Nth position. That is, the frequency points are sorted according to the energy growth amplitude from largest to smallest to determine the number of candidate frequency points needed. For example, if 1000 candidate frequency points are needed, then N is 1000. The energy growth amplitude of the frequency point ranked 1000th is used as the amplitude threshold, and the frequency points after 1000th are filtered out.

[0081] It can be understood that the amplitude threshold is the energy growth amplitude of each frequency point arranged from largest to smallest, and the energy growth amplitude at the Nth position. In this way, the number of candidate frequency points required can be determined according to the required level of accuracy. If more emphasis is placed on computational efficiency, N can be appropriately reduced; if more emphasis is placed on the accuracy of the calculation results, N can be appropriately increased. The choice can be made flexibly according to actual needs.

[0082] 104. The relative time delay between any two adjacent microphones in the microphone array for each intermediate frequency point is determined as multiple first time delays.

[0083] It can be understood that the relative time delay between any two adjacent microphones in the corresponding microphone array transmitted from each intermediate frequency point is determined as a first time delay. One intermediate frequency point corresponds to one first time delay. Based on multiple intermediate frequency points, multiple first time delays are obtained.

[0084] It is understandable that in sound source localization, the cross-correlation method is generally used to calculate the time delay, and this application does not limit the specific algorithm used.

[0085] For example, the first time delay is obtained through the SRP-PHAT algorithm, specifically, the angle with the maximum cross-correlation is obtained through the following formula (4):

[0086]

[0087] Based on formula (4), the relative time delay between two adjacent microphones can be obtained as shown in formula (5):

[0088]

[0089] in, Let θ represent the angle with the highest cross-correlation, where θ represents the possible values ​​of the angle, such as 10, 20, 30, 40, 50...180, which is the independent variable. argmax represents the value of the independent variable that maximizes the expression. g(k,θ) represents the steering vector in the direction of θ. (Applicable to linear arrays), j represents the imaginary unit. ω kHere, ω represents the angular frequency, μ represents the spatial vector of the microphone, c represents the speed of sound, exp represents the exponent with the natural logarithm e as the base, μ0 represents the distance between any two adjacent microphones in the microphone array, and the superscript H indicates the complex conjugate transpose. x i (n,k) is the short-time Fourier transform of the sound signal, and the superscript * indicates conjugate. The first time delay corresponding to the intermediate frequency point k is indicated. It should be noted that the frequency point k involved in formulas (4) and (5) are all intermediate frequency points.

[0090] 105. Based on each first delay, determine multiple delay intervals, and the midpoint of each delay interval is the corresponding first delay.

[0091] It is understandable that a time delay interval is determined based on a mid-frequency point and the corresponding first time delay. Multiple time delay intervals are calculated through multiple mid-frequency points. That is, there is a one-to-one correspondence between the mid-frequency point and the first time delay, and a one-to-one correspondence between the first time delay and the time delay interval.

[0092] It can be understood that the midpoint of each delay interval is the corresponding first delay. That is, for a delay interval, the midpoint of the interval is the corresponding first delay. For example, if the first delay is 3, then the corresponding delay interval can be [1, 5], that is, the midpoint of 1 to 5 is 3.

[0093] 106. Determine a reference delay from multiple first delays. The reference delay corresponds to the delay interval that covers the most of the multiple first delays.

[0094] It can be understood that the reference delay is the delay corresponding to the target delay interval among multiple first delays, and the target delay interval is the delay interval that covers the most multiple first delays among multiple delay intervals.

[0095] For example, the delay 1 corresponding to frequency 1000 is 5, and the delay interval is [0, 10]. The delay 2 corresponding to frequency 2000 is 3, and the delay interval is [1, 5]. The delay 3 corresponding to frequency 5000 is 6, and the delay interval is [5, 7]. Then, the delays falling within [0, 10] include delay 1, delay 2, and delay 3, for a total of three delays. The delays falling within [1, 5] include delay 1 and delay 2, for a total of two delays. The delays falling within [5, 7] include delay 1 and delay 3, for a total of two delays. The number of delays falling within the delay interval [0, 10] is the largest. Therefore, the delay 5 corresponding to the delay interval [0, 10] is the reference delay.

[0096] It can be understood that the reference delay is the most likely relative delay among the first delays corresponding to all intermediate frequency points, specifically obtained through the following formula (6):

[0097]

[0098] in, This represents the set of intermediate frequency points, where any frequency point in this set is an intermediate frequency point. B(.) indicates that the result is 1 when the parameter within the parentheses is true, and 0 when the parameter within the parentheses is false. Representing the mid-frequency range, the more mid-frequency points a delay interval corresponding to a certain first delay covers, the more likely that the first delay corresponding to that delay interval is the relative delay of the actual sound signal arriving at two adjacent microphones. This is the reference delay.

[0099] 107. Determine the second time delay between any two adjacent microphones in the microphone array corresponding to each high-frequency point, and obtain multiple second time delays.

[0100] It can be understood that the relative time delay between any two adjacent microphones in the corresponding microphone array for each high-frequency point is determined as a second time delay. One high-frequency point corresponds to one second time delay, and multiple second time delays are obtained based on multiple high-frequency points.

[0101] It should be noted that the second delay is obtained through the SRP-PHAT algorithm. Specifically, it can be calculated using the above formulas (4) and (5). However, the frequency points k involved are all high-frequency points. For the specific process and related descriptions, please refer to the descriptions of the above formulas (4) and (5).

[0102] 108. Determine the frequency weight of each high-frequency point based on multiple second delays and reference delays.

[0103] Among them, the smaller the absolute value of the difference between a second delay and the reference delay, the greater the frequency point weight of the high-frequency point corresponding to the second delay.

[0104] It is understood that the frequency weight of each high-frequency point can be preset. For example, if the absolute value of the difference between the second delay and the reference delay belongs to the first range, then the frequency weight is the first weight; if the absolute value of the difference between the second delay and the reference delay belongs to the second range, then the frequency weight is the second weight; if the absolute value of the difference between the second delay and the reference delay belongs to the third range, then the frequency weight is the third weight. The first weight is greater than the second weight, the second weight is greater than the third weight, any value in the first range is less than any value in the second range, and any value in the second range is less than any value in the third range. The frequency weight of each high-frequency point can also be calculated by a function. Specifically, this application does not limit the specific implementation.

[0105] 109. Determine the objective function based on the sound source localization algorithm and the frequency weight of each high-frequency point.

[0106] The objective function is used to characterize the weighted sum of the angle functions for each high-frequency point, and the angle functions are determined based on the sound source localization algorithm.

[0107] 110. The angle that maximizes the value of the objective function is determined as the sound source angle.

[0108] It is understood that the specific sound source localization algorithm is not limited in the embodiments of this application.

[0109] For example, if the sound source localization algorithm is based on the SRP-PHAT algorithm, then the objective function is as shown in formula (7):

[0110]

[0111] The angle of the sound source can be determined by formula (8):

[0112]

[0113] in, Let W(n,k) represent the set of high-frequency points, where any frequency point in this set is a high-frequency point. Let W(n,k) represent the frequency point weight, and g(k,θ) represent the steering vector in the θ direction. (Applicable to linear arrays), j represents the imaginary unit. ω k Here, μ represents the angular frequency, c represents the speed of sound, exp represents the exponent with the natural logarithm e as the base, and the superscript H indicates the complex conjugate transpose. x i (n,k) is the short-time Fourier transform of the sound signal, and the superscript * indicates conjugate. The angle of the sound source is represented by θ, which represents the possible values ​​of the angle, such as 10, 20, 30, 40, 50...180, which is the independent variable. argmax represents the value of the independent variable that makes the expression reach its maximum value.

[0114] In this embodiment, sound signals are acquired using a microphone array; the energy growth amplitude of each frequency point of the sound signal is determined; based on the energy growth amplitude of each frequency point, multiple candidate frequency points are determined from each frequency point, wherein the energy growth amplitude of each candidate frequency point is greater than or equal to an amplitude threshold, and the multiple candidate frequency points include multiple mid-frequency frequency points and multiple high-frequency frequency points, wherein the frequency of each mid-frequency frequency point is greater than a low-frequency threshold and less than or equal to a mid-frequency threshold, and the frequency of each high-frequency frequency point is greater than a mid-frequency threshold; the relative time delay between each mid-frequency frequency point and any two adjacent microphones in the microphone array is determined as multiple first time delays; multiple time delay intervals are determined based on each first time delay, and the midpoint of each time delay interval is the corresponding first time delay; from the multiple first delays... A reference delay is determined, which corresponds to the delay interval that covers the most first delays among multiple delay intervals. The second delay between any two adjacent microphones in the microphone array corresponding to each high-frequency point is determined, resulting in multiple second delays. Based on the multiple second delays and the reference delay, the frequency weight of each high-frequency point is determined; the smaller the absolute value of the difference between a second delay and the reference delay, the greater the frequency weight of the high-frequency point corresponding to that second delay. Based on the sound source localization algorithm and the frequency weight of each high-frequency point, an objective function is determined. The objective function is used to characterize the weighted angle function of each high-frequency point, and the angle function is determined based on the sound source localization algorithm. The angle that maximizes the function value of the objective function is determined as the sound source angle. Thus, by focusing only on the mid-frequency and high-frequency bands for each frequency point corresponding to the sound signal and eliminating some low-frequency band frequencies to reduce the amount of computation, the efficiency of sound source localization is improved. By performing a rough estimation in the mid-frequency band to obtain the most likely relative time delay between the two microphones as a reference time delay, and performing a precise estimation in the high-frequency band, smaller weights are assigned to the frequency points corresponding to time delays that differ significantly from the reference time delay, reducing the impact of these frequency points on the results of sound source localization, thereby improving the accuracy of sound source angle estimation.

[0115] In some embodiments of this application, combined with Figure 1 ,like Figure 6 As shown, step 103 can be implemented through steps 103a to 103b.

[0116] 103a. Based on the energy growth rate of each frequency point, determine multiple candidate frequency points from each frequency point.

[0117] Among them, the energy growth rate of each candidate frequency point is greater than or equal to the amplitude threshold.

[0118] 103b. Delete the frequency points corresponding to transient noise from multiple candidate frequency points to obtain multiple candidate frequency points.

[0119] It is understandable that transient noise is characterized by high energy and short duration. Therefore, transient noise can be identified and its corresponding frequency points deleted through the following steps S1 and S2:

[0120] S1. Calculate the energy of each frame using the following formula, and determine the frames with local maxima of energy.

[0121] P t (n)=∑ k P(n,k),n v ={n|P t (n+1)-P t (n)<0,P t (n)-P t (n-1)>0}

[0122] Among them, P t (n) represents the energy of any frame, and n+1 and n-1 represent the frames adjacent to n, respectively. v The frame represents a short-lived frame with extremely high energy, where k is the frequency point.

[0123] S2, Determine n v Does it meet conditions 1 and 2 below?

[0124] Condition 1:

[0125] Condition 2:

[0126] Where Δn represents the custom frame range (i.e., the time window, the specific size of which is determined according to actual needs, such as the length of 70 speech frames), V1 represents the energy rise threshold, and V2 represents the energy fall threshold.

[0127] If n is determined in step S1 v If conditions 1 and 2 of step S2 are met, the number of frames in the candidate frequency points will be... Frequency points within the range are deleted to obtain multiple candidate frequency points.

[0128] In this embodiment, based on the energy growth amplitude of each frequency point, multiple candidate frequency points are determined from each frequency point, where the energy growth amplitude of each candidate frequency point is greater than or equal to an amplitude threshold. Frequency points corresponding to transient noise are then deleted from the multiple candidate frequency points to obtain multiple candidate frequency points. Thus, deleting transient noise ensures that it will not affect the localization result during subsequent sound source localization calculations, further improving the accuracy of sound source localization. Simultaneously, deleting transient noise also reduces the subsequent computational load, making sound source localization more efficient.

[0129] In some embodiments of this application, combined with Figure 1 ,like Figure 7As shown, step 105 above can be implemented through step 105a below.

[0130] 105a. Determine multiple delay intervals based on each intermediate frequency point and the corresponding first delay.

[0131] Among them, the higher the frequency of an intermediate frequency point, the smaller the interval length of the time delay interval corresponding to that intermediate frequency point.

[0132] It is understandable that a delay interval is determined based on each intermediate frequency point and its corresponding first delay, and multiple intermediate frequency points and their corresponding multiple first delays determine multiple delay intervals.

[0133] It is understandable that a time delay interval is determined based on the frequency value corresponding to an intermediate frequency point and the intermediate frequency point and the corresponding first time delay. The higher the frequency of the intermediate frequency point, the smaller the interval length of the time delay interval corresponding to that intermediate frequency point (after multiple experimental data verifications, this method can improve the accuracy of time delay estimation). The purpose is to alleviate the error of time delay estimation and make the estimation of the reference time delay more accurate.

[0134] Optionally, the length of the delay interval is inversely proportional to the frequency of the intermediate frequency point.

[0135] For example, the delay interval can be: Where τ represents the first time delay corresponding to the intermediate frequency point k, Let c represent a positive integer less than 10, c represent the speed of sound, and μ0 represent the distance between any two adjacent microphones in the microphone array. Specifically, as long as the length of the time delay interval is inversely proportional to the frequency of the intermediate frequency point, the formula is not specifically limited in the embodiments of this application.

[0136] Alternatively, the interval length is determined by the following formula:

[0137]

[0138] Where τ represents the first time delay corresponding to the intermediate frequency point k, represents a positive integer less than 10, c represents the speed of sound, and μ0 represents the distance between any two adjacent microphones in the microphone array.

[0139] It is understandable that, according to the formula The determined time delay interval is:

[0140] It is understandable that the difference in the first delay corresponding to each frequency point is relatively small, therefore the formula... In this equation, the main factor affecting the interval length is k; that is, the higher the frequency of k, the smaller the corresponding interval length. The formula for determining the interval length can be flexibly chosen according to actual needs.

[0141] In this embodiment, multiple delay intervals are determined based on each intermediate frequency point and its corresponding first delay; wherein, the higher the frequency of an intermediate frequency point, the smaller the interval length of the delay interval corresponding to that intermediate frequency point. This can mitigate delay estimation errors and make the estimation of the reference delay more accurate.

[0142] In some embodiments of this application, combined with Figure 1 ,like Figure 8 As shown, step 108 above can be implemented through step 108a below.

[0143] 108a. Based on multiple second delays and reference delays, determine the frequency weight of each high-frequency point through the target formula.

[0144] The target formula is:

[0145]

[0146] Where π represents the mathematical constant pi, and exp represents the exponent with the natural logarithm e as the base. This indicates the second time delay corresponding to the high-frequency point k. σ represents the reference time delay, c represents the standard deviation of W(n,k), μ0 represents the distance between any two adjacent microphones in the microphone array, and |.| represents taking the absolute value.

[0147] It is understandable that the target formula is a Gaussian function. For frequency points with a large difference from the reference time delay, the weight is set to 0 to reduce the impact of the frequency points corresponding to that time delay on the results during the sound source localization process.

[0148] In this embodiment, the frequency weight of each high-frequency point is determined by a target formula based on multiple second delays and a reference delay. Thus, for delays that differ significantly from the reference delay, the frequency weight of the corresponding frequency point is set to 0, meaning that this frequency point is not considered during sound source localization. This reduces the influence of irrelevant frequency points on the sound source localization results, thereby improving the accuracy of sound source localization.

[0149] Figure 9 This is a structural block diagram of a sound source localization device shown in an embodiment of this application, such as... Figure 9As shown, the device includes: an acquisition module 901 and a determination module 902; the acquisition module 901 is used to acquire sound signals through a microphone array; the determination module 902 is used to determine the energy growth amplitude of each frequency point of the sound signal; the determination module 902 is further used to determine multiple candidate frequency points from each frequency point based on the energy growth amplitude of each frequency point, wherein the energy growth amplitude of each candidate frequency point is greater than or equal to an amplitude threshold, the multiple candidate frequency points include multiple mid-frequency frequency points and multiple high-frequency frequency points, the frequency of each mid-frequency frequency point is greater than a low-frequency threshold and less than or equal to the mid-frequency threshold, and the frequency of each high-frequency frequency point is greater than the mid-frequency threshold; the determination module 902 is further used to determine the relative time delay between any two adjacent microphones in the microphone array for each mid-frequency frequency point as multiple first time delays; the determination module 902 is further used to determine multiple time delay intervals based on each first time delay, wherein the midpoint of each time delay interval is the corresponding first time delay. The determining module 902 is further configured to determine a reference delay from multiple first delays, wherein the reference delay corresponds to the delay interval that covers the most of the multiple first delays among the multiple delay intervals; the determining module 902 is further configured to determine the second delay between any two adjacent microphones in the microphone array corresponding to each high-frequency point, thereby obtaining multiple second delays; the determining module 902 is further configured to determine the frequency weight of each high-frequency point based on the multiple second delays and the reference delay, wherein the smaller the absolute value of the difference between a second delay and the reference delay, the greater the frequency weight of the high-frequency point corresponding to the second delay; the determining module 902 is further configured to determine an objective function based on the sound source localization algorithm and the frequency weight of each high-frequency point, wherein the objective function is used to characterize the weighted angle function of each high-frequency point, and the angle function is determined based on the sound source localization algorithm; the determining module is further configured to determine the angle that maximizes the function value of the objective function as the sound source angle.

[0150] In some embodiments of this application, the amplitude threshold is the energy growth amplitude at the Nth position when the energy growth amplitudes of each frequency point are arranged from largest to smallest, where N is an integer greater than 1.

[0151] In some embodiments of this application, the sound source localization device further includes a deletion module 903; the determination module 902 is specifically used to determine multiple candidate frequency points from each frequency point based on the energy growth amplitude of each frequency point, wherein the energy growth amplitude of each candidate frequency point is greater than or equal to an amplitude threshold; the deletion module 903 is used to delete the frequency points corresponding to transient noise from the multiple candidate frequency points to obtain multiple candidate frequency points.

[0152] In some embodiments of this application, the determining module 902 is specifically used to determine multiple delay intervals based on each intermediate frequency point and the corresponding first delay; wherein, the higher the frequency of an intermediate frequency point, the smaller the interval length of the delay interval corresponding to the intermediate frequency point.

[0153] In some embodiments of this application, the length of the delay interval is inversely proportional to the frequency of the intermediate frequency point.

[0154] In some embodiments of this application, the interval length is determined by the following formula:

[0155]

[0156] Where τ represents the first time delay corresponding to the intermediate frequency point k, represents a positive integer less than 10, c represents the speed of sound, and μ0 represents the distance between any two adjacent microphones in the microphone array.

[0157] In some embodiments of this application, the determining module 902 is specifically used to determine the frequency weight of each high-frequency point based on a target formula using multiple second delays and reference delays.

[0158] The target formula is:

[0159]

[0160] Where π represents the mathematical constant pi, and exp represents the exponent with the natural logarithm e as the base. This indicates the second time delay corresponding to the high-frequency point k. σ represents the reference time delay, c represents the standard deviation of W(n,k), μ0 represents the distance between any two adjacent microphones in the microphone array, and |.| represents taking the absolute value.

[0161] In some embodiments of this application, the microphone array is a linear array or a circular array.

[0162] It should be noted that the aforementioned sound source localization device can be the electronic device in the above method embodiment of this application, or it can be a functional module and / or functional entity in the electronic device that can realize the function of the device embodiment. This application embodiment does not limit it.

[0163] It should be noted that: such as Figure 9 As shown, the modules that must be included in the sound source localization device are indicated by solid lines, such as the acquisition module 901 and the determination module 902; the modules that may or may not be included in the sound source localization device are indicated by dashed lines, such as the deletion module 903.

[0164] In this embodiment, each module can implement the sound source localization method provided in the above method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0165] This application embodiment also provides an electronic device, which may include: a processor 1001, a memory 1002, and a program or instructions stored in the memory 1002 and executable on the processor 1001. When the program or instructions are executed by the processor 1001, they can implement the various processes of the sound source localization method provided in the above method embodiment and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0166] The present invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described sound source localization method and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0167] The computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0168] The present invention provides a computer program product, comprising: when the computer program product is run on a computer, causing the computer to implement the above-described sound source localization method.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0170] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.

Claims

1. A method for locating a sound source, characterized in that, include: Sound signals are acquired using a microphone array; Determine the energy increase amplitude at each frequency point of the sound signal; Based on the energy growth amplitude of each frequency point, multiple candidate frequency points are determined from each frequency point. The energy growth amplitude of each candidate frequency point is greater than or equal to the amplitude threshold. The multiple candidate frequency points include multiple mid-frequency points and multiple high-frequency points. The frequency of each mid-frequency point is greater than the low-frequency threshold and less than or equal to the mid-frequency threshold. The frequency of each high-frequency point is greater than the mid-frequency threshold. The relative time delay between any two adjacent microphones in the microphone array and each intermediate frequency point is transmitted to the microphone array is determined as multiple first time delays; Multiple delay intervals are determined based on each first delay, and the midpoint of each delay interval is the corresponding first delay. A reference delay is determined from multiple first delays, and the reference delay corresponds to the delay interval that covers the most of the multiple first delays among multiple delay intervals; Determine the second time delay between any two adjacent microphones in the microphone array corresponding to each high-frequency point to obtain multiple second time delays; Based on multiple second delays and reference delays, the frequency weight of each high-frequency point is determined. The smaller the absolute value of the difference between a second delay and a reference delay, the greater the frequency weight of the high-frequency point corresponding to the second delay. Based on the sound source localization algorithm and the frequency weight of each high-frequency point, an objective function is determined. The objective function is used to characterize the weighted angle function of each high-frequency point, and the angle function is determined based on the sound source localization algorithm. The angle that maximizes the objective function value is determined as the sound source angle.

2. The method according to claim 1, characterized in that, The amplitude threshold is the energy growth amplitude at the Nth position when the energy growth amplitudes of each frequency point are arranged from largest to smallest, where N is an integer greater than 1.

3. The method according to claim 1, characterized in that, The determination of multiple candidate frequency points based on the energy growth rate of each frequency point includes: Based on the energy growth amplitude of each frequency point, multiple candidate frequency points are determined from each frequency point, and the energy growth amplitude of each candidate frequency point is greater than or equal to the amplitude threshold. Delete the frequency points corresponding to transient noise from the plurality of candidate frequency points to obtain the plurality of candidate frequency points.

4. The method according to claim 1, characterized in that, The determination of multiple delay intervals based on each first delay includes: Multiple delay intervals are determined based on each intermediate frequency point and its corresponding first delay. The higher the frequency of a mid-frequency point, the smaller the interval length of the time delay interval corresponding to that mid-frequency point.

5. The method according to claim 4, characterized in that, The length of the time delay interval is inversely proportional to the frequency of the intermediate frequency point.

6. The method according to claim 4, characterized in that, The length of the interval is determined by the following formula: Where τ represents the first time delay corresponding to the intermediate frequency point k, denoted as a positive integer less than 10, c represents the speed of sound, and μ0 represents the distance between any two adjacent microphones in the microphone array.

7. The method according to claim 1, characterized in that, The step of determining the frequency weight of each high-frequency point based on the plurality of second delays and the reference delay includes: Based on the plurality of second delays and the reference delay, the frequency weight of each high-frequency point is determined by a target formula; The target formula is: Where π represents the mathematical constant pi, and exp represents the exponent with the natural logarithm e as the base. This indicates the second time delay corresponding to the high-frequency point k. σ represents the reference time delay, c represents the standard deviation of W(n,k), μ0 represents the distance between any two adjacent microphones in the microphone array, and |.| represents taking the absolute value.

8. The method according to any one of claims 1 to 7, characterized in that, The microphone array can be a linear array or a circular array.

9. A sound source localization device, characterized in that, include: The acquisition module is used to acquire sound signals through a microphone array; A determining module is used to determine the energy increase amplitude at each frequency point of the sound signal; The determining module is further configured to determine multiple candidate frequency points from the multiple frequency points based on the energy growth amplitude of each frequency point, wherein the energy growth amplitude of each candidate frequency point is greater than or equal to an amplitude threshold, the multiple candidate frequency points include multiple mid-frequency points and multiple high-frequency points, the frequency of each mid-frequency point is greater than a low-frequency threshold and less than or equal to a mid-frequency threshold, and the frequency of each high-frequency point is greater than the mid-frequency threshold; The determining module is further configured to determine the relative time delay between any two adjacent microphones in the microphone array for each intermediate frequency point as a plurality of first time delays; The determining module is further configured to determine multiple delay intervals based on each first delay, wherein the midpoint of each delay interval is the corresponding first delay; The determining module is further configured to determine a reference delay from the plurality of first delays, wherein the reference delay corresponds to the delay interval that covers the largest number of the plurality of first delays among the plurality of delay intervals; The determining module is further configured to determine the second time delay between any two adjacent microphones in the microphone array corresponding to each high-frequency point, thereby obtaining multiple second time delays; The determining module is further configured to determine the frequency weight of each high-frequency point based on the plurality of second delays and the reference delay, wherein the smaller the absolute value of the difference between a second delay and the reference delay, the greater the frequency weight of the high-frequency point corresponding to the second delay. The determining module is further configured to determine an objective function based on the sound source localization algorithm and the frequency weight of each high-frequency point. The objective function is used to characterize the weighted angle function of each high-frequency point, and the angle function is determined based on the sound source localization algorithm. The angle that maximizes the objective function value is determined as the sound source angle.

10. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the sound source localization method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Improved frequency domain SRP sound source orientation estimation method

    CN108445452A

  • Sound source positioning method and device

    CN111381211A