Tunable acoustic region detection

By employing multi-microphone acoustic processing and parametric beamforming technology, the challenge of speaker region detection in vehicle environments has been solved, improving detection accuracy and system performance, and enabling adaptation to environmental changes.

CN120958847APending Publication Date: 2025-11-14CERENCE OPERATING CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380095313.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-03-02
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively detect and locate a speaker's acoustic region, especially in in-vehicle environments, leading to inefficient configuration and control of voice interfaces and voice communication systems.

Method used

Acoustic processing is performed using multiple fixed-configuration microphones. The location of the sound source in a predetermined area is determined through parametric beamforming and tuning transformation. The parameters of the beamformer are optimized to adapt to environmental changes by utilizing the phase and amplitude relationship of the microphone signals, combined with an adaptive process and experimental data.

Benefits of technology

It improves the accuracy and efficiency of sound source area detection, reduces computational requirements, adapts to acoustic changes in the environment, and improves the performance of voice interfaces and communication systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120958847A_ABST
    Figure CN120958847A_ABST
Patent Text Reader

Abstract

A tunable region detection method utilizes a plurality of microphones in a fixed configuration in an environment. There are a plurality of regions in the environment and one or more predetermined locations in each region. A predetermined transfer function between the locations and the microphones is used to determine a beamforming energy for each of the locations based on the received microphone signals. These beamforming energies may be calculated using normalization of correlations between microphones. The beamforming energy is processed using a tunable transform to determine whether the sound source is in a particular region, thereby enabling adjustment of the detection method for situations including ambient acoustic changes.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] This invention relates to the detection and localization of sound sources, and more particularly to the acoustic region detection of sources such as the speaker.

[0002] To interact with an object that may be speaking in different areas of the environment, determining the speaker's area may be useful, for example, to configure or control a voice interface, such as for a voice assistant or voice communication system. For instance, in a vehicle environment, this area could correspond to one of a set number of different seating positions within the vehicle. Summary of the Invention

[0003] In one aspect, generally, a method for acoustic processing of an environment utilizes a plurality of microphones in a fixed configuration, the plurality of microphones providing a plurality of microphone signals for sensing the environment. The environment has a fixed plurality of predetermined locations (i.e., physical locations, optionally associated with orientation) and a plurality of predetermined regions (e.g., physical areas), each region having one or more locations located therein. The acoustic processing includes detecting sound sources in the predetermined regions of the environment. Sound source detection includes receiving the plurality of microphone signals and then processing these signals to determine phase and amplitude relationships between pairs of microphone signals at a plurality of frequencies. These relationships between the microphone pairs are processed according to a parametric beamforming process to generate a frequency representation of beamforming energy for each location in the environment. A parameter (e.g., “tunable”) transformation is then applied to the frequency representation of the beamforming energy at each location to generate a plurality of indication values, each of which is associated with a corresponding predetermined region. The parameter transformation is configured with parameter values ​​determined based on the performance of region detection for the environment. Whether a sound source is in the region is determined based on the indication value associated with a particular region.

[0004] Each aspect may include a combination of one or more of the following features.

[0005] Each location in the environment has a corresponding fixed physical location (e.g., three-dimensional coordinates) within the environment. In some examples, at least some of the locations have a corresponding orientation (i.e., the direction of source emission) within the environment.

[0006] Applying parameter transformation to the frequency representation of beamformer energy includes, for a region having at least one location located therein, combining the frequency representations of the beamformer energy at said locations.

[0007] Processing multiple microphone signals includes processing microphone signals over consecutive time periods, and repeatedly determining that a sound source is located in a specific area for multiple said time periods.

[0008] Determining the phase and amplitude relationship between microphone signal pairs involves determining the relationship based on microphone signals from a time period and one or more previous time periods (e.g., by averaging over multiple segments).

[0009] Determining the parameter values ​​for the parameterized beamforming process includes determining the transfer function for each microphone signal from each of multiple locations to multiple microphone signals.

[0010] The parameter values ​​for the parametric beamforming process are determined before receiving signals from multiple microphones. In some examples, the determination of the parameter values ​​for the parametric beamforming process is based at least in part on an analytical acoustic model of the environment. Alternatively or additionally, the determination of the parameter values ​​for the parametric beamforming process is based on experimental data collection in an environment with physical characteristics corresponding to the environment in which the sound source is detected (e.g., an environment with the same physical dimensions and microphone placement).

[0011] Determining the values ​​of the parameters for the parameterized beamforming process includes calculating a steering value between microphone signal pairs for each of a plurality of locations. In some instances, calculating the steering value between microphone signal pairs includes calculating the steering value for a subset of fewer than all microphone signal pairs (e.g., ignoring certain microphone pairs for one or more locations). In some instances, calculating the steering value between microphone signal pairs includes normalizing the magnitude of the steering value.

[0012] The method also includes determining, at least in part, the values ​​of parameters of a parameterized transformation of the frequency representation of beamformer energy for each location based on data including examples of microphone signals associated with a corresponding known source region collected in an environment having physical characteristics corresponding to the environment in which the sound source was detected.

[0013] The parameter values ​​for the parametric transformation are determined before receiving signals from multiple microphones.

[0014] The method also includes using an adaptive process (e.g., recursive least squares RLS) to determine the values ​​of parameters of a parameterized transform of the frequency representation of beamformer energy at each location using microphone signals previously collected in the environment where the sound source was detected.

[0015] The acoustic environment includes the vehicle's cabin and may also include at least one area outside the vehicle.

[0016] On the other hand, acoustic processing systems are typically configured to perform all the steps of any of the methods described above.

[0017] On the other hand, typically, non-transitory machine-readable media include instructions stored thereon, wherein executing the instructions on a data processor causes the processor to perform all the steps of any of the methods described above.

[0018] The advantage of transforming the frequency representation of tuned beamformer energy is that it adapts to acoustic variations in the environment (e.g., physical dimensions, surface reflections, reverberation, noise), which can improve the performance of area detection. Another advantage is that tuning allows for the consideration of a limited number of locations within each area, rather than having to determine the beamformer energy over the entire area of ​​the environment, thus reducing computational requirements.

[0019] Other features and advantages of the invention will be apparent from the following description and claims. Attached Figure Description

[0020] Figure 1 This is a schematic side view of a vehicle equipped with multiple microphones for monitoring the in-vehicle environment.

[0021] Figure 2 This is a schematic diagram of the vehicle's acoustic processing system;

[0022] Figure 3 This is a schematic diagram of a loudspeaker area detector; Detailed Implementation

[0023] refer to Figure 1 The vehicle is equipped with multiple microphones 132, 133 located within the vehicle environment 130 (e.g., the cabin, optionally including microphones outside the vehicle, not shown) that receive acoustic signals including speech generated by the occupant 120. The microphones provide signals to an audio system 136 that can transmit these signals to various audio-based or voice-based systems within (or away from) the vehicle, such as a voice interface 140 for voice assistance. Other systems using the microphone signals may include communication systems that allow the occupant to interact with remote persons via voice (e.g., via telephone), and in-vehicle communication systems that can enhance voice communication between occupants, such as using a speaker 133. The audio system 136 also provides the microphone signals to a speaker area detector 100, which determines the area (if any) within the vehicle where the passenger (hereinafter also referred to as the “user”) is speaking.

[0024] The speaker area detector 100 is used to determine the source area of ​​the user's voice in the vehicle environment 130. Such determination can then be used to control the process of audio interaction with the user, for example, to trigger an interaction when voice is detected, to control the directionality of audio signal acquisition and / or audio signal presentation for audio interaction, or to control the dialogue based on the speaker's location, for example, if multiple people can speak in the vehicle or the nature of the audio interaction can depend on the speaker's location (e.g., driver's position vs. passenger's position).

[0025] refer to Figure 2 One aspect of this system is that there is typically a finite set of predetermined source regions, for example, regions 122 labeled Z1 to Z5, which correspond to five seating positions in the cabin. In the following discussion, there are Z such predetermined source regions, as in... Figure 2 In the diagram, Z = 5. In at least some embodiments, the goal of the system is to determine which of these regions (if any) the user's voice originates from.

[0026] Another aspect of this system is that it can typically be configured for different vehicle environments, different microphone arrangements, and / or different area specifications within the environment (e.g., in size, location, typical speaker orientation, etc.). Figure 1 The image shows a total of nine microphones, 132 and 133. Figure 1 In the illustration, three microphones 132 have corresponding directional response patterns 134 (e.g., general omnidirectional patterns), and these microphones provide electrical microphone signals 151 representing the received acoustic signals. Optionally, some microphones undergo fixed array processing. For example, an array processor 138 can be used to process three microphones 133 (e.g., three microphones near the driver's position Z1) to generate two audio signals 152. For example, two different delay-sum beamformers are used to form two directional response patterns 135 corresponding to each audio signal 152, respectively. The microphone signals 151 and the processed audio signals 152 are considered as fixed acoustic sensors of the vehicle environment and are processed together as inputs to a manipulated response power processor 112. In the following discussion, these microphone and audio signals are referred to as x1 to x2. M ,exist Figure 1 In this arrangement, M=7. That is, the set of physical microphones generates a set of M microphone signals, each microphone signal being associated with a specific spatial response pattern in the vehicle.

[0027] The Speaker Region Detector 100 addresses the task of determining the region from which a speaker is speaking. The method described below utilizes one or more representative locations (i.e., a predetermined finite enumeration set) within the corresponding region (i.e., associated with the corresponding region). In the following description, there are P such locations. Typically, each location is associated with a specific physical location of the speaker in the environment, represented, for example, by three (or two) coordinates in the physical dimension. Each of these locations is represented by a location index 1 ≤ p ≤ P. Figure 2 In the diagram, P=14 indicates position 123, where each region has two or three positions.

[0028] Optionally, in addition to its physical location, each of the listed locations is also associated with one or more directions or orientations. When explicitly mentioned below, this direction is represented by a direction index d from a set of representative directions D. Furthermore, the reference to “location” below can therefore be optionally replaced below by a reference to “attitude” (p, d), representing both the physical location of the source and the direction of the source. That is, there exists an attitude P × D, not a location P.

[0029] For specific locations p and m in vehicle environment 130 th microphone signal x m 151, 152, the source signal s originating at this location to the microphone signal x m The acoustic and signal processing path can be represented (e.g., approximated) as a linear transfer function g. m In the frequency domain, for frequency K, k th The frequency index, the effect of the acoustic path, can be expressed as x. m (k)=g m (k,p)s(k), for example, can be computed via Fourier transform (e.g., Fast Fourier Transform, FFT). That is, for each combination of source location and microphone signal, there exists a different transfer function. Note that these transfer functions typically represent the gain and phase (which can be represented as complex quantities) at each frequency k, where both amplitude and phase are important in localization because differences in sound propagation distance can be reflected in both amplitude and phase differences. Various ways exist to estimate or approximate these transfer functions, including through experimental data collection (e.g., using physical acoustic environments, anatomically accurate simulations to transmit recorded speech, etc.) or through analytical techniques for acoustic signal propagation. In experimental and analytical techniques, it should be recognized that these determined transfer functions may be imprecise and may not match the effects of different numbers of passengers in a vehicle, reverberation effects, non-ideal user seating positions, etc. However, in the embodiments described below, it is assumed that these M×P transfer functions g m (k,p) is known and fixed for a specific vehicle configuration.

[0030] refer to Figure 3 The speaker region detector 100 receives M microphone signals 151, 152 and generates Z indicators (the number of regions) of the presence of speech in the corresponding regions, for example, representing a probability distribution over possible regions, or Z individual indicator probabilities of speech emanating from each region. In the embodiment described below, the region detector 100 includes two processing stages: a controlled response power (SRP) stage 112 and a tuning transformation stage 114. As described below, the controlled response power stage 112 manipulates the response power according to the transfer function g. m (k,p) is configured, and the tuning transform 114 is tuned based on experimental or synthesized speech data in the vehicle environment.

[0031] The following uses mathematical notation to specify microphone processing. While this processing can be performed digitally after analog-to-digital conversion (e.g., on a digital processor), other implementations can be used to perform the operation in analog form and / or in a mathematically approximate form. Typically, italic lowercase and uppercase letters represent scalar (potentially complex) quantities, such as a or A. Superscript * This indicates complex conjugate. Vectors are usually represented as columns and indicated by underscores, for example... a or A Finally, matrices are typically represented in uppercase, non-italic, for example, as A or Φ. Superscript T Represents the transpose of a vector or matrix, and the superscript... H This represents the transpose of the complex conjugate.

[0032] Figure 1 The SRP stage 112 shown receives M microphone signals 151 and 152 and outputs P SRP signals 161. The SRP stage can be implemented in the frequency domain, where the microphone time signal x m (t) is processed in consecutive time segments (e.g., "windows" or "frames"), and each segment is processed to produce a Fourier transform x. m (k), for k = 0, ..., K-1, is a complex function of frequency k (i.e., representing amplitude and phase). The following processing describes a time interval of such a process, which can be understood as repeating the process over consecutive time intervals. For each frequency index k, the microphone signal can be arranged as a vector. x (k)=[x1(k),x2(k),...,x M (k)] T .

[0033] For each frequency index, the spatial correlation (M×M complex) matrix for that time portion is calculated as follows: x (k) x H (k), and the spatially correlated smoothing estimate is updated to Where α∈[0,1) is chosen to adjust the system to smoothly estimate how fast the update is. The current smooth estimate is called as follows Φ xx And (m,n) entries are represented

[0034] An optional step is to perform element-wise smoothing estimation of Φ for all 1≤m, n≤M. xx Some entries in the database are cleared (i.e., set to zero). Where ψ mn ∈{0,1}. This mask can be represented in matrix notation as the smooth spatial correlation Φ of the groupings. xx,subset (k)=Φ xx (k)°Ψ. As discussed further below, one reason for setting certain terms to zero is that, for example, due to the distance between microphones, the corresponding microphone signal pairs may not provide useful information about the relative phase and / or amplitude, and the system may be more robust in these locations with zero rather than small random quantities.

[0035] The next step is to scale the smoothed spatial correlations of the groups for each frequency index k to produce a correlation matrix whose square root (which can be called the matrix's "Frobenius norm") is 1. This operation can be represented as the calculation of the norm as follows:

[0036]

[0037] Then perform the division:

[0038]

[0039] For all m and n, it can be represented by an element-wise matrix partition as follows:

[0040]

[0041] This term is called normalized smoothed spatial correlation.

[0042] After calculating the normalized spatial correlation matrix for each frequency index (i.e., for each pair of microphone signals), the next step is to calculate the "beamformer" energy (i.e., the squared amplitude of the complex-valued beamformer signal) at each frequency index for each position (or more generally, each orientation as a position-orientation pair). This step utilizes a pre-computed set of beamformer weights, which includes a complex w for each microphone m, position p, and frequency k combination. m (k,p). The pre-calculation of these weights is presented later in this paper. These quantities can be arranged in vector notation as follows: w (k,p)=[w1(k,p),w2(k,p),...,w M(k,p)] T For each position p, the manipulation matrix (i.e., the set of manipulation values ​​for several pairs of microphones) can be defined as W(k,p) = w (k,p) w H (k,p). The manipulation matrix is ​​grouped and magnitude normalized using the same sequence of operations applied to the spatial correlation matrix as described above:

[0043] as well as

[0044]

[0045] Finally, the beamformer energy at each position p and frequency k is calculated as follows:

[0046]

[0047] In the first method of area detection, the frequency- and location-dependent beamformer energy Φyy(k,p) is frequency-weighted and summed to produce the beamformer value of the frequency integral:

[0048]

[0049] Where κ(k)∈[0,1] is the real-valued weighting of the frequency index k. For example, the weighting can be zero at frequencies where no speech energy is found or very little speech energy is found and therefore little directional information can be extracted.

[0050] As mentioned above, there are typically more locations than regions, and each location is associated with a specific region. This association (e.g., assignment) is represented as z = ζ(p), where z is the region of location p. (Note that when azimuth d is also used, the assignment function is the same for all azimuths at a given location). Each region z has a predetermined function P applied to some or all frequency-weighted beamformer values. z And the relevant region assignment for all locations to produce the corresponding value ξ z One use of these values ​​is if ξ z If the threshold is exceeded, it is declared that there is speech activity in region z.

[0051] These regional functions P z It can take various forms. For all p, one option is as a region-specific linear combination of all GSRP(p) terms. In this case, for all p, all P =[P1,...,P Z ] TAll of these can be represented by a Z×P matrix, which combines all the P·D values ​​of GSRP(p) stacked in a vector into a Z-region output value.

[0052] Another option is to use a function of GSRP(p) for all p, where ζ(p) = z, that is, a function of all GSRP samples assigned to region z. An example would be to take the maximum GSRP value for all sampling source locations (and orientations) assigned to region z:

[0053]

[0054] There are other options for the function, such as those based on multiple logistic regression or feedforward neural networks (e.g., multilayer perceptrons).

[0055] In the second method, the integration over the frequency is delayed, and a region- and frequency-specific (i.e., "narrowband") function P is performed at each frequency in a manner similar to that described above. z (k), generating the corresponding value ξ z (k). Therefore, for each frequency slot k, there exist Z tuning functions P1(k), P2(k), ..., P Z (k). Then, the value ξ is weighted again according to the weight κ(k) in terms of frequency. z (k) Integrate:

[0056] Everything, for all of z.

[0057] There are at least two pre-computation stages. The first stage involves calculating the beamformer weighting w described above. The second stage involves calculating the function P described above. z This second stage can be referred to as the "tuning" stage. The tuning stage may optionally further include determining the relevant mask Ψ and / or the frequency weighting κ.

[0058] One method in the first phase utilizes the frequency k described earlier in this patent. th From the source at position (or orientation) pth to m th microphone signal transfer function g m (k,p). For a specific frequency and location, the vector of gain. g (k,p)=[g1(k,p),g2(k,p),...,g M (k,p)] T The beamformer weights are defined and pre-calculated in this method as follows:

[0059]

[0060] Where || represents the Euclidean (L2) norm.

[0061] In the second method, the first stage also utilizes a noise or reverberation model, for example, estimating m using the M×M×K complex noise power spectral density (PSD) for all m,n. th and n th Noise signal v in microphone signal m (k) and v n (k) The complex noise power spectral density (PSD) is composed of an M×M noise PSD matrix Φ for each k. vv (k) represents this. In this method, the beamformer weights can be calculated as:

[0062]

[0063] In a variation of the second approach, noise can be monitored during runtime, for example, when people are not speaking in the car, and / or reverberation can be monitored during their speech. In such a variation, the noise PSD matrix is ​​estimated during runtime, and the beamformer weights are updated accordingly.

[0064] The second stage includes estimating the function P. z The parameter values. The set of parameters to be estimated is denoted as θ. P The values ​​of these parameters are determined using indicators. d =(d1,d2,...,d Z The beamformer value Φyy(k,p) is determined from training data and its corresponding frequency and location-dependent beamformer values. The elements of this indicator take values ​​of 0.0 or 1.0 depending on the actual location of the source. These beamformer values ​​are based on environmental measurements or physical simulations (e.g., considering near-field and far-field effects, head shadows, etc.). For a specific parameter value θ... P The function P based on the beamformer value Φyy(k,p) z Generate output And by changing parameter values ​​to minimize some cost (e.g., loss) functions. For example, this cost could be defined as the mean square cost as follows:

[0065]

[0066] In some of the examples described above, the function P z Since the function is linear and the loss function is quadratic, a closed-form expression can be used to determine the solution for the optimal parameter values. Other methods can be used with different functional forms and different loss functions; for example, an iterative update process based on gradient descent can be used to determine the parameter values. This update can be performed independently for each frequency, or joint optimization can be used.

[0067] While the tuning method described above is within the context of data marked with indicators of the speaker's true region, similar approaches can be used to adapt tuning based on speech received at runtime, such as in decision-oriented methods. The "truth value" of the indicators can be based on runtime detection methods, possibly supplemented with other information such as dialogue state, or using a time zone variation model. Similarly, parameter optimization can be performed using real speech recordings from the car, incorporating new recordings into the optimization using a recursive least squares (RLS) algorithm.

[0068] As mentioned above, there are other parameters that can be adjusted, namely the correlation mask Ψ and / or the frequency weighting κ. One way to optimize the frequency weights is to include them in a gradient-based iterative update method. The correlation mask can be estimated based on various methods. For example, a threshold W(k,p) for the entries can be estimated, and the mask can be set based on whether the entries exceed that threshold. Other combined optimization methods can be used to select the entries to rely on; for example, iterative methods utilizing genetic algorithms can be used.

[0069] While the above description is primarily within the context of area detection in a vehicle cabin, it should be understood that areas and / or microphones outside the vehicle can be used to allow for the detection of speakers outside the vehicle, which may rely on both internal and external microphones. Furthermore, the method is not limited to vehicle applications and can be used for area detection in, for example, buildings, such as residential rooms with voice assistants or office meeting rooms where areas can be associated with video conference participants.

[0070] The above methods can be implemented using hardware, software, or a combination of both. Such hardware may include application-specific integrated circuits (ASICs) and field-programmable gate arrays (FPGAs). The software may include instructions stored on a non-transitory machine-readable medium, and the system may include a processor (e.g., a physical processor with a central processing unit) that executes the instructions to perform the methods.

[0071] Several embodiments of the invention have been described. However, it should be understood that the foregoing description is intended to illustrate, not limit, the scope of the invention, which is defined by the appended claims. Therefore, other embodiments are also within the scope of the appended claims. For example, various modifications can be made without departing from the scope of the invention. Furthermore, some of the steps described above may be sequential and therefore may be performed in a different order than that described.

Claims

1. A method for acoustic processing in an environment, wherein a plurality of microphones in the environment are in a fixed configuration for providing a plurality of microphone signals for sensing the environment, the environment having a fixed predetermined plurality of locations and a plurality of predetermined regions, wherein each region has one or more of the locations therein, wherein the acoustic processing includes detecting sound sources in the predetermined regions of the environment, the detection of the sound sources including: Receive signals from the plurality of microphones; Process the plurality of microphone signals to determine at least one of the phase and amplitude relationships between pairs of microphone signals at a plurality of frequencies; The relationship between microphone pairs is processed according to a parameterized beamforming process to generate a frequency representation of the beamforming energy for each of the plurality of locations in the environment. as well as The location of the sound source in a specific region is determined based on the beamforming energy at each of the locations.

2. The method of claim 1, wherein the acoustic environment comprises the vehicle cabin.

3. The method of claim 2, wherein the acoustic environment includes at least one area outside the vehicle.

4. The method of claim 1, further comprising applying a parameter transformation to a frequency representation of beamformer energy for each location to generate a plurality of indication values, each indication value being associated with a corresponding predetermined region, the parameter transformation being configured with values ​​of parameters determined based on the performance of region detection for the environment, and wherein determining that the sound source is in a particular region is based on the indication value associated with the region.

5. The method according to any one of claims 1 to 4, wherein, Each of the locations in the environment has a corresponding fixed physical location in the environment.

6. The method according to claim 5, wherein, Each of at least some of the locations has a corresponding orientation in the environment.

7. The method according to claim 4, wherein, Applying the parameter transformation to the frequency representation of beamformer energy includes, for at least one region having multiple locations therein, combining the frequency representation of beamformer energy for said location.

8. The method of claim 4, further comprising determining, at least in part, the values ​​of parameters of a parameterized transformation of the frequency representation of beamformer energy for each location based on data including examples of microphone signals, the examples of which are associated with a corresponding known source region collected in an environment having physical characteristics corresponding to the environment in which the sound source was detected.

9. The method according to claim 8, wherein, The values ​​of the parameters for performing parametric transformations are determined before receiving the signals from the plurality of microphones.

10. The method of claim 4, further comprising using an adaptive process to determine the values ​​of parameters of a parameterized transformation of the frequency representation of beamformer energy for each location using microphone signals previously collected in the environment where the sound source was detected.

11. The method according to any one of claims 1 to 4, wherein, Processing the plurality of microphone signals includes processing the microphone signals over consecutive time periods, and repeatedly determining that the sound source is in a specific region for the plurality of time periods.

12. The method according to any one of claims 1 to 4, wherein, Determining the phase and amplitude relationship between the microphone signal pairs includes determining the relationship between the time periods based on microphone signals from a time period and one or more previous time periods.

13. The method according to any one of claims 1 to 4, further comprising determining the values ​​of parameters of the parameterized beamforming process, including determining the transfer function from each of the plurality of locations to each of the plurality of microphone signals.

14. The method according to claim 13, wherein, The values ​​of the parameters for the parameterized beamforming process are determined before receiving the signals from the plurality of microphones.

15. The method of claim 13, wherein the determination of the values ​​of the parameters of the parameterized beamforming process is based at least in part on an analytical acoustic model of the environment.

16. The method according to claim 13, wherein, The values ​​of the parameters in the parameterized beamforming process are determined based on experimental data collection in an environment with physical characteristics corresponding to the environment in which the sound source was detected.

17. The method of claim 13, wherein determining the values ​​of the parameters of the parameterized beamforming process includes calculating manipulation values ​​between microphone signal pairs for each of the plurality of locations.

18. The method of claim 17, wherein calculating the manipulation value between microphone signal pairs comprises calculating the manipulation value for a subset of fewer than all of the microphone signal pairs.

19. The method of claim 17, wherein calculating the manipulation value between microphone signal pairs includes normalizing the manipulation value.

20. An acoustic processing system configured to perform all the steps of any one of claims 1 to 19.

21. A non-transitory machine-readable medium comprising instructions stored thereon, wherein Executing the instructions on the data processor causes the processor to perform all the steps of any one of claims 1 to 19.