Spatial audio-filtering in spatial audio capture

By estimating two direction-of-arrival parameters and their energy ratios for each frequency band, the method improves spatial audio capture by accurately filtering multiple sound sources, enhancing audio quality and user experience in complex environments.

JP2025148391AInactive Publication Date: 2025-10-07NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025113134
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-04
Filing Date
2025-07-03
Publication Date
2025-10-07
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing spatial audio capture methods struggle to accurately amplify and attenuate multiple sound sources in complex audio environments, leading to interference and reduced audio quality due to insufficient direction-of-arrival (DOA) estimation and energy ratio analysis.

Method used

Implement a method that estimates two direction-of-arrival parameters and their energy ratios for each frequency band, using multiple microphones to generate filters that enhance spatial audio filtering by combining these estimates, improving accuracy and reducing interference.

Benefits of technology

Enhances the perceived audio zoom effect and audio quality by accurately amplifying desired sound directions and attenuating others, even in complex environments with multiple sound sources, thereby improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025148391000001_ABST
    Figure 2025148391000001_ABST
Patent Text Reader

Abstract

To provide a device and a method for spatial audio-filtering in a spatial audio capture.SOLUTION: A device includes means configured to acquire a plurality of audio signals from a plurality of microphones, determine a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the plurality of audio signals on the basis of processing of the plurality of audio signals, determine a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the plurality of audio signals on the basis of processing of the plurality of audio signals, acquire an area for defining a direction and / or a range for a filter, and generate a filter to be applied to a plurality of audio signals. The device generates a filter gain / attenuation parameter on the basis of an area associated with the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application relates to an apparatus and method for spatial audio filtering within spatial audio capture. [Background technology]

[0002] Spatial audio capture using microphone arrays is often used in conjunction with video capture in many modern digital devices, such as mobile devices and cameras, and can be played back using headphones or loudspeakers to provide a user with the experience of the audio scene captured by the microphone array.

[0003] Parametric spatial audio capture methods enable spatial audio capture using a variety of microphone configurations and configurations and can therefore be used in consumer devices such as mobile phones. Parametric spatial audio capture methods are based on signal processing solutions that utilize available information from multiple microphones to analyze the spatial audio field around the device. Typically, these methods perceptually analyze microphone audio signals to determine relevant information within frequency bands. This information includes, for example, the direction of the dominant sound source (or sound source or audio object) and the relationship between source energy and overall band energy. Based on this determined information, spatial audio can be reproduced using, for example, headphones or loudspeakers. Ultimately, therefore, a user or listener can experience environmental audio as if they were present in the audio scene being recorded by the capture device.

[0004] The better the audio analysis and synthesis performance, the more realistic the results that the user or listener will experience. Summary of the Invention

[0005] According to a first aspect, there is provided an apparatus comprising means configured to perform the steps of: acquiring a plurality of audio signals from respective microphones; determining a first source direction parameter and a first source energy parameter in one or more frequency bands of the plurality of audio signals based on processing the plurality of audio signals; determining a second source direction parameter and a second source energy parameter in one or more frequency bands of the plurality of audio signals based on processing the plurality of audio signals; obtaining a region defining a direction and / or a range for a filter; and generating the filter to be applied to the plurality of audio signals, wherein filter gain / attenuation parameters are generated based on the region for the first source direction parameter, the first source energy parameter, the second source direction parameter, and the second source energy parameter.

[0006] The means may be configured to generate a filter to be applied to a plurality of audio signals, wherein the filter gain / attenuation parameters are generated based on a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a region associated with the second sound source energy parameter; generate a first band gain / attenuation value based on the first sound source direction parameter being within or outside the region; generate a second band gain / attenuation value based on the second sound source direction parameter being within or outside the region; and combine the first band gain / attenuation value and the second band gain / attenuation value to generate a composite band gain / attenuation value.

[0007] The means configured to obtain a region defining a direction and / or range for the filter may be configured to obtain at least one of a direction and range defining the region, an in-band gain / attenuation coefficient based on the sound source direction parameter being within the region, an out-of-band gain / attenuation coefficient based on the sound source direction parameter being outside the region, a direction and range defining the region together with an in-band gain / attenuation coefficient based on the sound source direction parameter being within the region, an edge zone band gain / attenuation coefficient based on the sound source direction parameter being within the edge zone region, an out-of-band gain / attenuation coefficient based on the sound source direction parameter being outside the region, and a further range defining the edge zone region.

[0008] The means configured to generate the filters to be applied to the plurality of audio signals, wherein filter gain / attenuation parameters are generated based on the region in relation to the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter, may be configured to: generate a first temporal gain / attenuation value based on a time average of the mean band value of the first sound source energy parameter, generate a second temporal gain / attenuation value based on the number of times the first sound source direction parameter is within the region over a defined time period, generate a second temporal gain / attenuation value based on the time average of the mean band value of the second sound source direction parameter and the number of times the second sound source direction parameter is within the region within a defined time period, and generate a synthesized temporal gain / attenuation value based on a combination of the first temporal gain / attenuation value and the second time gain / attenuation value to generate a synthesized temporal gain / attenuation value.

[0009] means configured to generate a filter to be applied to a plurality of audio signals, said means comprising: The filter gain / attenuation parameters may be generated based on the first source direction parameter, the first source energy parameter, the second source direction parameter, and a region associated with the second source energy parameter, and configured to generate a composite frame average based on a combination of the frame-averaged first source energy parameter and the frame-averaged second source energy parameter. An apparatus is provided that includes means configured to perform generating frame-smoothed gain / attenuation based on the frame average and the number of times the first and second source direction parameters are within the filter region during a frame period.

[0010] means configured to generate a filter to be applied to a plurality of audio signals, wherein the filter gain / attenuation parameters are generated based on a first sound source direction parameter, a region associated with the first sound source energy parameter; the means further comprising: a second sound source direction parameter; and The second sound source energy parameter may be configured to generate a filter gain / attenuation based on a combination of the frame smoothing gain / attenuation, the composite temporal gain / attenuation value, and the composite band gain / attenuation value.

[0011] The processing of the plurality of audio signals may be configured to provide one or more modified audio signals based on the plurality of audio signals, and the means configured to determine a second sound source direction parameter and a second sound source energy parameter based on the processing of the plurality of audio signals in one or more frequency bands of the plurality of audio signals may be configured to determine the second sound source direction parameter and the second sound source energy parameter based on the modified audio signal in one or more frequency bands of the plurality of audio signals.

[0012] The means configured to provide one or more modified audio signals based on the plurality of audio signals may be further configured to generate the modified plurality of audio signals based on modifying the plurality of audio signals with a projection of the first sound source defined by the first sound source direction parameter, and to determine at least a second sound source direction parameter in one or more frequency bands of the plurality of audio signals based at least in part on at least a portion of the one or more modified audio signals, wherein the means configured to determine at least the second sound source direction parameter by processing the modified plurality of audio signals in the one or more frequency bands of the plurality of audio signals.

[0013] The means configured to obtain a region defining the direction and / or extent of the filter comprises: The area may be obtained based on user input.

[0014] According to a second aspect, there is provided a method for an apparatus, the method comprising the steps of obtaining a plurality of audio signals from respective microphones; determining a first source direction parameter and a first source energy parameter in one or more frequency bands of the plurality of audio signals based on processing the plurality of audio signals; determining a second source direction parameter and a second source energy parameter in one or more frequency bands of the plurality of audio signals based on processing the plurality of audio signals; obtaining a region defining a direction and / or a range for a filter; and generating a filter to be applied to the plurality of audio signals, wherein filter gain / attenuation parameters are generated based on a region relating to the first source direction parameter, the first source energy parameter, the second source direction parameter, and the second source energy parameter.

[0015] The step of generating a filter to be applied to the plurality of audio signals, wherein filter gain / attenuation parameters are generated based on a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a region associated with the second sound source energy parameter, can include generating a first band gain / attenuation value based on the first sound source direction parameter being within or outside the region, generating a second band gain / attenuation value based on the second sound source direction parameter being within or outside the region, and combining the first band gain / attenuation value and the second band gain / attenuation value to generate a combined band gain / attenuation value.

[0016] The step of obtaining a region defining a direction and / or range for the filter may include at least one of a direction and range defining the region, an in-band gain / attenuation coefficient based on the sound source direction parameter being within the region, and an out-of-band gain / attenuation coefficient based on the sound source direction parameter being outside the region, a direction and range defining the region, the in-band gain / attenuation coefficient based on the sound source direction parameter being within the region, the in-band gain / attenuation coefficient based on the sound source direction parameter being within the region, and the out-of-band gain / attenuation coefficient based on the sound source direction parameter being outside the region, and a further range defining an edge zone region, the out-of-band gain / attenuation coefficient based on the sound source direction parameter being within the edge zone region.

[0017] The step of generating a filter to be applied to a plurality of audio signals, wherein filter gain / attenuation parameters are generated based on a region associated with a first sound source direction parameter, a first sound source direction parameter, a second sound source direction parameter, and a second sound source energy parameter, may include generating a first temporal gain / attenuation value based on a time average of an average band value of the first sound source energy parameter; generating the first sound source direction parameter based on the time average of the average band value of the first sound source energy parameter; generating a number of times the second sound source direction parameter is within a region over a defined time period based on a time average of the time-averaged band value of the second sound source direction parameter and a number of times the second sound source direction parameter is within the region within a defined time period; and generating a composite temporal gain / attenuation value based on a combination of the first temporal gain / attenuation value and the second temporal gain / attenuation value to generate a composite temporal gain / attenuation value.

[0018] The step of generating a filter to be applied to the plurality of audio signals may include generating a filter gain / attenuation parameter based on a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a region associated with the second sound source energy parameter, the filter gain / attenuation parameter may include generating a synthesized frame average value based on a combination of the frame-averaged first sound source energy parameter and the frame-averaged second sound source energy parameter, and generating a frame-smoothed gain / attenuation based on the synthesized frame average value and the number of times the first and second sound source direction parameters fall within the filter region over a frame period.

[0019] Generating filters to be applied to the plurality of audio signals, wherein filter gain / attenuation parameters are generated based on a first source direction parameter, a first source energy parameter, a second source direction parameter, and a region associated with the second source energy parameter, can include generating filter gains / attenuation for bands based on a combination of frame smoothing gains / attenuation, synthesized time gains / attenuation values, and synthesized band gains / attenuation values.

[0020] Processing the plurality of audio signals may include providing one or more modified audio signals based on the plurality of audio signals, and determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the plurality of audio signals based on the processing of the plurality of audio signals may include determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the plurality of audio signals based on the modified audio signals.

[0021] Providing one or more modified audio signals based on the plurality of audio signals may include generating the modified plurality of audio signals based on modifying the plurality of audio signals with a projection of a first sound source defined by a first sound source direction parameter. Determining at least a second sound source direction parameter in one or more frequency bands of the plurality of audio signals based at least in part on the one or more modified audio signals may include determining at least a second sound source direction parameter in the one or more frequency bands of the plurality of audio signals by processing the modified plurality of audio signals.

[0022] Obtaining a region that defines the direction and / or extent of the filter may include obtaining the region based on user input.

[0023] According to a third aspect, there is provided an apparatus comprising at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code causing, using the at least one processor, the apparatus to at least obtain a plurality of audio signals from respective ones of a plurality of microphones, determine a first sound source direction parameter and a first sound source energy parameter based on processing the plurality of audio signals in one or more frequency bands of the plurality of audio signals, determine a second sound source direction parameter and a second sound source energy parameter based on processing the plurality of audio signals, obtain a region defining a direction and / or range for a filter, and generate a filter to be applied to the plurality of audio signals, wherein filter gain / attenuation parameters are generated based on the region relating to the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter.

[0024] An apparatus for generating filters to be applied to multiple audio signals, wherein filter gain / attenuation parameters are generated based on a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a region associated with the second sound source energy parameter, the apparatus is configured to generate a first band gain / attenuation value based on the first sound source direction parameter being within or outside the region, generate a second band gain / attenuation value based on the second sound source direction parameter being within or outside the region, and combine the first band gain / attenuation value and the second band gain / attenuation value to generate a combined band gain / attenuation value.

[0025] The device for obtaining a region defining a direction and / or range of a filter may obtain at least one of: a direction and range defining a region having in-band gain / attenuation coefficients based on the sound source direction parameter being within the region; a direction and range defining a region having out-of-band gain / attenuation coefficients based on the sound source direction parameter being outside the region and in-band gain / attenuation coefficients based on the sound source direction parameter being within the region; and a further range defining an edge zone region having out-of-band gain / attenuation coefficients based on the sound source direction parameter being outside the region and an edge zone band gain / attenuation coefficient based on the sound source direction parameter being within the edge zone region.

[0026] An apparatus caused to generate filters to be applied to multiple audio signals, wherein filter gain / attenuation parameters are generated based on a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a region associated with the second sound source energy parameter. The apparatus may perform the following steps: generating a first temporal gain / attenuation value based on a time average of mean band values ​​of the first sound source energy parameter, wherein the number of times the first sound source direction parameter occurs in the region is the number of times it occurs in a defined time; generating a second temporal gain / attenuation value based on a time average of mean band values ​​of a second sound source energy parameter, wherein the number of times the second sound source direction parameter occurs in the region is the number of times it occurs in the defined time; and generating a composite temporal gain / attenuation value based on a combination of the first temporal gain / attenuation value and the second temporal gain / attenuation value to generate a composite temporal gain / attenuation value.

[0027] An apparatus for generating a filter to be applied to multiple audio signals, wherein filter gain / attenuation parameters are generated based on a first source direction parameter, a first source energy parameter, a second source direction parameter, and a region for the second source energy parameter, the apparatus may perform the steps of generating a synthesized frame average value based on a combination of the frame-averaged first source energy parameter and the frame-averaged second source energy parameter, and generating a frame-smoothed gain / attenuation value based on the synthesized frame average value and the number of times the first and second source direction parameters fall within the filter region over a frame period.

[0028] An apparatus for generating filters to be applied to multiple audio signals, wherein filter gain / attenuation parameters are generated based on a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a region related to the second sound source energy parameter, the apparatus may perform the step of generating filter gain / attenuation for bands based on a combination of frame smoothing gain / attenuation, a synthesis time gain / attenuation value, and a synthesis band gain / attenuation value.

[0029] The processing of the multiple audio signals may be configured to provide one or more modified audio signals based on the multiple audio signals, and the apparatus for determining a second sound source direction parameter and a second sound source energy parameter based on the processing of the multiple audio signals, in one or more frequency bands of the multiple audio signals, may determine the second sound source direction parameter and the second sound source energy parameter based on the modified audio signal, in one or more frequency bands of the multiple audio signals.

[0030] The apparatus for providing one or more modified audio signals based on the plurality of audio signals may further perform the step of generating the modified plurality of audio signals based on modifying the plurality of audio signals with a projection of a first sound source defined by a first sound source direction parameter. The apparatus for determining at least a second sound source direction parameter in one or more frequency bands of the plurality of audio signals based at least in part on at least one modified audio signal determines at least the second sound source direction parameter in one or more frequency bands of the plurality of audio signals by processing the modified plurality of audio signals.

[0031] The device that obtains the region that defines the direction and / or extent of the filter may have the region obtained based on user input.

[0032] According to a fourth aspect, there is provided an apparatus comprising: means for acquiring a plurality of audio signals from respective microphones; means for determining a first source direction parameter and a first source energy parameter in one or more frequency bands of the plurality of audio signals based on processing the plurality of audio signals; means for determining a second source direction parameter and a second source energy parameter in one or more frequency bands of the plurality of audio signals based on processing the plurality of audio signals; means for obtaining a region defining a direction and / or a range for a filter; and means for generating the filter to be applied to the plurality of audio signals, wherein filter gain / attenuation parameters are determined based on the first source direction parameter, the first source energy parameter, The second sound source direction parameter is generated based on the region with respect to the second sound source energy parameter.

[0033] According to a fifth aspect, there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] to cause an apparatus to at least: acquire a plurality of audio signals from respective microphones; determine a first source direction parameter and a first source energy parameter in one or more frequency bands of the plurality of audio signals based on processing the plurality of audio signals; determine a second source direction parameter and a second source energy parameter in one or more frequency bands of the plurality of audio signals based on processing the plurality of audio signals; obtain a region defining a direction and / or range for a filter; and generate the filter to be applied to the plurality of audio signals, wherein: A filter gain / attenuation parameter is generated based on the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and a region associated with the second sound source energy parameter.

[0034] According to a sixth aspect, there is provided a non-transitory computer-readable medium comprising program instructions to cause an apparatus to at least: obtain multiple audio signals from respective multiple microphones; determine a first source direction parameter and a first source energy parameter in one or more frequency bands of the multiple audio signals based on processing the multiple audio signals; determine a second source direction parameter and a second source energy parameter in the one or more frequency bands of the multiple audio signals based on processing the multiple audio signals; obtain a region defining a direction and / or range for a filter; and generate a filter to be applied to the multiple audio signals, wherein filter gain / attenuation parameters are generated based on a region for the first source direction parameter, the first source energy parameter, the second source direction parameter, and the second source energy parameter.

[0035] According to a seventh aspect, there is provided an acquisition circuit configured to acquire a plurality of audio signals from a respective plurality of microphones, and in one or more frequency bands of the plurality of audio signals, An apparatus is provided, comprising: a determination circuit configured to determine a first source direction parameter and a first source energy parameter based on processing of the multiple audio signals, a determination circuit configured to determine a second source direction parameter and a second source energy parameter in one or more frequency bands of the multiple audio signals based on processing of the multiple audio signals, an acquisition circuit configured to acquire a region defining a direction and / or a range for a filter, and a generation circuit configured to generate the filter to be applied to the multiple audio signals, wherein filter gain / attenuation parameters are generated based on the region for the first source direction parameter, the first source energy parameter, the second source direction parameter, and the second source energy parameter.

[0036] According to an eighth aspect, there is provided a computer-readable medium comprising program instructions to cause an apparatus to at least: obtain multiple audio signals from respective multiple microphones; determine a first source direction parameter and a first source energy parameter in one or more frequency bands of the multiple audio signals based on processing the multiple audio signals; determine a second source direction parameter and a second source energy parameter based on processing the multiple audio signals; obtain a region defining a direction and / or range for a filter; and generate a filter to be applied to the multiple audio signals, wherein the filter gain / attenuation parameters are generated based on the region for the first source direction parameter, the first source energy parameter, the second source direction parameter, and the second source energy parameter.

[0037] The apparatus includes means for performing the operations described above.

[0038] The apparatus is configured to perform the method operations as described above.

[0039] The computer program product of the present application comprises program instructions for causing a computer to carry out the above-described method.

[0040] A computer program product stored on the medium can cause an apparatus to perform the methods described herein.

[0041] The electronic device may comprise an apparatus as described herein.

[0042] A chipset may include the devices described herein.

[0043] Embodiments of the present application aim to address problems associated with the state of the art. [Brief explanation of the drawings]

[0044] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings in which: [Figure 1] FIG. 1 illustrates a schematic diagram of an exemplary apparatus for implementing spatial capture and playback according to some embodiments. [Figure 2] FIG. 2 shows a flow diagram of the operation of the apparatus shown in FIG. 1 according to some embodiments. [Figure 3] FIG. 3 illustrates a schematic diagram of an exemplary spatial analyzer such as that shown in FIG. 1, according to some embodiments. [Figure 4] FIG. 4 illustrates a flow diagram of the operation of the exemplary spatial analyzer shown in FIG. 3, according to some embodiments. [Figure 5] FIG. 5 shows an example situation where a sound source is located inside or outside the zone of interest. [Figure 6] FIG. 6 shows an example graph of the signal level of the spatial filter. [Figure 7]FIG. 7 shows a flow diagram of a spatial filtering operation for determining that a sound source is within a zone of interest based on two sound source direction estimates, according to some embodiments. [Figure 8] FIG. 8 shows a flow diagram of spatial filtering based on two sound source direction estimates, according to some embodiments. [Figure 9] FIG. 9 illustrates a schematic diagram of an exemplary spatial synthesizer such as that shown in FIG. 2, according to some embodiments. [Figure 10] 10 and 11 show schematic diagrams of an exemplary system of devices suitable for implementing embodiments, including devices as shown in the previous figures. [Figure 11] 10 and 11 show schematic diagrams of an exemplary system of devices suitable for implementing embodiments, including devices as shown in the previous figures. [Figure 12] FIG. 12 shows a schematic diagram of an exemplary device suitable for implementing the apparatus shown. DETAILED DESCRIPTION OF THE INVENTION

[0045] The concepts described in further detail herein with respect to the following embodiments relate to capturing an audio scene. For example, the following embodiments can be implemented within a capture device configured to determine object / source-related audio signals. For example, in some embodiments, two source direction estimates for a sector / zone of interest and their associated direct ambient energy ratios can be used in determining a filter gain / attenuation amount to "filter" the object / source-related audio signals. This spatial filtering can be used instead of (or in addition to) conventional beamforming to generate object audio signals. While the following embodiments describe filter gain parameters, these same approaches can be used to generate filter attenuation parameters.

[0046] Additionally, the following embodiments may also be implemented within playback devices where the captured audio is processed by "zoom" or "focus." Additionally, spatial filtering may be performed as any part of the spatial audio signal synthesis operation.

[0047] In the following description, the term sound source is used to describe a defined element (artificial or real) within a sound field (or audio scene). The term sound source can also be defined as an audio object or an audio source, and these terms are interchangeable with regard to understanding the implementation of the examples described herein.

[0048] Embodiments herein relate to parametric audio capture devices and methods, such as spatial audio capture (SPAC) techniques. For each time-frequency tile, the device is configured to estimate the direction of the dominant sound source and the relative energy of the direct and ambient components of the sound source, expressed as a direct-to-total energy ratio.

[0049] The following examples are suitable for devices with challenging microphone configurations or configurations, such as those found in typical mobile devices, whose dimensions typically include at least one dimension that is shorter (or thinner) relative to other dimensions. In the examples shown herein, The captured spatial audio signal is a suitable input for a spatial synthesizer to generate a spatial audio signal, such as a binaural format audio signal for headphone listening, or a multi-channel signal format audio signal for loudspeaker listening.

[0050] In some embodiments, these examples may be implemented as part of a spatial capture front end for the Immersive Voice and Audio Services (IVAS) standard codec by generating IVAS-compatible audio signals and metadata.

[0051] An audio scene (spatial audio environment) can be complex and comprise several simultaneous audio or sound sources with different spectral characteristics. In addition, strong background noise can make it difficult to determine the direction of a sound source. This can cause problems when filtering audio technical fields (represented by the captured audio signal), which also means that sound elements within the audio technical field that would otherwise be filtered (or attenuated) from the audible sound field may leak into the processed output due to insufficient precision or reliability of the spatial audio analysis.

[0052] Furthermore, real-world audio recording conditions, such as simultaneous sound sources, echoes, and ambient sound environments, often make it difficult to amplify and / or attenuate desired sound directions with good audio quality. Typically, spatial audio capture methods determine only a single direction estimate per frequency band and pass it to a filter. Therefore, distinguishing between and amplifying / attenuating audio signal components associated with two simultaneous sound directions present within the same frequency band can be difficult or virtually impossible. Because the direction of at least one of the two simultaneous audio sources remains unknown, a further problem exists for so-called audio zoom or audio focusing algorithms, whose goal is to amplify audio signal components (sounds) arriving only from a specified direction and attenuate those from other directions. The "unknown" sound source direction may be located in or near the zoom direction, but cannot be amplified without proper DOA estimation. Correspondingly, efficient attenuation of other directions requires DOA estimates for both sound sources; otherwise, the algorithm may accidentally attenuate other sound sources in or near the zoom direction based on a single DOA estimate for a sound source located in a direction far from the zoom direction.

[0053] The embodiments described herein aim to improve how sound sources can be amplified and / or attenuated as requested by the user by implementing an improved two-way direction estimation method for each frequency band. The estimation method provides additional information about the audio environment and sound source direction for filtering. In other words, it provides two direction estimates and their direct surrounding energy ratios for each subband, enabling more efficient spatial filtering. The increased efficiency is based on combining calculated filtering gains corresponding to both (all) DOA estimates and their energy ratios. This in turn increases and strengthens the perceived audio zoom effect, allowing audio zoom to be used in more complex sound environments in terms of the number and location of sound sources. The embodiments further aim to improve perceived audio quality due to improved derivation of filtering gains / attenuation amounts. The improvement results from being able to take into account DOA estimates of at least one previous frame (e.g., DOA estimates from the last 40 frames) and the energy ratios of (all) both directions when forming the filtering gains for the current time frame.

[0054] Thus, embodiments aim to prevent "interfering" filter leakage into the output from directions that should have been filtered or attenuated. This therefore enhances the perceived audio zoom effect and prevents disrupting the user experience when several sound sources are present in the capture. Furthermore, the target (focal) direction can be efficiently amplified relative to other sound directions in a complex environment, again enhancing the zoom effect experience.

[0055] Therefore, the embodiments described herein relate to parametric spatial audio capture using multiple microphones. Furthermore, at least two parameters, direction and energy ratio, are estimated for each time-frequency tile based on the audio signals from the multiple microphones.

[0056] In these embodiments, the effect of the first estimated direction is taken into account when estimating the second direction to achieve improved accuracy in detecting multiple sound source directions, which in some embodiments can result in improved perceptual quality of the synthesized spatial audio.

[0057] It is therefore possible to use similar techniques as described in EP3791605, but implemented as described herein.

[0058] In effect, the embodiments described herein produce estimates of sound sources that are perceived as being more spatially stable and more accurate (with respect to their correct or actual location).

[0059] With reference to FIG. 1, a schematic diagram of an apparatus suitable for carrying out the embodiments described herein is shown.

[0060] In this example, an apparatus is shown that includes a microphone array 101. The microphone array 101 includes multiple (two or more) microphones configured to capture audio signals. The microphones in the microphone array can be of any suitable microphone type, arrangement, or configuration. The microphone audio signal 102 generated by the microphone array 101 can be passed to a spatial analyzer 103.

[0061] The phone device may comprise a spatial analyzer 103 configured to receive or otherwise acquire the microphone audio signal 102 and configured to spatially analyze the microphone audio signal to determine at least two dominant sounds or audio sources for each time-frequency block.

[0062] The spatial analyzer 103 may be the CPU of a mobile device or a computer in some embodiments. The spatial analyzer 103 is configured to generate a data stream containing the audio signal as well as metadata of the analyzed spatial information 104.

[0063] Depending on the use case, the data stream may be stored or compressed and transmitted elsewhere.

[0064] The apparatus further comprises a spatial synthesizer 105. The spatial synthesizer 105 is configured to obtain a data stream comprising the audio signal and metadata. In some embodiments, the spatial synthesizer 105 is implemented in the same apparatus as the spatial analyzer 103 (as shown herein in FIG. 1), but in some embodiments may also be implemented in a different apparatus or device.

[0065] The spatial synthesizer 105 may be implemented in a CPU or similar processor and is configured to generate an output audio signal 106 based on the audio signal and associated metadata from the data stream 104.

[0066] Furthermore, depending on the use case, output signal 106 can be in any suitable output format. For example, in some embodiments, the output format is a binaural headphone signal (e.g., the output device presenting the output audio signal is a set of headphones / earphones or the like) or a multi-channel loudspeaker audio signal (e.g., the output device is a set of speakers). Output device 107 (which, as noted above, may be, for example, headphones or loudspeakers) may be configured to receive output audio signal 106 and present the output to a listener or user.

[0067] These operations of the exemplary apparatus shown in Figure 1 may be illustrated by the flow diagram shown in Figure 2. Accordingly, the operation of the exemplary apparatus may be summarized as follows:

[0068] Step 201 obtains a microphone audio signal as shown in FIG.

[0069] The microphone audio signals are spatially analyzed to generate, for each time-frequency tile, a spatial audio signal and metadata including the direction and energy ratio of the first and second audio sources, as shown in FIG. 2 by step 203.

[0070] Spatial synthesis is applied to the spatial audio signals to generate appropriate output audio signals as shown in FIG. 2 by step 205 .

[0071] Step 207 outputs the output audio signal to an output device, as shown in FIG.

[0072] In some embodiments, spatial analysis can be used in conjunction with an IVAS codec. In this example, the spatial analysis output is in an IVAS-compatible MASA (Metadata Assisted Spatial Audio) format that can be directly fed to an IVAS encoder. The IVAS encoder generates an IVAS data stream. At the receiving end, an IVAS decoder can directly generate the desired output audio format. In other words, in such embodiments, there is no separate spatial synthesis block.

[0073] The spatial analyzer, indicated in FIG. 1 by reference numeral 103, is shown in more detail with respect to FIG.

[0074] In some embodiments, the spatial analyzer 103 comprises a stream audio signal generator 307. The stream audio signal generator 307 is configured to receive the microphone audio signal 102 and generate a stream audio signal 308 that is passed to the multiplexer 309. The audio stream signal is generated from the input microphone audio signal based on any suitable method. For example, in some embodiments, one or two microphone signals may be selected from the microphone audio signal 102. Alternatively, in some embodiments, the microphone audio signal 102 may be downsampled and / or compressed to generate the stream audio signal 308.

[0075] In the examples below, the spatial analysis is performed in the frequency domain, but it will be appreciated that in some embodiments the analysis can also be performed in the time domain using a time-domain sampled version of the microphone audio signal.

[0076] In some embodiments, the spatial analyzer 103 comprises a time-to-frequency transformer 301. The time-to-frequency transformer 301 is configured to receive the microphone audio signals 102 and transform them into the frequency domain. In some embodiments, before transformation, the time-domain microphone audio signals are represented as s i It can be expressed as (t). The transformation to the frequency domain can be performed by any suitable time-frequency transform, such as a Short-time Fourier transform (STFT) or a Quadrature mirror filter (QMF). The resulting time-frequency domain microphone signal 302 is iIt is denoted as (b,n), where i is the microphone channel index, b is the frequency bin index, and n is the time frame index. The value of b ranges from 0,...,B-1, where B is the number of bin indices at each time index n.

[0077] The frequency bins can be further combined into subbands k=0, .., K-1. Each subband consists of one or more frequency bins. Each subband k is divided into the lowest bin b k,low And the best bottle b k,high The width of the subbands is typically selected based on the characteristics of human hearing, and for example, the equivalent rectangular bandwidth (ERB) or the Bark scale can be used.

[0078] In some embodiments, the spatial analyzer 103 comprises a first direction analyzer 303. The first direction analyzer 303 is configured to receive the time-frequency domain microphone audio signal 302 and to generate an estimate of a first sound source for each time-frequency partition of a (first) first direction 314 and a (first) first ratio 316.

[0079] The first direction analyzer 303 is configured to generate an estimate for the first direction based on any suitable method, such as SPAC (as described in more detail in US9313599).

[0080] In some embodiments, for example, the most dominant direction of the temporal frame index is The time shift τ that maximizes the correlation between two (microphone audio signal) channels in subband k k It is estimated by searching for S i (b,n) is calculated by τ samples.

number

number

[0081] In the above formula, the "optimal" delay is searched between microphones 1 and 2. Re denotes the real part of the result, and * is the complex conjugate of the signal. The delay search range parameter D max is defined based on the distance between the microphones. In other words, τ k The value of is searched only within the physically possible range, taking into account the distance between the microphones and the speed of sound.

[0082] Then, the angle of the first direction is

number

[0083] As shown, there is still uncertainty in the sign of the angle. Above, we defined the directional analysis between microphone 1 and microphone 2. A similar procedure can then be repeated between other microphone pairs to resolve the ambiguity (and / or obtain the direction by reference to another axis). In other words, using information from other analysis pairs,

number

[0084] For example, if the microphone array includes three microphones, the first, second, and third microphones are arranged in a configuration with a first pair of microphones (the first and third microphones) spaced apart by a distance on a first axis and a second pair of microphones (the first and second microphones) spaced apart by a distance on a second axis (in this example, the first axis is perpendicular to the second axis). Additionally, in this example, the three microphones may be on the same third axis, defined as perpendicular to the first and second axes (perpendicular to the plane of the paper on which the figure is printed). Analysis of the delay between the second pair of microphones results in two alternative angles, α and −α. Analysis of the delay between the second pair of microphones can be used to determine which of the alternative angles is correct. In some embodiments, the information needed from this analysis is whether the sound arrives at microphone 1 or 3 first. If the sound arrives at microphone 3, angle α is correct. Otherwise, −α is selected.

[0085] Furthermore, based on the inference between several microphone pairs, the first spatial analyzer determines the correct direction angle.

number

[0086] In some embodiments where there are only two microphones, the directional ambiguity cannot be resolved. In such embodiments, the spatial analyzer is configured to define all sources as always in front of the device. This situation is the same when there are three or more microphones, but their positions do not allow for, for example, behind-the-scenes analysis.

[0087] Although not disclosed herein, multiple pairs of microphones on the vertical axis can determine elevation and azimuth estimates.

[0088] The first direction analyzer 303 further includes, for example:

number

[0089] Values ​​range from -1 to 1, and are typically further restricted to 0 to 1.

[0090] In some embodiments, the first direction analyzer 303 is configured to generate a modified time-frequency microphone audio signal 304. The modified time-frequency microphone audio signal 304 is one in which the first source component is removed from the microphone signal.

[0091] So, for example, for the first microphone pair (microphones 1 and 2). For each subband k, the second microphone signal is sample shifted to obtain a shifted second microphone signal, which is the delay that provides the highest correlation for subband k.

[0092] The estimate of the source components is the average of these time-aligned signals.

number

[0093] In some embodiments, any other suitable method for determining source components may be used.

[0094] Once an estimate of the source component has been determined (e.g., in the example equation above), it can be removed from the microphone audio signal. On the other hand, simultaneous sources are not in phase, and therefore are attenuated. Now, the (shifted and unshifted) microphone signals

number

number

number

[0095] These modified signals

number

number

[0096] In some embodiments, the spatial analyzer 103 comprises a second direction analyzer 305. The second direction analyzer 305 is configured to estimate the time-frequency microphone audio signal 302, the modified time-frequency microphone audio signal 304, the first direction 314, and the first ratio 316 to generate second direction 324 and second ratio 326 estimates.

[0097] The estimation of the second direction parameter value may employ the same subband structure as the first direction estimation; Similar operations can be followed as described above for the first direction estimation.

[0098] It is therefore possible to estimate the second directional parameter. Modified time-frequency microphone audio signal 304

number

number

[0099] Furthermore, in some embodiments, the energy ratio is limited such that the sum of the first and second ratios should not be greater than two.

[0100] In some embodiments, the second ratio is

number

number

[0101] In the above example, there are several microphone pairs, so the corrected signal is

number

[0102] The first direction estimate 314, the first ratio estimate 316, the second direction estimate 324, and the second ratio estimate 326 are passed to a multiplexer (mux) 309 configured to generate the data stream 104 from combining the estimates with the streamed audio signal 308.

[0103] With reference to FIG. 4, a flow chart summarizing an exemplary operation of the spatial analyzer shown in FIG. 3 is shown.

[0104] A microphone audio signal is obtained by step 401 as shown in FIG.

[0105] Step 402 then generates a stream audio signal from the microphone audio signal, as shown in FIG.

[0106] The microphone audio signal may further be time-to-frequency domain transformed as shown in FIG. 4 by step 403 .

[0107] A first direction and a first ratio parameter estimate may then be determined, as shown in FIG. 4, by step 405 .

[0108] Then, step 407 allows the time-frequency domain microphone audio signal to be modified (removing the first source component) as shown in FIG.

[0109] The modified time-frequency domain microphone audio signal is then analyzed to determine a second direction and a second ratio parameter estimate, as shown in FIG. 4, per step 409.

[0110] Then, in step 411, the first direction, first ratio, second direction, and second ratio parameter estimates and the stream audio signal are multiplexed to generate a data stream (which may be a MASA format data stream), as shown in FIG.

[0111] In the following example, a spatial filtering method and apparatus is described in which several gain parameters are determined or calculated and set to adjust the filtering process. These gains can be divided into per-band gains, history-based (temporal) gains, and frame-based smoothing gains.

[0112] In the following examples, the two estimated directions of arrival (DOA) per subband are given a direct-to-ambient (DA) ratio estimate, which basically indicates how much of the corresponding direction estimate is considered the "direct" signal portion and how much is considered the "ambient" signal portion. In these examples, the term direct refers to the signal arriving directly from the sound source, and ambient refers to the echo and background noise present in the environment. The direct and ambient components of the signal for each subband b can have the range [0,1],

number

[0113] In some embodiments, the method begins after obtaining the direction and extent of the spatial filtering zone (which may also be defined as a focal sector of interest or a zoom sector) by checking through the subbands whether either or both of the two direction estimates are located inside the sector of interest. In the following example, the spatial filtering is positive notch filtering, where the audio signal within the sector of interest is increased relative to the audio signal outside the sector of interest. However, in some embodiments, the spatial filtering is negative notch filtering, where the audio signal within the sector of interest is decreased compared to the audio signal outside the sector of interest. The difference between the two is It will be appreciated that if the sector gain is greater than the out-sector gain that results in a positive spatial notch filter, or if the sector gain is less than the out-sector gain that results in a negative spatial notch filter.

[0114] A simplified illustration of these three main scenarios is shown with respect to FIG.

[0115] In this example, sounds are amplified within the sector and attenuated outside the sector, but the process is also significantly affected by the DA ratio of the direction estimate.

[0116] For example, the DA ratio estimate can be thought of as a weight on the actual direction estimate. The numbers in the table below are just examples to demonstrate the basic principles of their effect on deriving the example filter gain G(b). The first two columns show the cases where either of the two sources is estimated as ambient-like, which means that the direction estimate should not be used as such for filtering. [Table 1]

[0117] Therefore, a low DA ratio value can indicate that the corresponding direction estimate may not be caused by a real source, and in some cases there may be no active direct source or only one source during capture. In some embodiments, sector edges can also have regions where the applied subband gains are linearly smoothed to avoid abrupt gain changes at the sector edges.

[0118] Thus, as shown in FIG. 5, there is a first scenario 501, where both sound sources are within the sector, resulting in filtering gains corresponding to each direction estimate g1(b) and g2(b) both being greater than 1, and therefore resulting in a spatial gain G(b) greater than 1.

[0119] A second scenario 503 is shown, where one of the sources is within the sector filtering gain corresponding to one direction estimate (first g1(b)) and the other (second g2(b)) is greater than 1, thus resulting in a spatial gain G(b) that is close to 1.

[0120] Furthermore, a third scenario 505 is shown in which both sound sources are outside the sector, resulting in a filtering gain corresponding to each direction estimate g1(b), g2(b) being less than 1, and therefore a spatial gain G(b) less than 1.

[0121] In some embodiments, the energy of subband b of the input signal spectrum X(b) before any energy adjustment can be estimated as follows:

number

number

number

[0122] In some embodiments, a band gain is derived for each subband b based on the band's direction estimates d1 and d2. The direction estimate may be located inside the focus sector, outside the focus sector, or in a region near the sector edge (the so-called edge zone). The direct energy component for the first direction estimate d1 for subband b may be modified as follows:

number

number

number

number

[0123] For band b after energy adjustment, the target energy, which is initialized to 0 before the first frame, can be defined as:

number

number

[0124] To take into account the second direction estimate d2, the g2(b) gain value is calculated similarly to the g1(b) value, and then the gain is calculated as the overall band gain

number

[0125] Furthermore, in some embodiments, a temporal filtering gain is calculated for each subband for both direction estimates d1 and d2 to smooth the filtering gain over time. This prevents unnatural pumps and notches from appearing across the filter gain. In many cases, the estimated source DA ratio values ​​may vary across subbands, so averaging the DA ratio across the entire filtering frequency range provides a good estimate of how ambient the sound environment is at the current time frame f. The ratio average is calculated for each frame for the first direction estimate as follows:

number

number

number

number

[0126] When the history section is filled with such flags, the number of "true" flags in each subband b of d1, N1T(b) is denoted by the pseudo scaling variable

number

number

[0127] The number of direction estimates within a sector for each subband b in the past N1T(b) is

number

[0128] The temporal gain for direction estimate d2 is calculated in the same way as for d1, and the actual temporal filter gain is multiplied by

number

[0129] In some embodiments, the direction estimate across all subbands within a single time frame can vary significantly depending on the number and type of sound sources present in the sound environment. Therefore, an additional frame smoothing gain is required to smooth the spectrum to prevent sudden peaks and notches in the spectral envelope in each frame. First, the sum of the ratio means of d1 and d2 is given by

number

number

number

[0130] The previously derived attenuation state is the actual filter smoothing gain for each subband,

number

number

number

[0131] Once all the different gain types, bandwidth gain, time gain and frame gain, have been calculated, The actual output filter gain is

number

[0132] An example of the benefits of implementing embodiments described herein is shown in FIG. 6. Specifically, FIG. 6 illustrates the output signal level in dB of a known spatial filter that uses only a single direction estimate per subband 601, illustrating a spatial filtering approach according to some embodiments 603. In this example, the audio focus direction is set directly in front of the device, and the signal consists of a speaker speaking first in front of the device, then moving behind the device in the center of the signal, and finally returning to the front of the device again. Additionally, music is played from a speaker located to the left of the capture device. It can be seen that, on average, embodiments amplify audio from the front by approximately 2-3 dB compared to known methods.

[0133] Additionally, the embodiments also attenuate audio from behind the device by 2-3 dB more compared to known spatial filtering methods, meaning that the embodiments increase the overall focus effect gain by an average of 4-6 dB overall. This is a clearly audible and significant difference that improves the perceived audio zoom experience in most cases. As long as direction estimates d1 and d2 can be estimated from the capture, the spatial filter can always improve its performance compared to having only estimate d1.

[0134] With reference to FIG. 7, an overview of the operation of the embodiments described herein is shown.

[0135] The first action is to calculate or determine direction estimates for d1 and d2 for subband b, as shown in FIG. 7, per step 701.

[0136] Next, a first check can be performed to determine whether d1 is within a sector, as shown in FIG. 7, via step 703.

[0137] If d1 is in the sector, a further check can be made to determine if d2 is in the sector, as shown in FIG. 7, via step 705.

[0138] If both d1 and d2 are in the sector, then subband b is amplified according to the DA ratio of the associated estimates of both d1 and d2, as shown in diagram 707.

[0139] If d1 is not in the sector, a further check can be made to determine if d2 is in the sector, as shown in FIG. 7, via step 709.

[0140] If d1 is in the sector but d2 is not, or if d1 is not in the sector but d2 is in the sector, then subband b can be amplified according to the in-sector estimated DA ratio and attenuated according to the out-sector estimated DA ratio, as shown in FIG. 7 by step 711.

[0141] If both d1 and d2 are outside the sector, then subband b is attenuated according to the DA ratio of the associated estimates of both d1 and d2, as shown in diagram 713. With reference to Figure 8, a flow diagram illustrating the generation of gains according to some embodiments is shown.

[0142] Therefore, in some embodiments, the band gain g(b) is calculated in both directions by step 801 as shown in FIG.

number

[0143] Then, in some embodiments, the band gains are calculated by step 803 as the combined band gains, as shown in FIG.

number

[0144] Next, in step 805, the temporal gain g1 is calculated as shown in FIG. t (b), g2 t (b) is generated for each subband.

[0145] The temporal gains are then calculated by step 807 as the combined temporal gains, as shown in FIG.

number

[0146] Then, the frame smoothing gain g1 s (b), g2 s (b) may be determined for each subband and direction by step 809 as shown in FIG.

[0147] The frame smoothing gain is then calculated by step 811 as the combined frame smoothing gain .times. ...

number

[0148] Then, as shown in FIG. 8 by step 813, the combined frame smoothing gain, combined time gain, and combined bandwidth gain are calculated.

number

[0149] With reference to FIG. 9, an exemplary spatial synthesizer 105 as shown in FIG. 1 is shown.

[0150] The spatial synthesizer 105, in some embodiments, comprises a demultiplexer 1201. The demultiplexer (Demux) 1201, in some embodiments, receives the data stream 104 and separates the data stream into a streamed audio signal 1208 and spatial parameter estimates, such as a first direction 1214 estimate, a first ratio 1216 estimate, a second direction 1224 estimate, and a second <ratio> {ratio} 1226 estimate.

[0151] These are then passed to the spatial processor / synthesizer 1203 .

[0152] The spatial synthesizer 105 comprises a spatial processor / synthesizer 1203 and is configured to receive the estimates and the stream audio signals and to render an output audio signal. The spatial processing / synthesis may be any suitable two-way based synthesis, such as that described in EP3791605.

[0153] 10 and 11 show an end-to-end implementation of an embodiment. With respect to Fig. 10, it is shown that there is a capture device 1101 and a playback device 1111 communicating over a transport / storage channel 1105.

[0154] The capture device 1101 is configured as described above and is configured to transmit filtered audio 1109. In addition, filter direction / range information 1107 can be received from the playback device 1111.

[0155] 11, there is shown a capture device 1101 configured to transmit unfiltered audio 1119 that is received by a playback device 1111. The playback device comprises a spatial filter 1103 configured to apply spatial filtering as described in the embodiments described herein.

[0156] 12, an exemplary electronic device is shown that may be used as a computer, encoder processor, decoder processor, or any of the functional blocks described herein. The device may be any suitable electronic device or apparatus. For example, in some embodiments, device 1600 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc.

[0157] In some embodiments, device 1600 includes at least one processor or central processing unit 1607. Processor 1607 may be configured to execute various program code, such as the methods described herein.

[0158] In some embodiments, the device 1600 comprises a memory 1611 .

[0159] In some embodiments, at least one processor 1607 is coupled to a memory 1611. The memory 1611 may be any suitable storage means. In some embodiments, the memory 1611 comprises a program code section for storing program code executable on the processor 1607. Additionally, in some embodiments, the memory 1611 may further comprise a stored data section for storing data, e.g., data that has been processed or is to be processed in accordance with embodiments described herein. The executed program code stored in the program code section and the data stored in the stored data section may be retrieved by the processor 1607 via the memory-processor coupling as needed.

[0160] In some embodiments, device 1600 comprises a user interface 1605. User interface 1605, in some embodiments, may be coupled to a processor 1607. In some embodiments, processor 1607 may control the operation of user interface 1605 and receive input from user interface 1605. In some embodiments, user interface 1605 may allow a user to input commands into device 1600, for example, via a keypad. In some embodiments, user interface 1605 may allow a user to obtain information from device 1600. For example, user interface 1605 may comprise a display configured to display information from device 1600 to the user. User interface 1605, in some embodiments, may comprise a touch screen or touch interface that can both allow information to be input into device 1600 and further display information to the user of device 1600.

[0161] In some embodiments, apparatus 1600 comprises an input / output port 1609. In some embodiments, input / output port 1609 comprises a transceiver. The transceiver in such embodiments may be coupled to processor 1607 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means may in some embodiments be configured to communicate with other electronic devices or apparatuses via a wired or wired connection.

[0162] The transceiver may communicate with the further device by any suitable known communication protocol, for example, in some embodiments the transceiver may use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth®, or an infrared data channel (IRDA).

[0163] The transceiver input / output port 1609 may be configured to transmit / receive audio signals, bitstreams, and in some embodiments, perform the operations and methods as described above by using a processor 1607 executing appropriate code.

[0164] In general, various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. Although various aspects of the present invention may be illustrated and contemplated as block diagrams, flow charts, or using some other graphical representation, it is to be appreciated that these blocks, devices, systems, techniques, or methods contemplated herein may be implemented, by way of non-limiting example, in hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controllers, or other computing devices, or any combination thereof.

[0165]

[0013] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device, such as within a processor entity, or by hardware, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that any blocks of logic flows, such as those shown in the figures, may represent program steps, or interconnected logic circuits, blocks and functions, or combinations of program steps and logic circuits, blocks and functions. Software may be stored on physical media, such as memory chips, or memory blocks implemented within a processor, magnetic media, and optical media.

[0166] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed and removable memory, etc. The data processor may be of any type suitable for the local technology environment and may include, by way of non-limiting examples, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processor architecture.

[0167] Embodiments of the present invention can be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is a large-scale, highly automated process. Complex and powerful software tools are available to convert logic-level designs into semiconductor circuit designs ready to be etched and formed on semiconductor substrates.

[0168] Programs such as those offered by Synopsys, Inc. of Mountain View, California, and Cadence Design, of San Jose, California, automatically route conductors and locate components on semiconductor chips using well-established design rules and pre-stored libraries of design modules. Once the design of a semiconductor circuit is complete, the resulting design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be sent to a semiconductor manufacturing facility or "fab" for fabrication.

[0169] The foregoing description has provided a full and informative description of exemplary embodiments of the present invention, by way of illustrative and non-limiting example. However, various modifications and adaptations will become apparent to those skilled in the art in view of the foregoing description upon perusal of the accompanying drawings and the appended claims. However, all such similar modifications of the teachings of this invention will still fall within the scope of the present invention, as defined in the appended claims.

Claims

1. 1. An apparatus comprising at least one processor and at least one memory containing computer program code, The at least one memory and the computer program code, using the at least one processor, cause the device to at least: obtaining a plurality of audio signals from a respective plurality of microphones; determining a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the plurality of audio signals based on processing the plurality of audio signals; determining second sound source direction parameters and second sound source energy parameters in the one or more frequency bands of the plurality of audio signals based on processing the plurality of audio signals; obtaining a region defining a direction and / or range for a filter; generating the filters to be applied to the plurality of audio signals, The filter gain / attenuation parameter is the first sound source direction parameter, the first sound source energy parameter; the second sound source direction parameter, the second sound source energy parameter based on the region for An apparatus configured to cause a The generated filter causes the device to generate a first band gain / attenuation value based on the first sound source direction parameter being within or outside the region; generating a second band gain / attenuation value based on the second sound source direction parameter being within or outside the region; combining the first band gain / attenuation value and the second band gain / attenuation value to generate a combined band gain / attenuation value; Device.

2. The acquired region is then transmitted to the device. a direction and range defining said region, together with in-band gain / attenuation coefficients based on sound source direction parameters within said region; an out-of-band gain / attenuation factor based on the sound source direction parameters outside the region; a direction and range defining said region, together with an in-band gain / attenuation coefficient based on said sound source direction parameters within said region; an out-of-band gain / attenuation factor based on the sound source direction parameters outside the region; a further region within the edge zone region that defines the edge zone region together with an edge zone gain / attenuation coefficient based on the sound source direction parameter; obtain at least one of 10. The apparatus of claim 1.

3. The generated filter causes the device to generate a first temporal gain / attenuation value based on a time average of the mean band value of the first sound source energy parameter and the number of times the first sound source direction parameter is within the region over a defined time period; generating a second temporal gain / attenuation value based on a time average of the average band value of the second sound source energy parameter and a number of times the second sound source direction parameter is within the region over the defined time period; generating a composite temporal gain / attenuation value based on a combination of the first temporal gain / attenuation value and the second temporal gain / attenuation value to generate a composite temporal gain / attenuation value; 10. The apparatus of claim 1.

4. The generated filters are applied to the plurality of audio signals; the filter gain / attenuation parameters cause the device to generate a synthesized frame-averaged value based on a combination of a frame-averaged first sound source energy parameter and a frame-averaged second sound source energy parameter; generating a frame smoothing gain / attenuation based on the combined frame averaged value and the number of times the first and second sound source direction parameters are within the region of the filter over a frame period; 10. The apparatus of claim 1.

5. the generated filters are to be applied to the audio signals; the filter gain / attenuation parameters cause the device to generate a filter gain / attenuation for the frequency band based on a combination of the frame smoothing gain / attenuation, a composite temporal gain / attenuation value, and a composite band gain / attenuation value; 5. The apparatus of claim 4.

6. The step of processing the plurality of audio signals includes: causing the device to provide one or more modified audio signals based on the plurality of audio signals; causing the apparatus to determine, in the one or more frequency bands of the plurality of audio signals, second sound source direction parameters and second sound source energy parameters based on processing the plurality of audio signals; causing the apparatus to determine a second sound source direction parameter and a second sound source energy parameter based on the modified audio signal in the one or more frequency bands of the plurality of audio signals; 10. The apparatus of claim 1.

7. 7. The apparatus of claim 6, wherein the provided one or more modified audio signals cause the apparatus to generate modified multiple audio signals based on modifying the multiple audio signals with a projection of a first sound source defined by the first sound source direction parameter.

8. 8. The apparatus of claim 7, wherein the provided one or more modified audio signals further cause the apparatus to determine the second sound source direction parameter by processing the modified audio signals.

9. The obtained region defining the direction and / or range for the filter may be: The device of claim 1 based on user input.

10. 1. A method for an apparatus, the method comprising: acquiring a plurality of audio signals from a respective plurality of microphones; determining a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the plurality of audio signals based on processing the plurality of audio signals; determining second sound source direction parameters and second sound source energy parameters in the one or more frequency bands of the plurality of audio signals based on processing the plurality of audio signals; obtaining a region defining a direction and / or range for a filter; generating the filters to be applied to the plurality of audio signals, The filter gain / attenuation parameter is the first sound source direction parameter, the first sound source energy parameter; the second sound source direction parameter, and the second sound source energy parameter based on the region for A method comprising: generating the filters to be applied to the plurality of audio signals, wherein filter gain / attenuation parameters are generated based on the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the region relative to the second sound source energy parameter; generating a first band gain / attenuation value based on the first sound source direction parameter within or outside the region; generating a second band gain / attenuation value based on the second sound source direction parameter being within or outside the region; combining the first band gain / attenuation value and the second band gain / attenuation value to generate a combined band gain / attenuation value; A method comprising:

11. The step of obtaining the region defining the direction and / or the range for the filter comprises: a direction and range defining said region, together with in-band gain / attenuation coefficients based on sound source direction parameters within said region; an out-of-band gain / attenuation factor based on the source direction parameters within the region; a direction and range defining said region, together with in-band gain / attenuation coefficients based on said sound source direction parameters within said region; an out-of-band gain / attenuation factor based on the sound source direction parameters outside the region; a further region within the edge zone region that defines the edge zone region together with an edge zone gain / attenuation coefficient based on the sound source direction parameter; at least one of: The method of claim 10.

12. generating the filters to be applied to the plurality of audio signals, The filter gain / attenuation parameter is the first sound source direction parameter, the first sound source energy parameter; the second sound source direction parameter, and the second sound source energy parameter based on the region for an average band value of the first sound source energy parameter; generating a first temporal gain / attenuation value based on a time average of the first sound source direction parameter and the number of times the first sound source direction parameter is within the region over a defined period of time; generating a second temporal gain / attenuation value based on a time average of the mean band value of the second sound source energy parameter and the number of times the second sound source direction parameter is present within the region over the defined time period; generating a composite temporal gain / attenuation value based on a combination of the first temporal gain / attenuation value and the second temporal gain / attenuation value to generate a composite temporal gain / attenuation value; The method of claim 10, comprising:

13. generating the filters to be applied to the plurality of audio signals, The filter gain / attenuation parameter is the first sound source direction parameter, the first sound source energy parameter; the second sound source direction parameter, the second sound source energy parameter based on the region for generating a synthesized frame-averaged value based on a combination of the frame-averaged first sound source energy parameter and the frame-averaged second sound source energy parameter; generating a frame smoothing gain / attenuation based on the combined frame averaged value and the number of times the first and second sound source direction parameters fall within the region of the filter over a frame period; The method of claim 11 , comprising:

14. the generating a filter is applied to the plurality of audio signals; 14. The method of claim 13, wherein generating filter gain / attenuation parameters comprises generating filter gain / attenuation for the frequency band based on a combination of the frame-smoothed gain / attenuation, a composite temporal gain / attenuation value, and a composite band gain / attenuation value.

15. The step of processing the plurality of audio signals includes: providing one or more modified audio signals based on the plurality of audio signals; determining second sound source direction parameters and second sound source energy parameters in the one or more frequency bands of the plurality of audio signals based on processing the plurality of audio signals; Including, determining a second sound source direction parameter and a second sound source energy parameter based on the modified audio signal in the one or more frequency bands of the plurality of audio signals; The method of claim 10.

16. 16. The method of claim 15, wherein providing one or more modified audio signals based on the plurality of audio signals comprises generating the modified plurality of audio signals based on modifying the plurality of audio signals with a projection of a first sound source defined by the first sound source direction parameters.

17. providing one or more modified audio signals based on the plurality of audio signals, determining at least a second sound source direction parameter in the one or more frequency bands of the plurality of audio signals based at least in part on the one or more modified audio signals; determining the at least second sound source direction parameter by processing the modified plurality of audio signals in the one or more frequency bands of the plurality of audio signals; 17. The method of claim 16, further comprising:

18. The method of claim 10 , wherein obtaining the region defining the direction and / or the range for the filter comprises obtaining the region based on user input.