Spatial Audio Filtering in Spatial Audio Capture

The method enhances spatial audio capture by estimating two sound source directions and energy ratios for each frequency band, improving audio quality and stability in complex environments by preventing interference and enhancing the audio zoom effect.

JP7708729B2Active Publication Date: 2025-07-15NOKIA TECHNOLOGIES OY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022159369
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-10-04
Filing Date
2022-10-03
Publication Date
2025-07-15
Estimated Expiration
2042-10-03

AI Technical Summary

Technical Problem

Existing spatial audio capture technologies face challenges in accurately distinguishing and amplifying or attenuating multiple sound sources in complex audio environments, leading to interference and reduced audio quality due to insufficient direction-of-arrival (DOA) estimates and energy ratio calculations.

Method used

Implementing a method that estimates two sound source directions and their energy ratios for each frequency band, using a spatial analyzer to generate filters based on these parameters, enhancing the audio zoom effect and preventing interference by combining DOA estimation values and energy ratios across frames.

Benefits of technology

Improves the perceived audio quality by efficiently amplifying or attenuating sound sources in complex environments, providing a more accurate and stable spatial audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708729000050
    Figure 0007708729000050
  • Figure 0007708729000051
    Figure 0007708729000051
  • Figure 0007708729000052
    Figure 0007708729000052
Patent Text Reader

Abstract

To provide an apparatus and method to generate such a spatial audio capture that makes a result experienced by a listener more realistic.SOLUTION: A spatial analyzer obtains two or more audio signals from two or more microphones, subjects the obtained audio signals to time-frequency region conversion, and generates a stream audio signal. The spatial analyzer further determines a first sound source direction parameter, a first sound source energy ratio parameter, a second sound source direction parameter, and a second sound source energy ratio parameter based on the audio signals subjected to the time-frequency region conversion, and generates a data stream by multiplexing the parameters and the stream audio signal.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to an apparatus and method for spatial audio filtering within spatial audio capture.

Background Art

[0002] Spatial audio capture using microphone arrays is often utilized in many up-to-date digital devices such as mobile devices and cameras, together with video capture. Spatial audio capture can be reproduced using headphones or loudspeakers in order to provide the user with an experience of the audio scene captured by the microphone array.

[0003] Parametric spatial audio capture methods enable spatial audio capture using diverse microphone configurations and setups, and thus can be used in consumer devices such as mobile phones. Parametric spatial audio capture methods are based on signal processing solutions for analyzing the spatial audio field around the device, utilizing the available information from multiple microphones. Typically, these methods perceptually analyze the microphone audio signals to determine relevant information within the frequency band. This information includes, for example, the direction of the dominant sound source (or sound source or audio object), and the relationship between the source energy and the overall band energy. Based on this determined information, spatial audio can be reproduced using, for example, headphones or loudspeakers. Thus, ultimately, the user or listener can experience the ambient audio as if present in the audio scene being recorded by the capture device.

[0004] The better the audio analysis and synthesis performance, the more realistic the result experienced by the user or listener.

Summary of the Invention

[0005] According to a first aspect, there is provided a step of obtaining a plurality of audio signals from a plurality of microphones respectively, a step of determining a first sound source direction parameter and a first sound source energy parameter based on the processing of the plurality of audio signals in one or more frequency bands of the plurality of audio signals, a step of determining a second sound source direction parameter and a second sound source energy parameter based on the processing of the plurality of audio signals in one or more frequency bands of the plurality of audio signals, a step of obtaining a region defining a direction and / or a range for a filter, and a step of generating the filter to be applied to the plurality of audio signals, wherein a filter gain / attenuation parameter is generated based on the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the region related to the second sound source energy parameter. An apparatus is provided that includes means configured to perform the steps.

[0006] Means configured to generate a filter to be applied to a plurality of audio signals, wherein a filter gain / attenuation parameter is generated based on a region related to a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a second sound source energy parameter. The first sound source direction parameter generates a first band gain / attenuation value based on whether it is inside or outside the region, the second sound source direction parameter generates a second band gain / attenuation value based on whether it is inside or outside the region, and the first band gain / attenuation value and the second band gain / attenuation value can be combined to generate a combined band gain / attenuation value.

[0007] Means configured to obtain a region defining a direction and / or a range for a filter may obtain at least one of the direction and range defining the region, an in-band gain / attenuation coefficient based on the sound source direction parameter being within the region, an out-of-band gain / attenuation coefficient based on the sound source direction parameter being outside the region, the direction and range defining the region together with the in-band gain / attenuation coefficient based on the sound source direction parameter being within the region, the direction and range defining the region together with an edge zone band gain / attenuation coefficient based on the sound source direction parameter being within an edge zone region, an out-of-band gain / attenuation coefficient based on the sound source direction parameter being outside the region, and a further range defining the edge zone region.

[0008] Means configured to generate the filter applied to the plurality of audio signals, wherein filter gain / attenuation parameters are generated based on the region in relation to the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter, may generate a first time gain / attenuation value based on a time average of the average band value of the first sound source energy parameter, the first sound source direction parameter being within the region over a defined time period, generate a second time gain / attenuation value based on a time average of the average band value of the second sound source direction parameter and the number of times the second sound source direction parameter is present within a defined time period, and generate a combined time gain / attenuation value based on a combination of the first time gain / attenuation value and the second time gain / attenuation value, and may be configured to generate the combined time gain / attenuation value.

[0009] Means configured to generate a filter to be applied to a plurality of audio signals, The filter gain / attenuation parameter is generated based on regions related to the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter, and may be configured to generate a composite frame average value based on a combination of the frame-averaged first sound source energy parameter and the frame-averaged second sound source energy parameter. An apparatus is provided that includes means configured to perform generating frame-smoothed gain / attenuation based on the frame average value and the number of times the first and second sound source direction parameters are within a filter region during a frame period.

[0010] Means configured to generate a filter applied to a plurality of audio signals, wherein the filter gain / attenuation parameter is generated based on regions related to the first sound source direction parameter and the first sound source energy parameter, the means including the second sound source direction parameter, and the second sound source energy parameter may be configured to generate filter gain / attenuation based on a combination of frame-smoothed gain / attenuation, a composite time gain / attenuation value, and a composite band gain / attenuation value.

[0011] Processing of the plurality of audio signals may be configured to provide one or more modified audio signals based on the plurality of audio signals, and means configured to determine the second sound source direction parameter and the second sound source energy parameter based on processing of the plurality of audio signals in one or more frequency bands of the plurality of audio signals may be configured to determine the second sound source direction parameter and the second sound source energy parameter based on the modified audio signals in one or more frequency bands of the plurality of audio signals.

[0012] Means configured to provide one or more modified audio signals based on a plurality of audio signals may be further configured to generate a plurality of modified audio signals based on modifying the plurality of audio signals using a projection of a first sound source defined by a first sound source direction parameter, and in at least one frequency band of the plurality of audio signals, at least a second sound source direction parameter is at least partially based on at least a portion of the one or more modified audio signals, the means configured as such processes the plurality of modified audio signals in at least one frequency band of the plurality of audio signals to determine at least the second sound source direction parameter.

[0013] Means configured to obtain a region defining a direction and / or range of a filter may be configured to obtain the region based on a user input.

[0014] According to a second aspect, a method for an apparatus, the method comprising: obtaining a plurality of audio signals from respective ones of a plurality of microphones; determining, in at least one frequency band of the plurality of audio signals, a first sound source direction parameter and a first sound source energy parameter based on processing the plurality of audio signals; determining, in at least one frequency band of the plurality of audio signals, a second sound source direction parameter and a second sound source energy parameter based on processing the plurality of audio signals; obtaining a region defining a direction and / or range for a filter; and generating a filter to be applied to the plurality of audio signals, wherein a filter gain / attenuation parameter is generated based on a region related to the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter.

[0015] Generating a filter applied to a plurality of audio signals, wherein the filter gain / attenuation parameter is generated based on a region related to a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a second sound source energy parameter, the step may include: generating a first band gain / attenuation value based on whether the first sound source direction parameter is within or outside the region; generating a second band gain / attenuation value based on whether the second sound source direction parameter is within or outside the region; and combining the first band gain / attenuation value and the second band gain / attenuation value to generate a combined band gain / attenuation value.

[0016] The step of obtaining a region defining a direction and / or a range for the filter may include at least one of: the direction and range defining the region, the in-band gain / attenuation coefficient based on the sound source direction parameter being within the region, the out-of-band gain / attenuation coefficient based on the sound source direction parameter being outside the region; the direction and range defining the region, the in-band gain / attenuation coefficient based on the sound source direction parameter being within the region, the in-band gain / attenuation coefficient based on the sound source direction parameter being within the region, and the out-of-band gain / attenuation coefficient based on the sound source direction parameter being outside the region together; and the direction and range defining the region, the out-of-band gain / attenuation coefficient based on the sound source direction parameter being within the edge zone region together with a further range defining the edge zone region.

[0017] A step of generating a filter to be applied to a plurality of audio signals, wherein the filter gain / attenuation parameter is generated based on a region related to a first sound source direction parameter, a first sound source direction parameter, a second sound source direction parameter, and a second sound source energy parameter, the step includes generating a first temporal gain / attenuation value based on a temporal average of the average band value of the first sound source energy parameter; generating a first sound source direction parameter based on a temporal average of the average band value of the first sound source energy parameter; generating a number of times the second sound source direction parameter is within the region over a time period defined by the second sound source direction parameter based on a temporal average of the temporal average band value of the second sound source direction parameter and the number of times the second sound source direction parameter exists within the defined time period; and generating a combined temporal gain / attenuation value based on a combination of the first temporal gain / attenuation value and a second temporal gain / attenuation value to generate a combined temporal gain / attenuation value for generating a combined value of temporal gain / attenuation.

[0018] The step of generating a filter to be applied to a plurality of audio signals, wherein the filter gain / attenuation parameter generated based on a region related to a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a second sound source energy parameter includes generating a combined frame average value based on a combination of the frame-averaged first sound source energy parameter and the frame-averaged second sound source energy parameter; and generating frame-smoothed gain / attenuation based on the combined frame average value and the number of times the first and second sound source direction parameters are within the filter region over the frame period.

[0019] A step of generating a filter applied to a plurality of audio signals, wherein the filter gain / attenuation parameter is generated based on a region related to a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a second sound source energy parameter, the step may include a step of generating a filter gain / attenuation for a band based on a combination of a frame smoothing gain / attenuation, a synthesis time gain / attenuation value, and a synthesis band gain / attenuation value.

[0020] The step of processing a plurality of audio signals may include a step of providing one or more modified audio signals based on the plurality of audio signals, and determining a second sound source direction parameter and a second sound source energy parameter based on the processing of the plurality of audio signals in one or more frequency bands of the plurality of audio signals may include determining a second sound source direction parameter and a second sound source energy parameter based on the modified audio signals in one or more frequency bands of the plurality of audio signals.

[0021] The step of providing one or more modified audio signals based on the plurality of audio signals may include a step of generating a plurality of modified audio signals based on modifying the plurality of audio signals using a projection of a first sound source defined by a first sound source direction parameter. In one or more frequency bands of the plurality of audio signals, the step of determining at least a second sound source direction parameter based at least in part on one or more modified audio signals may include determining at least a second sound source direction parameter in one or more frequency bands of the plurality of audio signals by processing the plurality of modified audio signals.

[0022] The step of obtaining a region defining the direction and / or range of the filter may include a step of obtaining the region based on user input.

[0023] According to a third aspect, there is provided an apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being operative, using the at least one processor, to cause the apparatus to, at least: obtain a plurality of audio signals from respective plurality of microphones; determine a first sound source direction parameter and a first sound source energy parameter based on processing of the plurality of audio signals in one or more frequency bands of the plurality of audio signals; determine a second sound source direction parameter and a second sound source energy parameter based on processing of the plurality of audio signals; obtain a region defining a direction and / or a range for a filter; generate a filter to be applied to the plurality of audio signals, wherein a filter gain / attenuation parameter is generated based on the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the region related to the second sound source energy parameter.

[0024] An apparatus for generating a filter to be applied to a plurality of audio signals, wherein a filter gain / attenuation parameter is generated based on a region related to a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a second sound source energy parameter, the apparatus is operative to generate a first band gain / attenuation value based on whether the first sound source direction parameter is within or outside the region, generate a second band gain / attenuation value based on whether the second sound source direction parameter is within or outside the region, and combine the first band gain / attenuation value and the second band gain / attenuation value to generate a combined band gain / attenuation value.

[0025] An apparatus for obtaining a region that defines the direction and / or range of a filter has a direction and range that define a region having an in-band gain / attenuation coefficient based on the sound source direction parameter being within the region, an out-of-band gain / attenuation coefficient based on the sound source direction parameter being outside the region, a direction and range that define a region having an in-band gain / attenuation coefficient based on the sound source direction parameter being within the region, an out-of-band gain / attenuation coefficient based on the sound source direction parameter being outside the region, and an edge zone gain / attenuation coefficient based on a sound source direction parameter within an edge zone region, and can obtain at least one of a further range that defines an edge zone region.

[0026] An apparatus caused to generate a filter to be applied to a plurality of audio signals, wherein a filter gain / attenuation parameter is generated based on a region related to a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a second sound source energy parameter, the apparatus comprising: generating a first temporal gain / attenuation value based on a temporal average of an average band value of the first sound source energy parameter, wherein the number of times the first sound source direction parameter is present within the region is the number of times present within a defined time; generating a second temporal gain / attenuation value based on a temporal average of an average band value of the second sound source energy parameter, wherein the number of times the second sound source direction parameter is present within the region is the number of times present within a defined time; and generating a combined temporal gain / attenuation value based on a combination of the first temporal gain / attenuation value and the second temporal gain / attenuation value to generate a combined temporal gain / attenuation value.

[0027] An apparatus for generating a filter to be applied to a plurality of audio signals, wherein a filter gain / attenuation parameter is generated based on regions related to a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a second sound source energy parameter, the apparatus can perform the steps of: generating a combined frame average value based on a combination of the frame-averaged first sound source energy parameter and the frame-averaged second sound source energy parameter; and generating a frame-smoothed gain / attenuation based on the combined frame average value and the number of times the first and second sound source direction parameters are within a filter region over a frame period.

[0028] An apparatus for generating a filter to be applied to a plurality of audio signals, wherein a filter gain / attenuation parameter is generated based on regions related to a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a second sound source energy parameter, the apparatus can perform the step of generating a filter gain / attenuation for a band based on a combination of a frame-smoothed gain / attenuation, a combined time gain / attenuation value, and a combined band gain / attenuation value.

[0029] Processing of the plurality of audio signals can be configured to provide one or more modified audio signals based on the plurality of audio signals, and an apparatus for determining a second sound source direction parameter and a second sound source energy parameter based on processing of the plurality of audio signals in one or more frequency bands of the plurality of audio signals can determine the second sound source direction parameter and the second sound source energy parameter based on the modified audio signals in one or more frequency bands of the plurality of audio signals.

[0030] A device caused to provide one or more modified audio signals based on a plurality of audio signals further performs a step of generating a plurality of modified audio signals based on modifying the plurality of audio signals using a projection of a first sound source defined by a first sound source direction parameter. A device that determines at least a second sound source direction parameter, at least partially based on at least one or more of the modified audio signals, in one or more frequency bands of the plurality of audio signals processes the plurality of modified audio signals to determine at least the second sound source direction parameter in one or more frequency bands of the plurality of audio signals.

[0031] A device that obtains a region defining a direction and / or range of a filter can cause the region to be obtained based on user input.

[0032] According to a fourth aspect, there is provided a device comprising means for obtaining a plurality of audio signals from respective ones of a plurality of microphones, means for determining a first sound source direction parameter and a first sound source energy parameter based on processing of the plurality of audio signals in one or more frequency bands of the plurality of audio signals, means for determining a second sound source direction parameter and a second sound source energy parameter based on processing of the plurality of audio signals in one or more frequency bands of the plurality of audio signals, means for obtaining a region defining a direction and / or range for a filter, and means for generating the filter to be applied to the plurality of audio signals. Here, the filter gain / attenuation parameter is generated based on the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the region regarding the second sound source energy parameter.

[0033] According to a fifth aspect, the apparatus is caused to at least perform: obtaining a plurality of audio signals from respective plurality of microphones; determining a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the plurality of audio signals based on processing of the plurality of audio signals; determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the plurality of audio signals based on processing of the plurality of audio signals; obtaining a region defining a direction and / or a range for a filter; and generating the filter to be applied to the plurality of audio signals. Here, A filter gain / attenuation parameter is generated based on the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and a region related to the second sound source energy parameter.

[0034] According to a sixth aspect, there is provided a non-transitory computer-readable medium including program instructions for causing an apparatus to at least perform: obtaining a plurality of audio signals from respective plurality of microphones; determining a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the plurality of audio signals based on processing of the plurality of audio signals; determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the plurality of audio signals based on processing of the plurality of audio signals; obtaining a region defining a direction and / or a range for a filter; and generating the filter to be applied to the plurality of audio signals. Here, the filter gain / attenuation parameter is generated based on the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and a region related to the second sound source energy parameter.

[0035] According to a seventh aspect, an acquisition circuit configured to acquire a plurality of audio signals from a plurality of microphones respectively, and in one or more frequency bands of the plurality of audio signals, a determination circuit configured to determine a first sound source direction parameter and a first sound source energy parameter based on the processing of the plurality of audio signals; a determination circuit configured to determine a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the plurality of audio signals based on the processing of the plurality of audio signals; an acquisition circuit configured to acquire a region defining a direction and / or a range for a filter; and a generation circuit configured to generate the filter to be applied to the plurality of audio signals. Here, the filter gain / attenuation parameter is generated based on the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the region regarding the second sound source energy parameter,

[0036] According to an eighth aspect, there is provided a computer-readable medium including program instructions for causing a device to at least execute acquiring a plurality of audio signals from a plurality of microphones respectively, determining a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the plurality of audio signals based on the processing of the plurality of audio signals, determining a second sound source direction parameter and a second sound source energy parameter based on the processing of the plurality of audio signals, acquiring a region defining a direction and / or a range for a filter, and generating a filter to be applied to the plurality of audio signals. Here, the filter gain / attenuation parameter is generated based on the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the region regarding the second sound source energy parameter.

[0037] The device of the present application includes means for performing the operations as described above.

[0038] The apparatus of the present application is configured to execute the operations of the method as described above.

[0039] The computer program of the present application includes program instructions for causing a computer to execute the method as described above.

[0040] A computer program product stored on a medium can cause the apparatus to execute the method described herein.

[0041] An electronic device can include an apparatus as described herein.

[0042] A chipset can include the apparatus described herein.

[0043] Embodiments of the present application aim to address problems related to the state of the art. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] To better understand the present application, reference is now made, by way of example, to the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

[0045] The concepts described in more detail herein with respect to the following embodiments relate to the capture of an audio scene. For example, the following embodiments can be implemented within a capture device side configured to determine object / source related audio signals. For example, in some embodiments, two source direction estimates and their associated direct ambient energy ratios for a sector / zone of interest can be used in determining filter gains / attenuations to "filter" the object / source related audio signals. This spatial filtering can be used instead of (or in addition to) conventional beamforming to generate an object audio signal. The following embodiments describe filter gain parameters, but these same approaches can be used to generate filter attenuation parameters.

[0046] Furthermore, the following embodiments can also be implemented within a playback device where the captured audio is processed by "zoom" or "focus". Further, the spatial filtering can be implemented as any part of the spatial audio signal synthesis operation.

[0047] In the following description, the term sound source is used to describe a defined (artificial or actual) element within a sound field (or audio scene). The term sound source can also be defined as an audio object or audio source, and these terms are interchangeable with respect to understanding the example implementations described herein.

[0048] Embodiments herein relate to parametric audio capture devices and methods such as spatial audio capture (SPAC) technology. For each time-frequency tile, the device is configured to estimate the direction of the dominant sound source and the relative energy of the direct and ambient components of the sound source, which are represented as a direct-to-total energy ratio.

[0049] The following examples are suitable for devices having a difficult microphone configuration or configuration as found in typical mobile devices, and the dimensions of a mobile device typically include at least one short (or thin) dimension relative to other dimensions. In the examples shown herein, The captured spatial audio signal is a suitable input for a spatial synthesizer to generate a spatial audio signal such as a binaural format audio signal for headphone listening or to generate a multi-channel signal format audio signal for loudspeaker listening.

[0050] In some embodiments, these examples can be implemented as part of a spatial capture front end for an immersive voice and audio service (IVAS) standard codec by generating IVAS-compatible audio signals and metadata.

[0051] An audio scene (spatial audio environment) can be complex and can comprise several simultaneous audio or sound sources having different spectral characteristics. Additionally, strong background noise can make it difficult to determine the direction of the sound sources. This can cause problems when filtering the audio technology field (represented by the captured audio signals), which also means that audio elements within the audio technology field that should be filtered (or attenuated) from the audible sound field can leak into the processed output due to insufficient accuracy and reliability of the spatial audio analysis.

[0052] Furthermore, real-world audio recording situations such as simultaneous sound sources, echoes, and ambient sound environments often make it difficult to amplify and / or attenuate the desired sound direction with good audio quality. Typically, in spatial audio capture methods, only a single direction estimate per frequency band is determined and passed to the filter. Therefore, it can be difficult or virtually impossible to distinguish and thus amplify / attenuate the audio signal components associated with two simultaneous sound directions within the same frequency band. Since the direction of at least one of the two simultaneous audio sources remains unknown, there can be additional problems for so-called audio zoom or audio focusing algorithms, the purpose of which is to amplify the audio signal components (sounds) arriving only from a specified direction and attenuate other directions. The "unknown" sound source direction can be located at or near the zoom direction but cannot be amplified without an appropriate DOA estimate. Correspondingly, an efficient attenuation amount for other directions requires the DOA estimates of both sound sources; otherwise, the algorithm can accidentally attenuate other sound sources located in other directions far from the zoom direction based on a single DOA estimate of another sound source located in the other direction.

[0053] The embodiments described herein aim to improve a method by which a sound source can be amplified and / or attenuated as required by a user by implementing an improved two - direction estimation method for each frequency band. The estimation method provides additional information about the audio environment for filtering and the sound source direction. In other words, it provides (a plurality of) two direction estimation values and their direct surrounding energy ratios for each sub - band, enabling more efficient spatial filtering. The increased efficiency is based on combining the calculated filtering gains corresponding to both (all) DOA estimation values and their energy ratios. This, in turn, increases and enhances the perceived audio zoom effect and enables the audio zoom to be used in more complex sound environments with respect to the number and position of sound sources. The embodiments further aim to improve the perceived audio quality due to an improved derivation of the filtering gain / attenuation amount. The improvement results from being able to take into account the DOA estimation values of at least one previous frame (e.g., DOA estimation values from the last 40 frames) and (all) the energy ratios in both directions when forming the filtering gain for the current time - frame.

[0054] Accordingly, the embodiments aim to prevent "interference" filter leakage to the output from the direction that should be filtered or attenuated. Thus, this enhances the perceived audio zoom effect and prevents confusing the user experience when several sound sources are present within the capture. Further, the target (focus) direction can be efficiently amplified relative to other sound directions in a complex environment, again enhancing the zoom - effect experience.

[0055] Accordingly, the embodiments described herein relate to parametric spatial audio capture using multiple microphones. Further, at least two direction and energy - ratio parameters are estimated for each time - frequency tile based on the audio signals from the multiple microphones.

[0056] In these embodiments, in order to achieve an improvement in the plurality of sound source direction detection accuracies, when estimating the second direction, the effect of the first estimated direction is considered. This can lead to an improvement in the perceived quality of the synthesized spatial audio in some embodiments.

[0057] Thus, it is possible to use a similar technique as described in EP3791605, but it can be implemented as described herein.

[0058] In practice, the embodiments described herein generate estimates of sound sources that are perceived to be more spatially stable and more accurate (with respect to their correct or actual positions).

[0059] With respect to FIG. 1, a schematic diagram of an apparatus suitable for implementing the embodiments described herein is shown.

[0060] In this example, an apparatus comprising a microphone array 101 is shown. The microphone array 101 comprises a plurality (two or more) of microphones configured to capture audio signals. The microphones within the microphone array can be of any suitable microphone type, arrangement, or disposition. The microphone audio signal 102 generated by the microphone array 101 can be passed to a spatial analyzer 103.

[0061] The hon device can comprise a spatial analyzer 103 configured to receive or otherwise obtain the microphone audio signal 102 and configured to spatially analyze the microphone audio signal to determine at least two dominant sounds or audio sources for each time - frequency block.

[0062] The spatial analyzer can be the CPU of a mobile device or a computer in some embodiments. The spatial analyzer 103 is configured to generate a data stream that includes an audio signal as well as metadata of the analyzed spatial information 104.

[0063] Depending on the use case, the data stream can be stored or compressed and sent to another location.

[0064] The apparatus further comprises a spatial synthesizer 105. The spatial synthesizer 105 is configured to acquire a data stream that includes an audio signal and metadata. In some embodiments, the spatial synthesizer 105 is implemented within the same apparatus as the spatial analyzer 103 (as shown in FIG. 1 herein), but in some embodiments, it can further be implemented within a different apparatus or device.

[0065] The spatial synthesizer 105 can be implemented within a CPU or a similar processor. The spatial synthesizer 105 is configured to generate an output audio signal 106 based on the audio signal and associated metadata from the data stream 104.

[0066] Furthermore, depending on the use case, the output signal 106 can be in any suitable output format. For example, in some embodiments, the output format is a binaural headphone signal (similarly, the output device presenting the output audio signal is a set of headphones / earphones or the like), or a multi-channel loudspeaker audio signal (similarly, the output device is a set of speakers). The output device 107 (which can be, for example, headphones or loudspeakers as described above) can be configured to receive the output audio signal 106 and present the output to a listener or user.

[0067] These operations of the exemplary apparatus shown in FIG. 1 can be illustrated by the flowchart shown in FIG. 2. Thus, the operations of the exemplary apparatus are summarized as follows.

[0068] In step 201, a microphone audio signal as shown in FIG. 2 is acquired.

[0069] The microphone audio signal is spatially analyzed to generate a spatial audio signal and metadata including the directions and energy ratios of the first and second audio sources for each time-frequency tile as shown in FIG. 2 by step 203.

[0070] Spatial synthesis is applied to the spatial audio signal to generate an appropriate output audio signal as shown in FIG. 2 by step 205.

[0071] In step 207, the output audio signal is output to an output device as shown in FIG. 2.

[0072] In some embodiments, the spatial analysis can be used in connection with an IVAS codec. In this example, the spatial analysis output is in an IVAS-compatible MASA (Metadata-Assisted Spatial Audio) format that can be supplied directly to the IVAS encoder. The IVAS encoder generates an IVAS data stream. At the receiving end, the IVAS decoder can directly generate the desired output audio format. In other words, in such embodiments, there is no separate spatial synthesis block.

[0073] The spatial analyzer shown in FIG. 1 by reference numeral 103 is shown in further detail with respect to FIG. 3.

[0074] In some embodiments, the spatial analyzer 103 comprises a stream (carrier) audio signal generator 307. The stream audio signal generator 307 is configured to receive the microphone audio signal 102 and generate a stream audio signal 308 that is passed to the multiplexer 309. The audio stream signal is generated from the input microphone audio signal based on any suitable method. For example, in some embodiments, one or two microphone signals may be selected from the microphone audio signal 102. Alternatively, in some embodiments, the microphone audio signal 102 may be downsampled and / or compressed to generate the stream audio signal 308.

[0075] In the following example, the spatial analysis is performed in the frequency domain, but it is understood that in some embodiments the analysis can also be performed in the time domain using a time domain sampled version of the microphone audio signal.

[0076] In some embodiments, the spatial analyzer 103 comprises a time-frequency converter 301. The time-frequency converter 301 is configured to receive the microphone audio signals 102 and convert them into the frequency domain. In some embodiments, prior to conversion, the time-domain microphone audio signals can be represented as s i (t), where t is the time index and i is the microphone channel index. The conversion to the frequency domain can be performed by any suitable time-frequency conversion such as STFT (Short-time Fourier transform) or QMF (Quadrature mirror filter). The resulting time-frequency domain microphone signal 302 is S iIt is represented as (b, n). Here, i is the microphone channel index, b is the frequency bin index, and n is the time frame index. The value of b is in the range of 0, .., B - 1, where B is the number of bin indices at each time index n.

[0077] The frequency bins can further be grouped into sub - bands k = 0, .., K - 1. Each sub - band consists of one or more frequency bins. Each sub - band k has a lowest bin b k,low and a highest bin b k,high The width of the sub - band is typically selected based on the characteristics of human hearing, for example, the equivalent rectangular bandwidth (ERB) or Bark scale can be used.

[0078] In some embodiments, the spatial analyzer 103 includes a first - direction analyzer 303. The first - direction analyzer 303 is configured to receive the time - frequency domain microphone audio signal 302 and generate an estimated value of the first sound source for each time - frequency partition of the (first) first direction 314 and the (first) first ratio 316.

[0079] The first - direction analyzer 303 is configured to generate an estimated value for the first direction based on any suitable method such as SPAC (as described in more detail in US9313599).

[0080] In some embodiments, for example, the most dominant direction in the time frame index is the time shift τ that maximizes the correlation between two (microphone audio signal) channels of sub - band k k is estimated by searching. S i S(b, n) is shifted by τ samples as

Number

Number

[0081] In the above formula, the "optimal" delay is searched between microphones 1 and 2. Re represents the real part of the result, and * represents the complex conjugate of the signal. Delay search range parameter D max is defined based on the distance between the microphones. In other words, τ k The value of is searched only within the physically possible range considering the distance between the microphones and the speed of sound.

[0082] Next, the angle in the first direction is

Number

[0083] As shown, the uncertainty of the sign of the angle still exists. Above, the direction analysis between microphone 1 and microphone 2 was defined. Next, the same procedure can be repeated between other microphone pairs to resolve the ambiguity (and / or obtain the direction with reference to another axis). In other words, using information from other analysis pairs,

Number

[0084] For example, when a microphone array includes three microphones, the first microphone, the second microphone, and the third microphone are arranged in a configuration where there is a first pair of microphones (the first microphone and the third microphone) spaced apart by a distance along a first axis, and a second pair of microphones (the first microphone and the second microphone) spaced apart by a distance along a second axis (in this example, the first axis is perpendicular to the second axis). Further, in this example, the three microphones can be on the same third axis defined as being perpendicular to the first and second axes (perpendicular to the plane of the paper on which the figure is printed). Analysis of the delay between the second pair of microphones yields two alternative angles, α and -α. Using the analysis of the delay between the second pair of microphones, it is possible to determine which of the alternative angles is correct. In some embodiments, the information required from this analysis is whether the sound arrives at microphone 1 or 3 first. If the sound reaches microphone 3, the angle α is correct. Otherwise, -α is selected.

[0085] Further, based on inferences between some pairs of microphones, the first spatial analyzer can

Number

[0086] In some embodiments with limited microphone configurations or arrangements, for example where only two microphones are present, the ambiguity in direction cannot be resolved. In such embodiments, the spatial analyzer is configured to define that all sources are always in front of the device. This situation is the same when there are three or more microphones, but their positions do not allow, for example, back analysis.

[0087] Although not disclosed herein, multiple pairs of microphones on the vertical axis can determine elevation and azimuth estimates.

[0088] The first direction analyzer 303 can further, for example, [Number] use the correlation value after normalizing it to determine or estimate the energy ratio corresponding to the angle.

[0089] The value is from -1 to 1 and is typically further restricted to 0 to 1.

[0090] In some embodiments, the first direction analyzer 303 is configured to generate a modified time-frequency microphone audio signal 304. The modified time-frequency microphone audio signal 304 is one from which the first sound source component is removed from the microphone signal.

[0091] Thus, for example, with respect to the first microphone pair (microphones 1 and 2). For each sub-band k, the delay that provides the highest correlation for sub-band k, the second microphone signal is the shifted sample by which the second microphone signal is shifted to obtain the shifted second microphone signal.

[0092] The estimated value of the sound source component is the average of these time-aligned signals [Number] and can be determined as.

[0093] In some embodiments, any other suitable method for determining the sound source component can be used.

[0094] (For example, in the formula of the above example) When the estimated value of the sound source component is determined, this can be removed from the microphone audio signal. On the other hand, the simultaneous sound sources are not in phase, and thus the simultaneous sound sources are attenuated. In this way, the (shifted and non-shifted) microphone signals [Number] can be reduced from. Further, the shifted and corrected microphone audio signal is shifted back to [Number] , samples [Number] to obtain

[0095] These corrected signals [Number] and [Number] can then be passed to the second direction analyzer 305.

[0096] In some embodiments, the spatial analyzer 103 comprises the second direction analyzer 305. The second direction analyzer 305 is configured to estimate the time-frequency microphone audio signal 302, the corrected time-frequency microphone audio signal 304, the first direction 314, and the first ratio 316, and generate estimates of the second direction 324 and the second ratio 326.

[0097] The estimation of the second direction parameter value can adopt the same sub-band structure as the first direction estimation, and can follow the same operations as described above for the first direction estimation.

[0098] Thus, it is possible to estimate the second direction parameter. In such embodiments, the corrected time-frequency microphone audio signal 304 [Number] and [Number] Rather, it is used to determine the direction estimate, rather than the time-frequency microphone audio signal 302.

[0099] Furthermore, in some embodiments, although the energy ratios are limited, the sum of the first and second ratios should not be more than two.

[0100] In some embodiments, the second ratio is

Number

Number

[0101] In the above example, since there are several microphone pairs, the modified signal must be calculated separately for each pair, i.e.,

Number

[0102] The first direction estimate value 314, the first ratio estimate value 316, the second direction estimate value 324, and the second ratio estimate value 326 are passed to a multiplexer (mux) 309 configured to generate the data stream 104 by combining the estimate values and the stream audio signal 308.

[0103] Regarding FIG. 4, a flowchart summarizing the exemplary operation of the spatial analyzer shown in FIG. 3 is shown.

[0104] The microphone audio signal is obtained as shown in FIG. 4 by step 401.

[0105] Next, in step 402, as shown in FIG. 4, a stream audio signal is generated from the microphone audio signal.

[0106] The microphone audio signal can further be time-frequency domain transformed, as shown in FIG. 4, in step 403.

[0107] Next, in step 405, as shown in FIG. 4, the first direction and the first ratio parameter estimated values can be determined.

[0108] Next, in step 407, as shown in FIG. 4, the time-frequency domain microphone audio signal can be modified (the first source component removed).

[0109] Next, in step 409, as shown in FIG. 4, the modified time-frequency domain microphone audio signal is analyzed to determine the second direction and the second ratio parameter estimated values.

[0110] Next, in step 411, as shown in FIG. 4, the first direction, the first ratio, the second direction, and the second ratio parameter estimated values and the stream audio signal are multiplexed to generate a data stream (which can be a MASA format data stream).

[0111] In the following example, a spatial filtering method and apparatus are described in which several gain parameters are determined or calculated and set to adjust a filtering process. These gains can be divided into band-by-band gains, history-based (temporal) gains, and frame-based smoothing gains.

[0112] In the following example, two estimated directions of arrival (DOAs) for each sub-band give a direct-to-ambient (DA) ratio estimate, which basically indicates what portion of the corresponding direction estimate is considered the "direct" signal part and what portion is considered the "ambient" signal part. In these examples, the term direct refers to signals arriving directly from the sound source, and ambient refers to echoes and background noise present in the environment. The direct and ambient components of the signal for each sub-band b can have a range of [0,1], [Number] are defined as follows.

[0113] In some embodiments, the method starts after obtaining the direction and range of a spatial filtering zone (which can also be defined as a sector of interest or a zoom sector of focus) by checking whether either or both of the two direction estimates are not located inside the sector of interest through the sub-bands. In the following example, the spatial filtering is positive notch filtering where the audio signal within the sector of interest is increased relative to the audio signal outside the sector of interest. However, in some embodiments, the spatial filtering is negative notch filtering and the audio signal within the sector of interest is decreased compared to the audio signal outside the sector of interest. The difference between the two is whether the sector gain is greater than the out-of-sector gain that results in a positive spatial notch filter, or whether the sector gain is less than the out-of-sector gain that results in a negative spatial notch filter.

[0114] Simplified diagrams of these three main scenarios are shown with respect to Figure 5.

[0115] In this example, the sound is amplified within the sector and attenuated outside the sector, but the processing is also significantly affected by the DA ratio of the direction estimation.

[0116] For example, the DA ratio estimated value can be considered as a weight for the actual direction estimated value. The numbers in the following table are merely examples for demonstrating the basic principles of their effects on deriving the filter example gain G(b). The first two columns show the case where either of the two sources is estimated as ambient-like sound, which means that the direction estimate should not be used as such for filtering.

Table 1

[0117] Therefore, a low DA ratio value can indicate that the corresponding direction estimate may not be caused by the actual sound source, and in some cases, there is no active direct sound source during capture, or there is only one sound source. In some embodiments, the sector edge can also have a region where the applied subband gain is linearly smoothed to avoid a sharp gain change at the sector edge.

[0118] Therefore, as shown in FIG. 5, there is a first scenario 501 where both sound sources are within the sector, and as a result, filtering gains corresponding to each direction estimate g1(b) occur, both g2(b) being greater than 1, and thus, the spatial gain G(b) results in a value greater than 1.

[0119] A second scenario 503 is shown, where one of the sound sources is within the sector filtering gain corresponding to one direction estimate (the first g1(b)), and the other (the second g2(b)) is greater than 1, and thus, the spatial gain G(b) results in a value approximating 1.

[0120] Furthermore, a third scenario 505 is shown, where both sound sources are outside the sector, and as a result, filtering gains corresponding to each direction estimate g1(b) are obtained, g2(b) being less than 1, and thus, the spatial gain G(b) becomes a value less than 1.

[0121] In some embodiments, the energy of sub - band b of the input signal spectrum X(b) before any energy adjustment can be estimated as follows.

Number

Number

Number

[0122] In some embodiments, the band gain is derived for each sub - band b based on the direction estimates d1 and d2 of the band. The direction - estimate values can be located in the inner side of the focus sector, the outer side of the focus sector, or in the region near the sector edge (so - called edge zone). The direct - energy component for the first direction estimate d1 for sub - band b can be modified as follows.

Number

Number

Number

Number

[0123] For the band b after energy adjustment, the target energy initialized to 0 before the first frame can be defined as follows.

Number

Number

[0124] To take into account the second - direction estimation d2, the g2(b) gain value is calculated in the same way as the g1(b) value, and then the gain is multiplied to obtain the overall band - gain

Number

[0125] Furthermore, in some embodiments, to smooth the filtering gain over time, the time - filtering gain is calculated for each sub - band for both direction estimations d1 and d2. This prevents unnatural pumps or notches from occurring in the overall filter gain. In many cases, the estimated sound - source DA ratio value can vary across sub - bands, so averaging the DA ratio over the entire filtering frequency range provides a good estimate of how much the sound environment is in the ambient environment at the current - time frame f. The ratio average value is calculated for each frame for the first - direction estimation as follows.

Number

Number

Number

Number

[0126] When the history segment is filled with such flags, the number of "true" flags in each sub - band b of d1, N1T(b) is used to obtain a temporary scaling variable

Number

Number

[0127] The number of in-sector direction estimations in each sub-band b at past N1T(b) is

Number

[0128] The time gain for the direction estimation value d2 is calculated in the same way as that for d1, and the actual time filter gain is obtained by multiplication

Number

[0129] In some embodiments, the direction estimation across all sub-bands within a single time frame can vary significantly depending on the number and type of sound sources present within the sound environment. Therefore, an additional frame smoothing gain is required to smooth the spectrum in order to prevent sudden pumps and notches within the spectral envelope in each frame. First, the sum of the ratio means of d1 and d2 can be calculated as

Number

Number

Number

[0130] The previously derived attenuation state is used to calculate the actual filter smoothing gain for each sub - band

Number

Number

Number

[0131] Once all different gain types such as band gain, time gain, and frame gain are calculated, the actual output filter gain can be determined or calculated for each sub - band b, as

Number

[0132] An example of the advantages of implementing the embodiments described in this specification is shown in FIG. 6. Specifically, FIG. 6 shows the output signal levels in dB of a known spatial filter that uses only single-direction estimation for each sub-band 601, and shows a spatial filter approach according to some embodiments 603. In this example, the audio focus direction is directly set to the front of the device, and the signal consists of a speaker that first speaks in front of the device, then moves behind the device at the center of the signal, and finally returns to the front of the device again. Further, the music is reproduced from a speaker located on the left side of the capture device. On average, it can be seen that the embodiments amplify the audio from the front by about 2-3 dB compared to known methods.

[0133] In addition, the embodiments also attenuate the audio from behind the device by 2-3 dB more compared to known spatial filtering methods, which means that the embodiments increase the overall focus effect gain by an average of 4-6 dB overall. This is a clearly audible and significant difference that, in most cases, improves the perceived audio zoom experience. As long as the direction estimations d1 and d2 can be estimated from the capture, the spatial filter can always improve its performance compared to the case where it has only the estimation d1.

[0134] Regarding FIG. 7, an overview of the operation of the embodiments described in this specification is shown.

[0135] The first operation is, as shown in step 701 in FIG. 7, to calculate or determine the direction estimation values of d1 and d2 for sub-band b.

[0136] Next, as shown in step 703 in FIG. 7, a first check can be performed to determine whether d1 is within the sector.

[0137] If d1 is within the sector, as shown in step 705 in FIG. 7, a further check can be made to determine whether d2 is within the sector.

[0138] If both d1 and d2 are within the sector, sub-band b is amplified according to the DA ratio of the associated estimated values of both d1 and d2 as shown in FIG. 707.

[0139] If d1 is not within the sector, further checks can be made in step 709 to determine whether d2 is within the sector as shown in FIG. 7.

[0140] If d1 is within the sector but d2 is not, or if d1 is not within the sector but d2 is within the sector, sub-band b is amplified according to the DA ratio of the in-sector estimate, and sub-band b can be attenuated according to the DA ratio of the out-of-sector estimate as shown in FIG. 7 by step 711.

[0141] If both d1 and d2 are outside the sector, sub-band b is attenuated according to the DA ratio of the associated estimated values of both d1 and d2 as shown in FIG. 713. With respect to FIG. 8, a flowchart showing the generation of gain according to some embodiments is shown.

[0142] Thus, in some embodiments, the band gain g(b) is calculated in both directions in step 801 as shown in FIG. 8.

Number

[0143] Next, in some embodiments, the band gains are multiplied together in step 803 to generate the combined band gain as shown in FIG. 8.

Number

[0144] Next, in step 805, time gains g1 t (b), g2 t (b) are generated for each sub-band.

[0145] Next, the temporal gain can be multiplied together to generate the combined temporal gain as shown in FIG. 8 by step 807. [Number]

[0146] Next, the frame smoothing gains g1 s (b), g2 s (b) can be determined for each subband and direction as shown in FIG. 8 by step 809.

[0147] Next, the frame smoothing gain can be multiplied together to generate the combined frame smoothing gain as shown in FIG. 8 by step 811. [Number]

[0148] Next, as shown in FIG. 8 by step 813, the combined frame smoothing gain, combined temporal gain, and combined band gain [Number] can be multiplied to generate the overall filter gain for subband b.

[0149] With respect to FIG. 9, an exemplary spatial synthesizer 105 as shown in FIG. 1 is shown.

[0150] The spatial synthesizer 105 includes a demultiplexer 1201 in some embodiments. The demultiplexer (Demux) 1201 receives the data stream 104 in some embodiments and separates the data stream into a stream audio signal 1208 and spatial parameter estimates such as a first direction 1214 estimate, a first ratio 1216 estimate, a second direction 1224 estimate, and a second <ratio> estimate 1226.

[0151] These are then passed to the spatial processor / synthesizer 1203.

[0152] The spatial synthesizer 105 comprises a spatial processor / synthesizer 1203, is configured to receive the estimated values and the stereo audio signal, and render an output audio signal. The spatial processing / synthesis can be any suitable two-way based synthesis as described in EP3791605.

[0153] Figures 10 and 11 show end-to-end implementations of embodiments. With respect to Figure 10, it is shown that there are a capture device 1101 and a playback device 1111 that communicate via a transport / storage channel 1105.

[0154] The capture device 1101 is configured as described above and is configured to transmit filtered audio 1109. Additionally, filter orientation / range information 1107 can be received from the playback device 1111.

[0155] With respect to Figure 11, a capture device 1101 is shown that is configured to transmit unfiltered audio 1119 received by the playback device 1111. The playback device comprises a spatial filter 1103 configured to apply spatial filtering as described in the embodiments described herein.

[0156] With respect to Figure 12, an exemplary electronic device that can be used as a computer, an encoder processor, a decoder processor, or any of the functional blocks described herein is shown. The device can be any suitable electronic device or apparatus. For example, in some embodiments, the device 1600 is a mobile device, a user equipment, a tablet computer, a computer, an audio playback device, etc.

[0157] In some embodiments, device 1600 comprises at least one processor or central processing unit 1607. The processor 1607 can be configured to execute various program codes, such as the methods described herein.

[0158] In some embodiments, device 1600 comprises a memory 1611.

[0159] In some embodiments, at least one processor 1607 is coupled to the memory 1611. The memory 1611 can be any suitable storage means. In some embodiments, the memory 1611 comprises a program code section for storing program code executable on the processor 1607. Further, in some embodiments, the memory 1611 can further comprise a stored data section for storing data, such as data processed or to be processed according to the embodiments described herein. The executed program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1607 via the memory-processor coupling as needed.

[0160] In some embodiments, device 1600 comprises a user interface 1605. The user interface 1605 can be coupled to a processor 1607 in some embodiments. In some embodiments, the processor 1607 can control the operation of the user interface 1605 and receive inputs from the user interface 1605. In some embodiments, the user interface 1605 can enable a user to input commands to the device 1600, for example, via a keypad. In some embodiments, the user interface 1605 can enable a user to obtain information from the device 1600. For example, the user interface 1605 may comprise a display configured to display information from the device 1600 to the user. The user interface 1605 can, in some embodiments, comprise a touch screen or touch interface that is capable of both enabling information to be input into the device 1600 and further displaying information to a user of the device 1600.

[0161] In some embodiments, device 1600 comprises an input / output port 1609. In some embodiments, the input / output port 1609 comprises a transceiver. The transceiver in such embodiments is coupled to the processor 1607 and can be configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means can, in some embodiments, be configured to communicate with other electronic devices or devices via a wired or wired - optical connection.

[0162] The transceiver can communicate with additional devices via any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a wireless local area network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth®, or an infrared data communication path (IRDA).

[0163] The transceiver input / output port 1609 can be configured to transmit / receive audio signals, bitstreams, and in some embodiments, execute the operations and methods as described above by using a processor 1607 that executes suitable code.

[0164] Generally, various embodiments of the present invention can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software executable by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Various aspects of the present invention can be illustrated and described as block diagrams, flowcharts, or in some other graphical representation, but the blocks, devices, systems, techniques, or methods contemplated herein are, by way of non-limiting example, implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers, or other computing devices, or any combination thereof, as will be fully understood.

[0165] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device, such as within a processor entity, or by hardware, or by a combination of software and hardware. Further in this regard, note that any block of the logical flow as in the figures can represent a program step, or interconnected logical circuits, blocks and functions, or a combination of program steps and logical circuits, blocks and functions. The software can be stored in a physical medium such as a memory chip, or a memory block implemented within a processor, a magnetic medium, and an optical medium.

[0166] The memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and can include, by way of non-limiting example, one or more of a general-purpose computer, a dedicated computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processor architecture.

[0167] Embodiments of the present invention can be implemented in various components such as integrated circuit modules. The design of integrated circuits is by a large-scale and highly automated process. Complex and powerful software tools are available to convert a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0168] Programs such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design of San Jose, California automatically route conductors and locate components on a semiconductor chip using well-established design rules and a library of pre-stored design modules. Once the design of a semiconductor circuit is complete, the design obtained in a standardized electronic format (e.g., Opus, GDSII, etc.) can be sent to a semiconductor manufacturing facility or "fab" for fabrication.

[0169] The foregoing description has provided a complete and beneficial illustration of exemplary embodiments of the present invention as examples that are illustrative and non-limiting. However, various modifications and adaptations will become apparent to those skilled in the art upon a review of the foregoing description in consideration of the accompanying drawings and the appended claims. However, all such similar modifications of the teachings of the present invention are still included within the scope of the present invention as defined by the appended patent claims.

Claims

1. An apparatus comprising at least one processor and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured to cause the at least one processor to perform at least:[[]] obtaining a plurality of audio signals from respective plurality of microphones; determining a first sound source direction parameter and a first sound source energy parameter based on processing of the plurality of audio signals in one or more frequency bands of the plurality of audio signals; determining a second sound source direction parameter and a second sound source energy parameter based on processing of the plurality of audio signals in the one or more frequency bands of the plurality of audio signals; obtaining a region defining a direction and / or a range for a filter; generating the filter to be applied to the plurality of audio signals, wherein a filter gain / attenuation parameter is generated based on the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, the second sound source energy parameter and the region, such that the generated filter causes the apparatus to generate a first band gain / attenuation value based on the first sound source direction parameter inside or outside the region, generate a second band gain / attenuation value based on whether the second sound source direction parameter is inside or outside the region, and combine the first band gain / attenuation value and the second band gain / attenuation value to generate a combined band gain / attenuation value. The apparatus

2. The obtained region causes the apparatus to define a direction and range of the region together with an in-band gain / attenuation coefficient based on a sound source direction parameter inside the region, an out-of-band gain / attenuation coefficient based on the sound source direction parameter outside the region, define a direction and range of the region together with an in-band gain / attenuation coefficient based on the sound source direction parameter inside the region, an out-of-band gain / attenuation coefficient based on the sound source direction parameter outside the region, and define a further range of an edge zone region together with an edge zone gain / attenuation coefficient based on the sound source direction parameter inside the edge zone region. ​ causing at least one of them to be obtained The apparatus according to claim 1

3. The generated filter causes the apparatus to generate a first temporal gain / attenuation value based on a time average of the average band value of the first sound source energy parameter and the number of times the first sound source direction parameter is within the region over a defined period generating a second temporal gain / attenuation value based on a time average of the average band value of the second sound source energy parameter and the number of times the second sound source direction parameter exists within the region beyond the defined period causing a combined temporal gain / attenuation value to be generated based on a combination of the first temporal gain / attenuation value and the second temporal gain / attenuation value for generating a combined temporal gain / attenuation value The apparatus according to claim 1

4. The generated filter is applied to the plurality of audio signals the filter gain / attenuation parameter causes the apparatus to generate a combined frame-averaged value based on a combination of a frame-averaged first sound source energy parameter and a frame-averaged second sound source energy parameter generating frame-smoothed gain / attenuation based on the combined frame-averaged value and the number of times the first and second sound source direction parameters are within the region of the filter over a frame period The apparatus according to claim 1

5. The generated filter is to be applied to the plurality of audio signals the filter gain / attenuation parameter causes the apparatus to generate filter gain / attenuation for the frequency band based on a combination of the frame-smoothed gain / attenuation, the combined temporal gain / attenuation value, and the combined band gain / attenuation value The apparatus according to claim 4

6. The step of processing the plurality of audio signals includes causing the apparatus to provide one or more modified audio signals based on the plurality of audio signals causing the apparatus to determine a second sound source direction parameter and a second sound source energy parameter based on processing of the plurality of audio signals in one or more of the frequency bands of the plurality of audio signals causing the apparatus to determine a second sound source direction parameter and a second sound source energy parameter based on the modified audio signals in one or more of the frequency bands of the plurality of audio signals The apparatus according to claim 1

7. The provided one or more modified audio signals cause the apparatus to generate a plurality of modified audio signals based on modifying the plurality of audio signals using a projection of a first sound source defined by the first sound source direction parameter, the apparatus according to claim 6.

8. The provided one or more modified audio signals further cause the apparatus to determine the second sound source direction parameter by processing the plurality of modified audio signals, the apparatus according to claim 7.

9. The obtained region defining the direction and / or range for the filter is based on user input, the apparatus according to claim 1.

10. A method for an apparatus, the method comprising: obtaining a plurality of audio signals from respective ones of a plurality of microphones; determining a first sound source direction parameter and a first sound source energy parameter based on processing of the plurality of audio signals in one or more frequency bands of the plurality of audio signals; determining a second sound source direction parameter and a second sound source energy parameter based on processing of the plurality of audio signals in the one or more frequency bands of the plurality of audio signals; obtaining a region defining a direction and / or range for a filter; generating the filter to be applied to the plurality of audio signals, wherein a filter gain / attenuation parameter is generated based on the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter with respect to the region; including: generating the filter to be applied to the plurality of audio signals, wherein a filter gain / attenuation parameter is generated based on the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the region with respect to the second sound source energy parameter, the step comprising: generating a first band gain / attenuation value based on the first sound source direction parameter being within or outside the region; generating a second band gain / attenuation value based on the second sound source direction parameter being within or outside the region; A step of synthesizing the first band gain / attenuation value and the second band gain / attenuation value to generate a synthesized band gain / attenuation value A method comprising the above

11. The step of obtaining the region defining the direction and / or the range for the filter comprises The direction and range defining the region, together with the in-band gain / attenuation coefficient based on the sound source direction parameter within the region The out-of-band gain / attenuation coefficient based on the sound source direction parameter within the region The direction and range defining the region, together with the in-band gain / attenuation coefficient based on the sound source direction parameter within the region The out-of-band gain / attenuation coefficient based on the sound source direction parameter outside the region A further range defining the edge zone region, together with the edge zone gain / attenuation coefficient based on the sound source direction parameter within the edge zone region Comprising at least one of the above The method according to claim 10

12. The step of generating the filter applied to the plurality of audio signals, wherein The filter gain / attenuation parameter is based on the first sound source direction parameter The first sound source energy parameter The second sound source direction parameter, and The second sound source energy parameter Generated based on the region related to the above, the step comprises Generating a first temporal gain / attenuation value based on the average band value of the first sound source energy parameter and the time average of the number of times the first sound source direction parameter is within the region over a defined period Generating a second temporal gain / attenuation value based on the time average of the average band value of the second sound source energy parameter and the number of times the second sound source direction parameter exists within the region over the defined period Generating a synthesized temporal gain / attenuation value based on a combination of the first temporal gain / attenuation value and the second temporal gain / attenuation value for generating the synthesized temporal gain / attenuation value The method according to claim 10, comprising the above

13. The step of generating the filter applied to the plurality of audio signals, wherein The filter gain / attenuation parameter is based on the first sound source direction parameter The first sound source energy parameter The second sound source direction parameter The second sound source energy parameter Generated based on the region related to the above, the step comprises ​ Generating a synthesized frame-averaged value based on a combination of a frame-averaged first sound source energy parameter and a frame-averaged second sound source energy parameter; Generating a frame smoothing gain / attenuation based on the synthesized frame-averaged value and the number of times the first and second sound source direction parameters are within the region of the filter over a frame period; The method according to claim 11, comprising:

14. The step of generating the filter is applicable to the plurality of audio signals; The method according to claim 13, wherein the step of generating a filter gain / attenuation parameter includes generating a filter gain / attenuation for the frequency band based on a combination of the frame smoothing gain / attenuation, a synthesized temporal gain / attenuation value, and a synthesized band gain / attenuation value.

15. The step of processing the plurality of audio signals includes: Providing one or more modified audio signals based on the plurality of audio signals; Determining a second sound source direction parameter and a second sound source energy parameter based on the processing of the plurality of audio signals in one or more of the frequency bands of the plurality of audio signals; comprising: Determining a second sound source direction parameter and a second sound source energy parameter based on the modified audio signal in one or more of the frequency bands of the plurality of audio signals. The method according to claim 10.

16. The method according to claim 15, wherein the step of providing one or more modified audio signals based on the plurality of audio signals includes generating the modified plurality of audio signals based on modifying the plurality of audio signals using a projection of a first sound source defined by the first sound source direction parameter.

17. The step of providing one or more modified audio signals based on the plurality of audio signals includes: Determining at least a second sound source direction parameter based on at least a part of the one or more modified audio signals in one or more of the frequency bands of the plurality of audio signals; Determining the at least second sound source direction parameter by processing the modified plurality of audio signals in the one or more frequency bands of the plurality of audio signals The method according to claim 16, further comprising.

18. The method according to claim 10, wherein the step of obtaining the region defining the direction and / or the range for the filter includes obtaining the region based on user input.

Citation Information

Patent Citations

  • Audio lens

    US20140348342A1

  • Multi-media content

    WO2021160465A1