Spatial audio filtering within spatial audio capture

By using an improved spatial audio filtering method, multi-microphone signal processing is employed to estimate the direction and energy ratio of the sound source and generate filter gain/attenuation parameters. This solves the problem of inaccurate sound source differentiation and amplification/attenuation in the existing audio environment, thereby improving audio quality and zoom effect.

CN122054036APending Publication Date: 2026-05-15NOKIA TECHNOLOGIES OY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2022-09-29
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately distinguish and amplify/attenuate multiple simultaneous sound sources in complex audio environments, resulting in poor audio zoom effects and decreased audio quality in environments with background noise and multiple sound sources.

Method used

By improving the spatial audio filtering method, based on the processing of signals from two or more microphones, the direction and energy ratio of multiple sound sources are estimated, and filter gain/attenuation parameters are generated to achieve more accurate audio filtering and zoom effects.

Benefits of technology

It improves audio zoom effect and perceived audio quality, prevents leakage of interfering sound sources, and enhances the stability and accuracy of audio experience in multi-sound source environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122054036A_ABST
    Figure CN122054036A_ABST
Patent Text Reader

Abstract

An apparatus comprises means configured to: obtain respective two or more audio signals from two or more microphones; determining a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the two or more audio signals based on the processing of the two or more audio signals; determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the two or more audio signals based on the processing of the two or more audio signals; obtaining a region defining a direction and / or range for the filter; and generating a filter to be applied to the two or more audio signals, where a filter gain / attenuation parameter is generated based on a relationship of the region to the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application No. 202211217607.6, entitled "Spatial audio filtering within spatial audio capture" (filed on September 29, 2022). Technical Field

[0002] This application relates to apparatus and methods for spatial audio filtering within spatial audio capture. Background Technology

[0003] Spatial audio capture is performed using microphone arrays in many modern digital devices, such as mobile devices and cameras, and is often used in conjunction with video capture. The spatial audio capture can be played back through headphones or speakers to provide the user with an experience of the audio scene captured by the microphone array.

[0004] Parametric spatial audio capture methods enable spatial audio capture using different microphone configurations and arrangements, and therefore can be used in consumer devices such as mobile phones. These methods are based on signal processing solutions that analyze the spatial audio field around the device using available information from multiple microphones. Typically, these methods perceptually analyze the microphone audio signals to determine relevant information within the frequency band. This information includes, for example, the direction of the primary sound source (or audio source or audio object) and the relationship between the source energy and the total frequency band energy. Based on this determined information, spatial audio can be reproduced, for example, using headphones or speakers. Ultimately, users or listeners can thus experience ambient audio as if they were present in the audio scene being recorded by the capture device.

[0005] The better the audio analysis and synthesis performance, the more realistic the result will be for the user or listener. Summary of the Invention

[0006] According to a first aspect, an apparatus is provided, comprising components configured to perform the following operations: acquiring corresponding two or more audio signals from two or more microphones; determining a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; obtaining a region defining the direction and / or range for a filter; and generating a filter to be applied to the two or more audio signals, wherein filter gain / attenuation parameters are generated based on the relationship between the region and the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter.

[0007] A component configured to generate a filter to be applied to two or more audio signals, wherein the component generating filter gain / attenuation parameters based on the relationship between a region and a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a second sound source energy parameter can be configured to: generate a first frequency band gain / attenuation value based on whether the first sound source direction parameter is within or outside the region; generate a second frequency band gain / attenuation value based on whether the second sound source direction parameter is within or outside the region; and combine the first frequency band gain / attenuation value and the second frequency band gain / attenuation value to generate a combined frequency band gain / attenuation value.

[0008] The component configured to obtain a region defining the direction and / or range for the filter can be configured to obtain at least one of the following: defining the direction and range of the region, together with in-band gain / attenuation factors within the region and out-of-band gain / attenuation factors outside the region based on the sound source direction parameters; defining the direction and range of the region, together with in-band gain / attenuation factors within the region and out-of-band gain / attenuation factors outside the region based on the sound source direction parameters; and defining a further range of the edge zone region, together with edge zone band gain / attenuation factors within the edge zone region based on the sound source direction parameters.

[0009] A component configured to generate a filter to be applied to two or more audio signals, wherein the component generating filter gain / attenuation parameters based on the relationship between a region and a first source direction parameter, a first source energy parameter, a second source direction parameter, and a second source energy parameter can be configured to: generate a first time gain / attenuation value based on the time average of the average frequency band value of the first source energy parameter within a defined time period and the number of times the first source direction parameter is within the region; generate a second time gain / attenuation value based on the time average of the average frequency band value of the second source energy parameter within a defined time period and the number of times the second source direction parameter is within the region; and generate a combined time gain / attenuation value based on the combination of the first time gain / attenuation value and the second time gain / attenuation value.

[0010] A component configured to generate a filter to be applied to two or more audio signals, wherein the component that generates filter gain / attenuation parameters based on the relationship between a region and a first source direction parameter, a first source energy parameter, a second source direction parameter, and a second source energy parameter can be configured to: generate a combined frame average based on a combination of a frame-averaged first source energy parameter and a frame-averaged second source energy parameter; and generate frame smoothing gain / attenuation based on the combined frame average and the number of times the first and second source direction parameters are within the filter region during a frame period.

[0011] The component configured to generate a filter to be applied to two or more audio signals, wherein the component that generates filter gain / attenuation parameters based on the relationship between a region and a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a second sound source energy parameter can be configured to generate filter gain / attenuation for a frequency band based on a combination of frame smoothing gain / attenuation, combined time gain / attenuation values, and combined frequency band gain / attenuation values.

[0012] The processing of two or more audio signals can be configured to provide one or more modified audio signals based on the two or more audio signals, and wherein the component configured to determine a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the two or more audio signals based on the processing of the two or more audio signals can be configured to: determine the second sound source direction parameter and the second sound source energy parameter in one or more frequency bands of the two or more audio signals based on the modified audio signals.

[0013] The component configured to provide one or more modified audio signals based on two or more audio signals may be further configured to: modify the two or more audio signals based on the projection of a first sound source defined by a first sound source direction parameter to generate two or more modified audio signals; and the component configured to determine at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals based at least in part on the one or more modified audio signals is configured to: determine at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals by processing the modified two or more audio signals.

[0014] A component configured to obtain a region that defines the direction and / or range for a filter can be configured to obtain the region based on user input.

[0015] According to a second aspect, a method for an apparatus is provided, the method comprising: acquiring corresponding two or more audio signals from two or more microphones; determining a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; obtaining a region defining the direction and / or range for a filter; and generating a filter to be applied to the two or more audio signals, wherein filter gain / attenuation parameters are generated based on the relationship between the region and the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter.

[0016] Generating a filter to be applied to two or more audio signals, wherein generating filter gain / attenuation parameters based on the relationship between a region and a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a second sound source energy parameter may include: generating a first frequency band gain / attenuation value based on whether the first sound source direction parameter is within or outside the region; generating a second frequency band gain / attenuation value based on whether the second sound source direction parameter is within or outside the region; and combining the first frequency band gain / attenuation value and the second frequency band gain / attenuation value to generate a combined frequency band gain / attenuation value.

[0017] Obtaining a region that defines the direction and / or range for the filter may include at least one of the following: defining the direction and range of the region, together with in-band gain / attenuation factors within the region and out-of-band gain / attenuation factors outside the region based on the source direction parameters; defining the direction and range of the region, together with in-band gain / attenuation factors within the region and out-of-band gain / attenuation factors outside the region based on the source direction parameters; and defining a further range of the edge zone region, together with edge zone band gain / attenuation factors within the edge zone region based on the source direction parameters.

[0018] Generating a filter to be applied to two or more audio signals, wherein generating filter gain / attenuation parameters based on the relationship between a region and a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a second sound source energy parameter may include: generating a first time gain / attenuation value based on the time average of the average frequency band value of the first sound source energy parameter within a defined time period and the number of times the first sound source direction parameter is within the region; generating a second time gain / attenuation value based on the time average of the average frequency band value of the second sound source energy parameter within a defined time period and the number of times the second sound source direction parameter is within the region; and generating a combined time gain / attenuation value based on a combination of the first time gain / attenuation value and the second time gain / attenuation value.

[0019] Generating a filter to be applied to two or more audio signals, wherein generating filter gain / attenuation parameters based on the relationship between a region and a first source direction parameter, a first source energy parameter, a second source direction parameter, and a second source energy parameter may include: generating a combined frame average based on a combination of a frame-averaged first source energy parameter and a frame-averaged second source energy parameter; and generating frame smoothing gain / attenuation based on the combined frame average and the number of times the first and second source direction parameters are within the filter region during the frame period.

[0020] Generating filters to be applied to two or more audio signals, wherein generating filter gain / attenuation parameters based on the relationship between the region and the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter may include: generating filter gain / attenuation for the frequency band based on a combination of frame smoothing gain / attenuation, combined time gain / attenuation values, and combined frequency band gain / attenuation values.

[0021] Processing two or more audio signals may include providing one or more modified audio signals based on the two or more audio signals, and determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the two or more audio signals based on the processing of the two or more audio signals may include: determining the second sound source direction parameter and the second sound source energy parameter in one or more frequency bands of the two or more audio signals based on the modified audio signals.

[0022] Providing one or more modified audio signals based on two or more audio signals may include: modifying the two or more audio signals based on the projection of a first sound source defined by a first sound source direction parameter to generate two or more modified audio signals; and determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, at least in part, based on the one or more modified audio signals, includes: determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals by processing the modified two or more audio signals.

[0023] Obtaining the region defined for the direction and / or range of the filter can include obtaining the region based on user input.

[0024] According to a third aspect, an apparatus is provided, comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured, together with the at least one processor, to cause the apparatus to at least perform: acquiring corresponding two or more audio signals from two or more microphones; determining a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; obtaining a region defining the direction and / or range for a filter; and generating a filter to be applied to the two or more audio signals, wherein filter gain / attenuation parameters are generated based on the relationship between the region and the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter.

[0025] The means for generating a filter to be applied to two or more audio signals, wherein the means for generating filter gain / attenuation parameters based on the relationship between a region and a first sound source direction parameter, a first sound source energy parameter, a second sound source direction parameter, and a second sound source energy parameter, is capable of: generating a first frequency band gain / attenuation value based on whether the first sound source direction parameter is within or outside the region; generating a second frequency band gain / attenuation value based on whether the second sound source direction parameter is within or outside the region; and combining the first frequency band gain / attenuation value and the second frequency band gain / attenuation value to generate a combined frequency band gain / attenuation value.

[0026] The means for obtaining a region defining the direction and / or range for a filter can be made to obtain at least one of the following: defining the direction and range of the region, together with in-band gain / attenuation factors within the region and out-of-band gain / attenuation factors outside the region based on the sound source direction parameters; defining the direction and range of the region, together with in-band gain / attenuation factors within the region and out-of-band gain / attenuation factors outside the region based on the sound source direction parameters; and defining a further range of the edge zone region, together with edge zone band gain / attenuation factors within the edge zone region based on the sound source direction parameters.

[0027] An apparatus for generating a filter to be applied to two or more audio signals, wherein an apparatus for generating filter gain / attenuation parameters based on the relationship between a region and a first source direction parameter, a first source energy parameter, a second source direction parameter, and a second source energy parameter, is configured to: generate a first time gain / attenuation value based on the time average of the average frequency band value of the first source energy parameter within a defined time period and the number of times the first source direction parameter is within the region; generate a second time gain / attenuation value based on the time average of the average frequency band value of the second source energy parameter within a defined time period and the number of times the second source direction parameter is within the region; and generate a combined time gain / attenuation value based on a combination of the first time gain / attenuation value and the second time gain / attenuation value.

[0028] The means for generating a filter to be applied to two or more audio signals, wherein the means for generating filter gain / attenuation parameters based on the relationship between a region and a first source direction parameter, a first source energy parameter, a second source direction parameter, and a second source energy parameter, can be made to: generate a combined frame average based on a combination of a frame-averaged first source energy parameter and a frame-averaged second source energy parameter; and generate frame smoothing gain / attenuation based on the combined frame average and the number of times the first and second source direction parameters are within the filter region during the frame period.

[0029] The means for generating filters to be applied to two or more audio signals, wherein the means for generating filter gain / attenuation parameters based on the relationship between a region and a first source direction parameter, a first source energy parameter, a second source direction parameter, and a second source energy parameter, can be made to generate filter gain / attenuation for a frequency band based on a combination of frame smoothing gain / attenuation, combined time gain / attenuation values, and combined band gain / attenuation values.

[0030] The processing of two or more audio signals can be configured to provide one or more modified audio signals based on the two or more audio signals, and wherein the means for determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the two or more audio signals based on the processing of the two or more audio signals can be configured to: determine the second sound source direction parameter and the second sound source energy parameter in one or more frequency bands of the two or more audio signals based on the modified audio signals.

[0031] The means of providing one or more modified audio signals based on two or more audio signals may be further made to: modify the two or more audio signals based on the projection of a first sound source defined by a first sound source direction parameter to generate two or more modified audio signals; and the means of determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals based at least in part on the one or more modified audio signals may be made to: determine at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals by processing the modified two or more audio signals.

[0032] The means by which a region is defined for the direction and / or range of a filter can be made to obtain the region based on user input.

[0033] According to a fourth aspect, an apparatus is provided, comprising: means for acquiring corresponding two or more audio signals from two or more microphones; means for determining a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; means for determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; means for obtaining a region defining the direction and / or range for a filter; and means for generating a filter to be applied to the two or more audio signals, wherein filter gain / attenuation parameters are generated based on the relationship between the region and the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter.

[0034] According to a fifth aspect, a computer program [or a computer-readable medium including program instructions] is provided, the instructions being configured to cause a device to perform at least the following operations: acquiring two or more corresponding audio signals from two or more microphones; determining a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; obtaining a region defining the direction and / or range for a filter; and generating a filter to be applied to the two or more audio signals, wherein filter gain / attenuation parameters are generated based on the relationship between the region and the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter.

[0035] According to a sixth aspect, a non-transitory computer-readable medium is provided, comprising program instructions for causing a device to perform at least the following operations: acquiring corresponding two or more audio signals from two or more microphones; determining a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; obtaining a region defining the direction and / or range for a filter; and generating a filter to be applied to the two or more audio signals, wherein filter gain / attenuation parameters are generated based on the relationship between the region and the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter.

[0036] According to a seventh aspect, an apparatus is provided, comprising: an acquisition circuit configured to acquire two or more corresponding audio signals from two or more microphones; a determination circuit configured to determine a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; a determination circuit configured to determine a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; an acquisition circuit configured to acquire a region defining a direction and / or range for a filter; and a generation circuit configured to generate a filter to be applied to the two or more audio signals, wherein filter gain / attenuation parameters are generated based on the relationship between the region and the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter.

[0037] According to an eighth aspect, a computer-readable medium including program instructions is provided for causing a device to perform at least the following operations: acquiring corresponding two or more audio signals from two or more microphones; determining a first sound source direction parameter and a first sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; determining a second sound source direction parameter and a second sound source energy parameter in one or more frequency bands of the two or more audio signals based on processing of the two or more audio signals; obtaining a region defining the direction and / or range for a filter; and generating a filter to be applied to the two or more audio signals, wherein filter gain / attenuation parameters are generated based on the relationship between the region and the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter.

[0038] An apparatus comprising components for performing actions as described above.

[0039] An apparatus configured to perform the actions described above.

[0040] A computer program comprising program instructions for causing a computer to perform the methods described above.

[0041] A computer program product stored on a medium can enable a device to perform the methods described herein.

[0042] An electronic device may include the means as described herein.

[0043] A chipset may include devices as described herein.

[0044] The embodiments of this application are intended to solve problems associated with the prior art. Attached Figure Description

[0045] To better understand this application, reference will now be made to the accompanying drawings by way of example, wherein:

[0046] Figure 1 An example apparatus for implementing spatial capture and playback is illustrated schematically according to some embodiments;

[0047] Figure 2 Illustrations based on some embodiments Figure 1 A flowchart illustrating the operation of the device shown;

[0048] Figure 3 Schematic illustration of some embodiments such as Figure 1 The example spatial analyzer shown;

[0049] Figure 4 Illustrations based on some embodiments Figure 3A flowchart illustrating the operation of the example spatial analyzer is shown below;

[0050] Figure 5 Examples are shown where the sound source is located inside or outside the region of interest.

[0051] Figure 6 Example graphs showing the signal levels of a spatial filter;

[0052] Figure 7 A flowchart is shown illustrating a spatial filtering operation for determining the location of a sound source within a region of interest based on two sound source direction estimates, according to some embodiments.

[0053] Figure 8 A flowchart illustrating spatial filtering based on two sound source direction estimations according to some embodiments is shown;

[0054] Figure 9 Schematic illustration of some embodiments such as Figure 2 The example spatial synthesizer shown;

[0055] Figure 10 and Figure 11 An example device system, including means suitable for implementing embodiments as shown in the preceding figures, is schematically illustrated.

[0056] Figure 12 An example device suitable for implementing the illustrated apparatus is shown schematically. Detailed Implementation

[0057] The concepts discussed in further detail herein with respect to the following embodiments relate to the capture of audio scenes. For example, the following embodiments can be implemented within a capture device configured to determine object / source-related audio signals. For instance, in some embodiments, two source direction estimates and their correlation with respect to the sector / region of interest directly versus the ambient energy ratio can be used to determine filter gain / attenuation to "filter" the object / source-related audio signal. This spatial filtering can be used in place of (or even supplement) traditional beamforming to generate the object audio signal. Filter gain parameters are discussed in the following embodiments, although these same methods can be used to generate filter attenuation parameters.

[0058] Furthermore, the following embodiments can also be implemented within a playback device that processes the captured audio through "zoom" or "focus". Additionally, spatial filtering can be implemented as an optional part of the spatial audio signal synthesis operation.

[0059] In the following description, the term "sound source" is used to describe an element (artificial or real) defined within a sound field (or audio scene). The term "sound source" can also be defined as an audio object or an audio source, and these terms are interchangeable in understanding the implementation of the examples described herein.

[0060] The embodiments described herein relate to parametric audio capture apparatus and methods, such as spatial audio capture (SPAC) techniques. For each time-frequency tile, the apparatus is configured to estimate the direction of the primary sound source and the relative energy of the direct component of the sound source and the ambient component, which is expressed as a direct-to-total energy ratio.

[0061] The following examples are suitable for devices with challenging microphone arrangements or configurations (such as those found in typical mobile devices), where the size of the mobile device typically includes at least one short (or thin) dimension relative to other dimensions. In the examples shown herein, the captured spatial audio signal is a suitable input to a spatial synthesizer to generate spatial audio signals, such as binaural format audio signals for headphone listening or multichannel signal format audio signals for speaker listening.

[0062] In some embodiments, these examples can be implemented as part of a spatial capture front-end for IVAS standard codecs by generating audio signals and metadata compatible with Immersive Speech and Audio Services (IVAS).

[0063] Audio scenes (spatial audio environments) can be complex and include several simultaneous audio sources or sound sources with different spectral characteristics. Furthermore, strong background noise can make it difficult to determine the direction of sound sources. This can cause problems when filtering the audio field (represented by the captured audio signal), meaning that sound elements within the audio field that are expected to be filtered out (or attenuated) from the audible sound field may leak into the processed output due to inaccurate or unreliable spatial audio analysis.

[0064] Furthermore, real-life audio recording scenarios such as simultaneous sound sources, echoes, and ambient sound environments often make amplifying and / or attenuating the desired sound direction with good audio quality challenging. Typically, in spatial audio capture methods, only a single direction estimate is determined for each frequency band and passed to the filter. Therefore, it is difficult, or practically impossible, to distinguish and thus amplify / attenuate audio signal components associated with two simultaneous sound directions present within the same frequency band. Since the direction of at least one of the two simultaneous audio sources remains unknown, so-called audio zoom or audio focus algorithms, where the goal is to amplify audio signal components (sounds) arriving only from a specified direction and attenuate audio signal components (sounds) from other directions, can encounter further problems. The direction of the "unknown" (one or more) sound sources may be located at or near the zoom direction but cannot be amplified without a correct DOA estimate. Correspondingly, effective attenuation in other directions requires DOA estimates for both sound sources; otherwise, the algorithm may inadvertently attenuate another sound source in or near the zoom direction based on a single DOA estimate of another sound source located in a direction far from the zoom direction.

[0065] The embodiments described herein aim to improve how sound sources can be amplified and / or attenuated as requested by the user by implementing an improved two-direction estimation method(s) for each frequency band. The estimation method provides additional information about the audio environment and the direction of the sound sources for filtering. In other words, by providing (multiple) two-direction estimates and their direct energy ratios to the environment for each subband, more effective spatial filtering is achieved. The increased efficiency is based on a calculated filter gain corresponding to the combination of (all) two DOA estimates and their energy ratios. This, in turn, increases and enhances the perceived audio zoom effect, enabling audio zoom to be used in sound environments with greater complexity in terms of the number and location of sound sources.

[0066] The embodiments further aim to improve perceived audio quality due to the improved derivation of filter gain / attenuation. The improvement stems from the ability to take into account at least one previous frame's DOA estimate (e.g., DOA estimate from the last 40 frames) and the energy ratio in both directions when forming the filter gain for the current time frame.

[0067] Therefore, the implementation is designed to prevent "interference" filters from leaking into the output from the direction they should be filtered or attenuated. This enhances the perceived audio zoom effect when multiple sound sources are present in the capture and prevents a confusing user experience. Furthermore, in complex environments, the target (focus) direction can be effectively amplified relative to other sound directions, further enhancing the zoom effect experience.

[0068] Therefore, the embodiments described herein relate to parametric spatial audio capture using two or more microphones. Furthermore, based on the audio signals from the two or more microphones, at least two directional parameters and an energy ratio parameter are estimated in each time-frequency tile.

[0069] In these embodiments, the influence of the first estimated direction is taken into account when estimating the second direction in order to improve the accuracy of multi-source direction detection. In some embodiments, this can lead to an improvement in the perceptual quality of the synthesized spatial audio.

[0070] Therefore, similar techniques, such as those described in EP3791605, can be used, but implemented in the manner described herein.

[0071] In practice, the embodiments described herein produce estimates of sound sources that are perceived as more spatially stable and accurate (relative to their correct or actual locations).

[0072] about Figure 1 The diagram illustrates an apparatus suitable for implementing the embodiments described herein.

[0073] In this example, an apparatus including a microphone array 101 is shown. The microphone array 101 includes a plurality of microphones (two or more) configured to capture audio signals. The microphones within the microphone array can be any suitable microphone type, arrangement, or configuration. The microphone audio signal 102 generated by the microphone array 101 can be transmitted to a spatial analyzer 103.

[0074] The device may include a spatial analyzer 103 configured to receive or otherwise acquire a microphone audio signal 102 and configured to perform spatial analysis on the microphone audio signal to identify at least two primary sound or audio sources for each time-frequency block.

[0075] In some embodiments, the spatial analyzer may be a mobile device or the CPU of a computer. The spatial analyzer 103 is configured to generate a data stream 104 that includes audio signals and metadata of the spatial information being analyzed.

[0076] Depending on the use case, the data stream can be stored or compressed and transmitted to another location.

[0077] Furthermore, the device includes a spatial synthesizer 105. The spatial synthesizer 105 is configured to acquire a data stream including audio signals and metadata. In some embodiments, the spatial synthesizer 105 is implemented within the same device as the spatial analyzer 103 (e.g., Figure 1 (as shown), but in some embodiments it can also be implemented in different devices or equipment.

[0078] Spatial synthesizer 105 can be implemented within a CPU or similar processor. Spatial synthesizer 105 is configured to generate an output audio signal 106 based on the audio signal from data stream 104 and associated metadata.

[0079] Furthermore, depending on the usage, the output signal 106 can be any suitable output format. For example, in some embodiments, the output format is a binaural headphone signal (where the output device presenting the output audio signal is a set of headphones / earbuds or the like) or a multi-channel speaker audio signal (where the output device is a set of speakers).

[0080] Output device 107 (as described above, which may be, for example, headphones or speakers) can be configured to receive output audio signal 106 and present the output to a listener or user.

[0081] Figure 1 These operations of the example device shown can be performed by Figure 2 The flowchart shown is used to illustrate this. Therefore, the operation of the example device is summarized as follows.

[0082] Obtain the microphone audio signal, such as Figure 2 Step 201 is shown in the diagram.

[0083] Spatial analysis is performed on the microphone audio signal to generate spatial audio signals and metadata, which includes the direction and energy ratio for the first and second audio sources for each time-frequency tile, such as... Figure 2 Step 203 is shown in the diagram.

[0084] Spatial synthesis is applied to spatial audio signals to generate suitable output audio signals, such as... Figure 2 Step 205 is shown in the diagram.

[0085] Output the audio signal to the output device, such as Figure 2 Step 207 is shown in the diagram.

[0086] In some embodiments, spatial analysis can be used in conjunction with an IVAS codec. In this example, the spatial analysis output is an IVAS-compatible MASA (Metadata-Assisted Spatial Audio) format, which can be directly fed to the IVAS encoder. The IVAS encoder generates an IVAS data stream. At the receiving end, the IVAS decoder is able to directly produce the desired output audio format. In other words, in this embodiment, there is no separate spatial synthesis block.

[0087] refer to Figure 3 It shows more details Figure 1 The spatial analyzer is shown by reference numeral 103 in the attached figure.

[0088] In some embodiments, the spatial analyzer 103 includes a streaming audio signal generator 307. The streaming audio signal generator 307 is configured to receive microphone audio signals 102 and generate one or more streaming audio signals 308 to be passed to the multiplexer 309. The audio streaming signal is generated from the input microphone audio signals using any suitable method. For example, in some embodiments, one or two microphone signals may be selected from the microphone audio signals 102. Alternatively, in some embodiments, the microphone audio signals 102 may be downsampled and / or compressed to generate the streaming audio signals 308.

[0089] In the following examples, spatial analysis is performed in the frequency domain; however, it should be understood that in some embodiments, analysis may also be performed in the time domain using a temporal sampled version of the microphone audio signal.

[0090] In some embodiments, the spatial analyzer 103 includes a time-frequency converter 301. The time-frequency converter 301 is configured to receive microphone audio signals 102 and convert them to the frequency domain. In some embodiments, the time-domain microphone audio signals may be represented as follows before the conversion: Where t is the time index and i is the microphone channel index. The transformation to the frequency domain can be achieved using any suitable time-frequency transform, such as STFT (Short Time Fourier Transform) or QMF (Quadrature Mirror Filter). The resulting time-frequency domain microphone signal 302 is represented as... Where i is the microphone channel index, b is the frequency bin index, and n is the time frame index. The value of b is in the range of 0, ..., B – 1, where B is the number of bin indices at each time index n.

[0091] Frequency bins can be further combined into subbands k = 0, ..., K – 1. Each subband consists of one or more frequency bins. Each subband k has a minimum number of bins. and the highest warehouse The width of the subband is usually chosen based on the characteristics of human hearing; for example, equivalent rectangular bandwidth (ERB) or Bark scale can be used.

[0092] In some embodiments, the spatial analyzer 103 includes a first direction analyzer 303. The first direction analyzer 303 is configured to receive a time-frequency domain microphone audio signal 302 and generate estimates of a first direction 314 and a first ratio 316 for each time-frequency tile for the first sound source.

[0093] The first direction analyzer 303 is configured to generate an estimate of the first direction based on any suitable method, such as SPAC (as described in more detail in US9313599).

[0094] In some embodiments, for example, by searching for a time shift for subband k that maximizes the correlation between the two (microphone audio signal) channels. To estimate the primary direction used for time frame indexing. Can be shifted The following are examples:

[0095]

[0096] Then, find the delay for each subband k. It maximizes the correlation between the two microphone channels:

[0097]

[0098] In the formula above, the "optimal" delay is searched between microphones 1 and 2. Re indicates the real part of the result, and * is the complex conjugate of the signal. The delay search range parameter is defined based on the distance between the microphones. In other words, considering the distance between microphones and the speed of sound, the search is limited to the physically possible range. The value of .

[0099] The angle in the first direction can then be defined as...

[0100]

[0101] As shown in the figure, the sign of the angle remains uncertain.

[0102] The directional analysis between microphones 1 and 2 has been defined above. A similar process can then be repeated between other microphone pairs to resolve ambiguities (and / or obtain the direction referenced to another axis). In other words, information from other analysis pairs can be used to eliminate ambiguities. The ambiguity of symbols in [the text].

[0103] For example, in a microphone array comprising three microphones, the first, second, and third microphones are arranged such that the first microphone pair (first and third microphones) is separated by a distance on a first axis and the second microphone pair (first and second microphones) is separated by a distance on a second axis (in this example, the first axis is perpendicular to the second axis). Furthermore, in this example, the three microphones may be located on the same third axis, which is defined as perpendicular to both the first and second axes (and perpendicular to the paper on which the diagram is printed). The delay between the second microphone pairs is analyzed, resulting in two alternative angles, α and -α. The delay between the second microphone pairs is then analyzed, and this can be used to determine which alternative angle is correct. In some embodiments, the information required for this analysis is whether the sound arrives at microphone 1 or microphone 3 first. If the sound arrives at microphone 3, then angle α is correct. If not, -α is chosen.

[0104] Furthermore, based on inferences between several microphone pairs, the first spatial analyzer can determine or estimate the correct orientation angle. .

[0105] In some embodiments with limited microphone configurations or arrangements, such as only two microphones, directional ambiguity cannot be resolved. In such embodiments, the spatial analyzer is configured to define that all sources are always in front of the device. The same is true when there are more than two microphones; however, their positions do not allow for, for example, front-to-back analysis.

[0106] Although not disclosed in this paper, multiple pairs of microphones on the vertical axis can determine elevation and azimuth estimates.

[0107] The first-direction analyzer 303 can also use, for example, normalized correlation values. To determine or estimate the angle Corresponding energy ratio ,For example:

[0108]

[0109] The value of is between -1 and 1, and is usually further restricted to between 0 and 1.

[0110] In some embodiments, the first direction analyzer 303 is configured to generate a modified time-frequency microphone audio signal 304. The modified time-frequency microphone audio signal 304 is a signal from which the first sound source component has been removed.

[0111] Therefore, for example, regarding the first microphone pair (microphone 1 and microphone 2), for subband k, the delay that provides the highest correlation is... For each sub-band k, the second microphone signal is shifted. One sample was used to obtain the shifted second microphone signal. .

[0112] The source components can be estimated as the average of these time-aligned signals:

[0113]

[0114] In some embodiments, any other suitable method may be used to determine the sound source components.

[0115] The estimates of the sound source components have already been determined (for example, in the example formula above). Then, it can be removed from the microphone audio signal. On the other hand, other simultaneous sound sources are not in phase, which causes them to be attenuated. Now, it can be subtracted from the (shifted and unshifted) microphone signal. :

[0116]

[0117]

[0118] In addition, the microphone audio signal after shifting modification Displaced backward 1 sample to obtain:

[0119]

[0120] Then, these modified signals and It can be passed to the second direction analyzer 305.

[0121] In some embodiments, the spatial analyzer 103 includes a second direction analyzer 305. The second direction analyzer 305 is configured to receive a time-frequency microphone audio signal 302 estimate, a modified time-frequency microphone audio signal 304 estimate, a first direction 314 estimate, and a first ratio 316 estimate, and generate a second direction 324 estimate and a second ratio 326 estimate.

[0122] The estimation of the second direction parameter value can adopt the same sub-band structure as the first direction estimation and follow similar operations as previously described for the first direction estimation.

[0123] Therefore, the second direction parameter can be estimated. and In this embodiment, a modified time-frequency microphone audio signal 304 is used. and Instead of the time-frequency microphone audio signal 302 and To determine the direction estimate.

[0124] Furthermore, in some embodiments, the energy ratio It is restricted because the sum of the first ratio and the second ratio should not exceed 1.

[0125] In some embodiments, the second ratio is subject to the following limitations:

[0126]

[0127] or

[0128]

[0129] The `min` function selects the smaller of the provided alternatives. Both alternatives have been found to provide good quality ratio values.

[0130] Note that in the example above, because there are several microphone pairs, the modified signal must be calculated separately for each pair; that is, when considering microphone pairs 1 and 3 or microphone pairs 1 and 2, They are not the same signal.

[0131] The first direction estimate 314, the first ratio estimate 316, the second direction estimate 324, and the second ratio estimate 326 are passed to a multiplexer (mux) 309, which is configured to generate a data stream 104 by combining these estimates with the streaming audio signal 308.

[0132] about Figure 4 This shows a summary Figure 3 The flowchart shows an example operation of the spatial analyzer.

[0133] Obtain the microphone audio signal, such as Figure 4 As shown in step 401.

[0134] Then, a streaming audio signal is generated from the microphone audio signal, such as... Figure 4 As shown in step 402.

[0135] In addition, time-frequency domain transformation can be performed on the microphone audio signal, such as... Figure 4 As shown in step 403.

[0136] Then, the first direction parameter estimate and the first ratio parameter estimate can be determined, such as Figure 4 As shown in step 405.

[0137] Then, the time-frequency domain microphone audio signal can be modified (to remove the first sound source component), such as... Figure 4 As shown in step 407.

[0138] Then, the modified time-frequency domain microphone audio signal is analyzed to determine the second direction parameter estimate and the second ratio parameter estimate, such as... Figure 4 As shown in step 409.

[0139] Then, the first direction parameter estimate, the first ratio parameter estimate, the second direction parameter estimate, the second ratio parameter estimate, and the streaming audio signal are multiplexed to generate a data stream (which can be a MASA format data stream), such as... Figure 4 As shown in step 411.

[0140] The following example describes a spatial filtering method and apparatus in which several gain parameters are determined or calculated and set to adjust the filtering process. These gains can be categorized as band-wise gain, history-based (time) gain, and frame-based smoothing gain.

[0141] In the following examples, direct vs. ambient (DA) ratio estimates are provided for the two estimated directions of origin (DOA) of each subband. This essentially indicates how much of the corresponding direction estimate is considered the "direct" signal portion and how much is considered the "ambient" signal portion. In these examples, the term "direct" refers to the signal arriving directly from the sound source, while "ambient" refers to echoes and background noise present in the environment. The direct and ambient components of the signal for each subband b can have a range of [0, 1] and are defined as follows:

[0142] ,

[0143] .

[0144] In some embodiments, the method begins, after obtaining the direction and extent of the spatially filtered region (which may also be defined as the focal sector or zoom sector), by checking whether these sub-bands are located within the sector of interest for either of the two direction estimates, none of them, or both. In the following example, the spatial filtering is a positive notch filter, where the audio signal within the sector of interest is amplified relative to the audio signal outside the sector of interest. However, in some embodiments, the spatial filtering is a negative notch filter, where the audio signal within the sector of interest is reduced relative to the audio signal outside the sector of interest. It will be understood that the difference between the two will be whether the sector gain is greater than the gain outside the sector (which produces a positive spatial notch filter) or whether the sector gain is less than the gain outside the sector (which produces a negative spatial notch filter).

[0145] about Figure 5 The diagram shows simplified illustrations of these three main scenarios.

[0146] In this example, the sound is amplified within the sector and attenuated outside the sector; however, this processing is also significantly affected by the DA ratio of the direction estimation.

[0147] For example, the DA ratio estimate can be considered as a weight used for actual direction estimation. The numbers in the table below are merely examples to illustrate the basic principle of their effect on the derived filter gain G(b). The first two rows show the case where either of the two sound sources is estimated to be similar to the ambient sound, meaning that its direction estimate should not be used for filtering in this way.

[0148]

[0149] Therefore, a low DA ratio value may indicate that the corresponding direction estimate may not be caused by a real sound source, because in some cases, no direct sound source is active during capture, or there is only one sound source. In some embodiments, the sector edges may also have regions where the applied sub-band gain is linearly smoothed to avoid abrupt gain changes at the sector edges.

[0150] Therefore, as Figure 5 As shown, there exists a first scenario 501 in which both sound sources are within the sector, which will result in the filter gains g1(b) and g2(b) corresponding to each direction estimate being greater than 1, and therefore, the spatial gain G(b) will produce a value greater than 1.

[0151] The second scenario 503 is shown, in which a sound source is in the sector, and the filter gain corresponding to one direction estimate (first g1(b)) is greater than 1, and the other (second g2(b)) is less than 1. Therefore, the spatial gain G(b) will produce a value close to 1.

[0152] A third scenario 505 is also shown, in which both sound sources are outside the sector, which will produce filter gains g1(b), g2(b) corresponding to each direction estimate that are less than 1. Therefore, the spatial gain G(b) will produce a value less than 1.

[0153] In some embodiments, the energy of subband b of the input signal spectrum X(b) can be estimated as follows before any energy adjustment:

[0154] ,

[0155] ,

[0156] Here, IIRFactor < 1.0 defines how much of the energy from the previous time frame is included to smooth the energy levels between time frames. Before the first frame, the energy at each subband b can be initialized as bandEne(b) = 0.

[0157] In some embodiments, the band gain is derived for each sub-band b based on the direction estimates d1 and d2. The direction estimates can be located within the focal sector, outside the focal sector, or in a region near the sector edge (the so-called edge zone). The direct energy component used for the first direction estimate d1 for sub-band b can be modified as follows:

[0158]

[0159] Here, inGain and outGain are adjustable and / or user-defined parameters used to control the focusing effect intensity of sound sources within and outside the focus sector, and

[0160] ,

[0161] Where angleDiff1 is the angular difference between the observed first direction estimate d1 and the edge of the sector, and edgeWidth is the width of the edge zone, for example, 20 degrees. Furthermore, in some embodiments, the environmental signal portion used for the first direction estimate of sub-band b can be modified as follows:

[0162] ,

[0163] Next, calculate the total energy adjustment for subband b:

[0164] .

[0165] The target energy for frequency band b after energy adjustment (which is initialized to 0 before the first frame) can be defined as:

[0166] ,

[0167] ,

[0168] Subsequently, the actual band gain value for sub-band b corresponding to the first direction estimate d1 is calculated as follows:

[0169]

[0170] To account for the second-direction estimate d2, the g2(b) gain value is calculated similarly to the g1(b) value. Then, these gains are multiplied together to obtain the total bandwidth gain:

[0171] .

[0172] Furthermore, in some embodiments, for both directional estimates d1 and d2, the time-filtered gain for each sub-band is calculated to smooth the filter gain over time. This prevents unnatural, abrupt pumps and notches in the total filter gain. In many cases, the estimated source DA ratio values ​​can vary across sub-bands, which is why averaging the DA ratio over the entire filter frequency range provides a good estimate of how similar the acoustic environment is to the surrounding environment at the current time frame f. For the first directional estimate, the average ratio is calculated at each frame as follows:

[0173] ,

[0174] Among them, b low It is the lowest frequency subband to be filtered, b high The highest frequency subband to be filtered. Furthermore, the average past ratio is tracked over a preferred number of previous frames (i.e., the history length, which can be a user-defined and / or adjustable parameter). Then, the calculated average ratio is further averaged over the history segments to obtain the time ratio mean:

[0175] ,

[0176] Here, `frames` represents the number of frames in the historical segment, for example, 60. For the second-direction estimate `d2`, the temporal ratio mean is further scaled as follows:

[0177] ,

[0178] It is better suited for weighted filtering purposes than the original DA-ratio scale. For each subband b and the two direction estimates d1 and d2, a Boolean flag (which indicates whether the direction estimate of the subband at the current frame f is within the focal sector) is also used to track the amount of past direction estimates within the focal sector.

[0179] .

[0180] Once the historical fragment is filled with such flags, the number of "true" flags at each subband b for d1 is N1. T (b) was used to obtain temporary scaling variables.

[0181] ,

[0182] Here, tempGain is an adjustable and / or user-defined parameter, with typical values ​​of [1.0, …, 6.0]. It can be seen that the scaling variable decreases as the "true" flag decreases, and vice versa. Finally, the time gain for d1 is calculated as…

[0183] ,

[0184] Here, bias is a constant between 0 and 1, used to control how much weight is given to the DA ratio value when deriving the time gain. Typically, the value can be set, for example, ~ 0.4 – 0.6.

[0185] The number of directional estimates within the sector at each sub-band b in the past is N1 T (b) can also be used to provide a so-called decay state for later use, as follows:

[0186] .

[0187] The time gain for direction estimation d2 is calculated similarly to the time gain used for d1. And the actual time filtering gain is obtained through multiplication.

[0188] .

[0189] In some embodiments, the direction estimation across all subbands within a single time frame can vary significantly depending on the number and type of sound sources present in the acoustic environment. Therefore, additional frame smoothing gain is needed to smooth the spectrum in order to prevent abrupt pumps and dips in the spectral envelope at each frame. First, the sum of the ratios of d1 and d2 can be calculated as:

[0190] ,

[0191] Next, the ratio of the intra-sector estimate Nin to the intra-frame estimate N in all directions is used to calculate the smoothing factor:

[0192] ,

[0193] Then it is applied to frame gain calculation.

[0194] ,

[0195] Here, smoothGain is an adjustable gain parameter with typical values ​​of [1.0, …, 2.0]. Higher values ​​provide more effective filtering performance, but they may cause unwanted gain level pumping, especially when there is large background noise in the capture.

[0196] The previously derived attenuation states were used to calculate the actual filter smoothing gain for each subband:

[0197] ,

[0198] in, This is an adjustable attenuation gain. Similarly, the smoothing gain for d2 is calculated, and the total smoothing gain is obtained through multiplication:

[0199] .

[0200] Once all the different gain types—band gain, time gain, and frame gain—have been calculated, the actual output filter gain can be determined or calculated for each subband b.

[0201] ,

[0202] Furthermore, the output is compressed and limited based on the available clearance in the subsequent processing chain.

[0203] exist Figure 6 Examples illustrating the advantages of the embodiments described herein are shown. Specifically, Figure 6 The output signal levels, in dB, are shown for a known spatial filter 601 using only a single direction estimate per subband and a spatial filter method 603 according to some embodiments. In this example, the audio focus direction is directly set to the front of the device, and the signal includes the speaker initially speaking in front of the device, then moving to the back of the device in the middle of the signal, and finally returning to the front of the device. Furthermore, music is played from a speaker located on the left side of the capture device. It can be seen that, compared to known methods, the embodiment amplifies the speech from the front by an average of approximately 2-3 dB.

[0204] Furthermore, compared to known spatial filtering methods, the embodiment also attenuates speech from behind the device by more than 2-3 dB, meaning the embodiment increases the overall focus effect gain by an average of 4-6 dB. This is a clearly audible and noticeable difference, improving the perceived audio zoom experience in most cases. As long as the direction estimates d1 and d2 can be estimated from the capture, the spatial filter can always improve its performance compared to estimating only d1.

[0205] about Figure 7This provides a summary of the operation of the embodiments described herein.

[0206] The first operation is to calculate or determine the direction estimates d1 and d2 used for subband b, such as Figure 7 Step 701 is shown.

[0207] Then, a first check can be performed to determine whether d1 is within the sector, such as... Figure 7 Step 703 is shown.

[0208] If d1 is within the sector, further checks can be performed to determine if d2 is within the sector, such as... Figure 7 Step 705 is shown.

[0209] If d1 and d2 are both within the sector, then the subband b is amplified based on the DA ratio of d1 and d2, as estimated by their correlation. Figure 7 Step 707 is shown.

[0210] If d1 is not within the sector, further checks can be performed to determine if d2 is within the sector, such as... Figure 7 Step 709 is shown.

[0211] If d1 is within the sector but d2 is not, or d1 is not within the sector but d2 is, then subband b can be amplified based on the estimated DA ratio within the sector, and subband b can be attenuated based on the estimated DA ratio outside the sector, such as... Figure 7 Step 711 is shown.

[0212] If d1 and d2 are both outside the sector, then the DA ratio of d1 and d2 is estimated to attenuate the subband b, such as... Figure 7 Step 713 is shown.

[0213] about Figure 8 The diagram shows a flowchart illustrating the generation gain according to some embodiments.

[0214] Therefore, in some embodiments, the band gain g(b) is calculated for both directions, i.e. and ,like Figure 8 Step 801 is shown.

[0215] Then, in some embodiments, the band gains are multiplied to generate a combined band gain. ,like Figure 8 Step 803 is shown.

[0216] Then, for each sub-band and direction, a time gain is generated. , ,like Figure 8 Step 805 is shown.

[0217] Then, the time gains can be multiplied to generate a combined time gain. ,like Figure 8 Step 807 is shown.

[0218] Then, the frame smoothing gain can be determined for each sub-band and orientation. , ,like Figure 8 Step 809 is shown.

[0219] Then, the frame smoothing gains can be multiplied to generate a combined frame smoothing gain. ,like Figure 8 Step 811 is shown.

[0220] Then, for subband b, the total filter gain for the subband can be generated by multiplying the combined frame smoothing gain, the combined time gain, and the combined band gain. ,like Figure 8 Step 813 is shown.

[0221] about Figure 9 This shows, as Figure 1 The example spatial synthesizer 105 is shown.

[0222] In some embodiments, the spatial synthesizer 105 includes a demultiplexer 1201. In some embodiments, the demultiplexer 1201 receives a data stream 104 and splits the data stream into a streaming audio signal 1208 and spatial parameter estimates, such as a first direction 1214 estimate, a first ratio 1216 estimate, a second direction 1224 estimate, and a second ratio 1226 estimate.

[0223] These are then passed to the space processor / synthesizer 1203.

[0224] Spatial synthesizer 105 includes spatial processor / synthesizer 1203 and is configured to receive these estimated and streamed audio signals and render the output audio signal. Spatial processing / synthesis can be any suitable two-way based synthesis, such as that described in EP3791605.

[0225] Figure 10 and Figure 11 An end-to-end implementation of the embodiment is shown. About Figure 10 The image shows a capture device 1101 and a playback device 1111, which communicate via a transmission / storage channel 1105.

[0226] The capture device 1101 is configured as described above and is configured to transmit filtered audio 1109. Additionally, filter orientation / range information 1107 can be received from the playback device 1111.

[0227] about Figure 11 The image shows a capture device 1101 configured to transmit unfiltered audio 1119 received by a playback device 1111. The playback device includes a spatial filter 1103 configured to apply spatial filtering as discussed in the embodiments described herein.

[0228] about Figure 12 This illustrates an example electronic device that can be used as a computer, encoder processor, decoder processor, or any functional block described herein. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1600 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc.

[0229] In some embodiments, device 1600 includes at least one processor or central processing unit 1607. Processor 1607 may be configured to execute various program codes, such as the methods described herein.

[0230] In some embodiments, device 1600 includes memory 1611. In some embodiments, at least one processor 1607 is coupled to memory 1611. Memory 1611 can be any suitable storage device. In some embodiments, memory 1611 includes a program code portion for storing program code that can be implemented on processor 1607. Furthermore, in some embodiments, memory 1611 may further include a storage data portion for storing data, such as data that has been processed or will be processed according to the embodiments described herein. The implemented program code stored in the program code portion and the data stored in the storage data portion can be retrieved by processor 1607 via memory-processor coupling when needed.

[0231] In some embodiments, device 1600 includes a user interface 1605. In some embodiments, user interface 1605 may be coupled to processor 1607. In some embodiments, processor 1607 may control the operation of user interface 1605 and receive input from user interface 1605. In some embodiments, user interface 1605 enables a user to input commands to device 1600, for example, via a keypad. In some embodiments, user interface 1605 enables a user to obtain information from device 1600. For example, user interface 1605 may include a display configured to show information from device 1600 to a user. In some embodiments, user interface 1605 may include a touchscreen or touch interface capable of allowing information to be input to device 1600 and further displaying the information to a user of device 1600.

[0232] In some embodiments, device 1600 includes an input / output port 1609. In some embodiments, input / output port 1609 includes a transceiver. In this embodiment, the transceiver may be coupled to processor 1607 and configured to communicate with other devices or electronic devices, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver device may be configured to communicate with other electronic devices or devices via wired or wired coupling.

[0233] The transceiver can communicate with other devices using any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol (such as IEEE 802.X), a suitable short-range radio frequency communication protocol (such as Bluetooth), or an Infrared Data Communication Path (IRDA).

[0234] The transceiver input / output port 1609 can be configured to send / receive audio signals, bit streams, and in some embodiments, the operations and methods described above are performed by using a processor 1607 that executes appropriate code.

[0235] Generally, various embodiments of the present invention can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while others may be implemented in firmware or software executable by a controller, microprocessor, or other computing device, although the invention is not limited thereto. Although various aspects of the invention may be illustrated and described as block diagrams, flowcharts, or represented using certain other graphical representations, it is well understood that such blocks, apparatuses, systems, techniques, or methods described herein may be implemented (by way of non-limiting example) in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0236] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device (such as in a processor entity), or by hardware, or by a combination of software and hardware. Further, in this respect, it should be noted that any block in the logical flow of the figures may represent a program step, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on a physical medium such as a memory chip, or a memory block implemented within a processor, magnetic media, or optical media.

[0237] The memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic storage devices and systems, optical storage devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and can include one or more of the following as non-limiting examples: general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits, and processors based on multi-core processor architectures.

[0238] Embodiments of the present invention can be practiced in a variety of elements, such as integrated circuit modules. The design of integrated circuits is largely a highly automated process. Sophisticated and powerful software tools can be used to transform logic-level designs into semiconductor circuit designs ready to be etched and formed on semiconductor substrates.

[0239] Programs (such as those provided by Synopsys in Mountain View, California, and Cadence Design in San Jose, California) use well-established design rules and pre-stored libraries of design modules to automatically route conductors and position components on semiconductor chips. Once the semiconductor circuit design is complete, the final design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be sent to a semiconductor manufacturing facility or “fab” for fabrication.

[0240] The foregoing description has provided a complete and informative description of exemplary embodiments of the invention by way of exemplary and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the art when read in conjunction with the accompanying drawings and appended claims, given the foregoing description. Nevertheless, all such modifications and similar alterations to the teachings of the invention will still fall within the scope of the invention as defined by the appended claims.

Claims

1. A method for spatial audio filtering, comprising: Obtain two or more audio signals; For one or more frequency bands of the two or more audio signals, obtain the first sound source direction parameter and the first sound source energy parameter; For one or more frequency bands of the two or more audio signals, obtain the second sound source direction parameter and the second sound source energy parameter; Obtain the region for spatial filtering; Based on the relationship between the region and the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter, a gain / attenuation parameter is obtained to adjust the spatial filter; as well as The adjusted spatial filter is applied to the two or more audio signals.

2. The method according to claim 1, wherein, Based on the region, obtaining the gain / attenuation parameters includes: Based on the first sound source direction parameter within or outside the region, a first frequency band gain / attenuation value is generated; Based on the second sound source direction parameter within or outside the region, a second frequency band gain / attenuation value is generated; and The first band gain / attenuation value and the second band gain / attenuation value are combined to generate a combined band gain / attenuation value.

3. The method according to claim 2, wherein, Obtaining the region that defines the direction and / or range for the spatial filter includes at least one of the following: The direction and range, together with the in-band gain / attenuation factor in the region based on the first sound source direction parameter and the out-of-band gain / attenuation factor outside the region based on the second sound source direction parameter; as well as Define the further extent of the edge zone region, along with the edge zone band gain / attenuation factor.

4. The method according to claim 1, wherein, Applying an adjusted spatial filter to the two or more audio signals includes: A first time gain / attenuation value is generated based on the time average of the average frequency band value of the first sound source energy parameter within the defined time period and the number of times the first sound source direction parameter is within or outside the region. A second time gain / attenuation value is generated based on the time average of the average frequency band value of the second sound source energy parameter within the defined time period and the frequency of the second sound source direction parameter within or outside the region; and A combined time gain / attenuation value is generated based on the combination of the first time gain / attenuation value and the second time gain / attenuation value.

5. The method according to claim 1, wherein, Applying an adjusted spatial filter to the two or more audio signals includes: A combined frame average value is generated based on a combination of the first sound source energy parameter and the second sound source energy parameter (frame averaged). Based on the combined frame average value and the number of times the first and second sound source direction parameters are within the filter region during the frame period, a frame smoothing gain / attenuation is generated.

6. The method according to claim 1, wherein, Obtaining the gain / attenuation parameters includes generating the gain / attenuation for one or more frequency bands based on a combination of frame smoothing gain / attenuation, combined time gain / attenuation values, and combined frequency band gain / attenuation values.

7. The method according to claim 1, wherein, The applied adjusted spatial filter includes: Process the two or more audio signals to provide one or more modified audio signals based on the two or more audio signals; and Based on the modified audio signal, the second sound source direction parameter and the second sound source energy parameter are determined in one or more frequency bands of the two or more audio signals.

8. The method according to claim 7, further comprising: Outputting one or more modified audio signals based on the two or more audio signals includes: generating two or more modified audio signals by modifying the two or more audio signals based on the projection of a first sound source defined by the first sound source direction parameter.

9. The method according to claim 8, further comprising: Outputting one or more modified audio signals based on the two or more audio signals, further comprising: determining at least a second sound source direction parameter in one or more frequency bands of the two or more audio signals, at least in part based on the one or more modified audio signals, including: determining at least the second sound source direction parameter in one or more frequency bands of the two or more audio signals by processing the modified two or more audio signals.

10. The method according to claim 1, wherein, Obtaining the region that defines the direction and / or range for the spatial filter includes: obtaining the region based on user input.

11. An apparatus for spatial audio filtering, comprising: At least one processor; as well as At least one memory including computer program code; The at least one memory and the computer program code are configured, together with the at least one processor, to make the device at least: Obtain two or more audio signals; For one or more frequency bands of the two or more audio signals, obtain the first sound source direction parameter and the first sound source energy parameter; For one or more frequency bands of the two or more audio signals, obtain the second sound source direction parameter and the second sound source energy parameter; Obtain the region for spatial filtering; Based on the relationship between the region and the first sound source direction parameter, the first sound source energy parameter, the second sound source direction parameter, and the second sound source energy parameter, a gain / attenuation parameter is obtained to adjust the spatial filter; as well as The adjusted spatial filter is applied to the two or more audio signals.

12. The apparatus according to claim 11, wherein, Based on the region, the gain / attenuation parameters are obtained so that the device: Based on the first sound source direction parameter within or outside the region, a first frequency band gain / attenuation value is generated; Based on the second sound source direction parameter within or outside the region, a second frequency band gain / attenuation value is generated; as well as The first band gain / attenuation value and the second band gain / attenuation value are combined to generate a combined band gain / attenuation value.

13. The apparatus according to claim 12, wherein, Obtaining the region that defines the direction and / or range for the spatial filter includes at least one of the following: The direction and range, together with the in-band gain / attenuation factor in the region based on the first sound source direction parameter and the out-of-band gain / attenuation factor outside the region based on the second sound source direction parameter; as well as Define the further extent of the edge zone region, along with the edge zone band gain / attenuation factor.

14. The apparatus according to claim 11, wherein, The device is made to apply an adjusted spatial filter to the two or more audio signals: A first time gain / attenuation value is generated based on the time average of the average frequency band value of the first sound source energy parameter within the defined time period and the number of times the first sound source direction parameter is within or outside the region. A second time gain / attenuation value is generated based on the time average of the average frequency band value of the second sound source energy parameter within the defined time period and the number of times the second sound source direction parameter is within or outside the region. as well as A combined time gain / attenuation value is generated based on the combination of the first time gain / attenuation value and the second time gain / attenuation value.

15. The apparatus according to claim 11, wherein, The device is made to apply an adjusted spatial filter to the two or more audio signals: A combined frame average value is generated based on a combination of the first sound source energy parameter and the second sound source energy parameter (frame averaged). Based on the combined frame average value and the number of times the first and second sound source direction parameters are within the filter region during the frame period, a frame smoothing gain / attenuation is generated.

16. The apparatus according to claim 11, wherein, Obtaining the gain / attenuation parameters enables the device to generate the gain / attenuation for one or more frequency bands based on a combination of frame smoothing gain / attenuation, combined time gain / attenuation values, and combined frequency band gain / attenuation values.

17. The apparatus according to claim 11, wherein, The device is made possible by applying an adjusted spatial filter: Process the two or more audio signals to provide one or more modified audio signals based on the two or more audio signals; as well as Based on the modified audio signal, the second sound source direction parameter and the second sound source energy parameter are determined in one or more frequency bands of the two or more audio signals.

18. The apparatus of claim 17, caused to perform at least one of the following: Based on processing the two or more audio signals using the projection of the first sound source defined by the first sound source direction parameter, two or more modified audio signals are generated; and Output one or more modified audio signals based on the two or more audio signals.

19. The apparatus according to claim 18, wherein, The device outputs the one or more modified audio signals to: determine at least a second sound source direction parameter in the one or more frequency bands of the two or more audio signals, at least in part, based on the one or more modified audio signals; and the device determines at least a second sound source direction parameter in the one or more frequency bands of the two or more audio signals by processing the two or more modified audio signals.

20. The apparatus according to claim 11, wherein, Obtaining the region that defines the direction and / or range for the spatial filter enables the device to: obtain the region based on user input.

Citation Information

Patent Citations

  • An apparatus, method and computer program for audio signal processing

    EP3791605A1