Audio rendering of spatial audio
By controlling regularization in spatial audio rendering based on input characteristics like bitrate and codec format, the method addresses quality issues in existing technologies, ensuring improved audio quality across different input conditions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2026-03-17
AI Technical Summary
Existing spatial audio rendering technologies face challenges in achieving optimal audio quality due to fixed regularization values that do not account for varying input signal characteristics such as bitrate and codec format, leading to issues like noise amplification and excessive decorrelation, particularly at different bitrates and signal-to-noise ratios.
A method and apparatus for controlling regularization in spatial audio rendering based on input characteristics, such as bitrate and codec format, to minimize artifacts and improve perceived audio quality by adjusting rendering gain.
The proposed solution enhances audio quality by reducing reverberation and noise amplification, resulting in more engaging and artifact-free spatial audio reproduction across varying input conditions.
Smart Images

Figure 2026509128000001_ABST
Abstract
Description
[Technical Field]
[0001] This application relates to an apparatus and method for audio rendering of spatial audio and for the application of regularization in rendering, but is not limited to configuring regularization of a mixing solution for rendering. [Background technology]
[0002] There are many ways to capture spatial audio. One option is to capture spatial audio using a microphone array as part of a mobile device, for example. The microphone signal can be used to perform a spatial analysis of the sound scene and determine spatial metadata in the frequency band. Furthermore, the microphone signal can be used to determine the transport audio signal. The spatial metadata and the transport audio signal can be combined to form a spatial audio stream.
[0003] Metadata-assisted spatial audio (MASA) is an example of a spatial audio stream. It is one of the input formats that future Immersive Speech and Audio Services (IVAS) codecs will likely support. It uses an audio signal along with corresponding spatial metadata (e.g., direction and direct-to-overall energy ratio in frequency bands) as well as descriptive metadata (e.g., additional information about the original captured audio signal and (transport) audio signal). A MASA stream can be obtained, for example, by capturing spatial audio using a mobile device's microphone, with the set of spatial metadata inferred based on the microphone signal. MASA streams can also be obtained from other sources, such as specific spatial audio microphones (e.g., ambisonics), studio mixes (e.g., 5.1 mixes), or other content, through appropriate format conversion. It is also possible to use the MASA tool within the codec for encoding multi-channel channel signals by converting a multi-channel signal to a MASA stream and then encoding that stream. [Overview of the project]
[0004] According to a first aspect, a method for generating an audio signal is provided, the method comprising: acquiring an input audio signal including at least two audio channels; acquiring at least one input characteristic parameter associated with the input audio signal; determining at least one control parameter at least based on the at least one input characteristic parameter; determining processing parameters at least based on at least two audio channels, wherein the determination of the processing parameters is controlled at least in part based on at least one control parameter; and generating an audio signal at least based on the at least two audio channels and the processing parameters.
[0005] The input audio signal may also be a spatial audio signal, which further includes at least one spatial parameter associated with at least two audio channels.
[0006] Determining processing parameters based on at least two audio channels may further include determining processing parameters based on at least one spatial parameter.
[0007] At least one input characteristic parameter may include at least one of the following: a bitrate associated with at least one input audio signal, a codec format that indicates the origin of at least one input audio signal, and the configuration of at least one input audio signal.
[0008] The configuration of at least one input audio signal may include at least one of the following: a source format that indicates the original format in which the input audio signal was created or the input format, a description of the transport channels, the number of channels, the channel distance, and the channel angle.
[0009] A codec format that indicates the origin of at least one input audio signal may include at least one of the following: a metadata-assisted spatial audio stream origin codec format value that indicates at least one input audio signal represents a metadata-assisted spatial audio signal; a multichannel audio stream origin codec format value that indicates at least one input audio signal represents a multichannel audio signal; an audio object codec format value that indicates the input audio signal represents an audio object; and an ambisonic format value that indicates the input audio signal represents an ambisonic audio signal.
[0010] At least one spatial parameter is information that describes the organization of sound in a space with respect to at least one input audio signal, including a direction parameter configured to indicate where the sound is coming from, and a ratio parameter configured to indicate a portion of the sound coming from that direction, and information that describes characteristics of an original multi-channel or multi-object sound scene, including at least one of channel level, object level, inter-channel correlation, inter-object correlation, and object direction, and may include at least one of the at least one processing coefficient related to obtaining a spatial audio format signal based at least on the at least one input audio signal.
[0011] A codec format indicating the origin of at least one input audio signal may include an indication that the origin is unknown or undefined.
[0012] Determining at least one control parameter based at least on at least one input characteristic parameter may include determining a first set of control parameter values based on a first input characteristic parameter among the at least one input characteristic parameter, and selecting one control parameter value from the first set of control parameter values based on a second input characteristic parameter among the at least one input characteristic parameter.
[0013] The first input characteristic parameter among the at least one input characteristic parameter may be a codec format, and the second input characteristic parameter among the at least one input characteristic parameter may be a bit rate.
[0014] Determining processing parameters based at least on at least two audio channels, wherein the determination of the processing parameters is controlled at least partially based on at least one control parameter, may include generating a processing matrix based at least on at least two audio channels, wherein the generation of the processing matrix may be controlled at least partially based on at least one control parameter.
[0015] Generating a processing matrix that is controlled at least partially based on a control parameter may include regularizing the generation of the processing matrix based at least on the control parameter.
[0016] Generating an audio signal based at least on at least two audio channels and processing parameters may include processing at least two audio channels using the regularized processing matrix to generate the audio signal.
[0017] Generating a processing matrix that is controlled at least partially based on at least one control parameter may include generating components of a diagonal matrix based at least on at least two control parameter values.
[0018] Obtaining at least one input characteristic parameter associated with an input audio signal may include inferring at least one input characteristic parameter from the input audio signal.
[0019] Obtaining at least one input characteristic parameter associated with an input audio signal may include receiving at least one input characteristic parameter as part of or as configuration information of the input audio signal.
[0020] Receiving at least one input characteristic parameter as part of the input audio signal may further include receiving an encoded input characteristic parameter as part of the input audio signal and decoding the encoded input characteristic parameter to obtain the input characteristic parameter.
[0021] Generating an audio signal based on at least two audio channels and processing parameters may include generating at least one of a binaural audio signal and a multi-channel audio signal.
[0022] According to a second embodiment, an apparatus for generating an audio signal is provided, the apparatus comprising means configured to acquire an input audio signal including at least two audio channels; acquire at least one input characteristic parameter associated with the input audio signal; determine at least one control parameter based at least on the at least one input characteristic parameter; determine processing parameters based at least on at least two audio channels, wherein the determination of the processing parameters is controlled at least in part on at least one control parameter; and generate an audio signal based at least on at least two audio channels and processing parameters.
[0023] The input audio signal may also be a spatial audio signal, which further includes at least one spatial parameter associated with at least two audio channels.
[0024] A means configured to determine processing parameters based on at least two audio channels may be further configured to determine processing parameters based on at least one spatial parameter.
[0025] At least one input characteristic parameter may include at least one of the following: a bitrate associated with at least one input audio signal, a codec format that indicates the origin of at least one input audio signal, and the configuration of at least one input audio signal.
[0026] The configuration of at least one input audio signal may include at least one of the following: a source format that indicates the original format in which the input audio signal was created or the input format, a description of the transport channels, the number of channels, the channel distance, and the channel angle.
[0027] A codec format that indicates the origin of at least one input audio signal may include at least one of the following: a metadata-assisted spatial audio stream origin codec format value that indicates at least one input audio signal represents a metadata-assisted spatial audio signal; a multichannel audio stream origin codec format value that indicates at least one input audio signal represents a multichannel audio signal; an audio object codec format value that indicates the input audio signal represents an audio object; and an ambisonic format value that indicates the input audio signal represents an ambisonic audio signal.
[0028] At least one spatial parameter may include information describing the spatial composition of sound with respect to at least one input audio signal, including a direction parameter configured to indicate where the sound is coming from and a ratio parameter configured to indicate the portion of the sound coming from that direction; information describing the characteristics of the original multi-channel or multi-object sound scene, including at least one of channel level, object level, inter-channel correlation, inter-object correlation, and object direction; and at least one of processing coefficients related to obtaining a spatial audio format signal based at least on the at least one input audio signal.
[0029] A codec format that indicates the origin of at least one input audio signal may include an indication that the origin is unknown or undefined.
[0030] Means configured to determine at least one control parameter based at least one input characteristic parameter may be further configured to determine a first set of control parameter values based on a first input characteristic parameter among the at least one input characteristic parameter, and to select one control parameter value from the first set of control parameter values based on a second input characteristic parameter among the at least one input characteristic parameter.
[0031] The first of at least one input characteristic parameter may be the codec format, and the second of at least one input characteristic parameter may be the bitrate.
[0032] Means configured to determine processing parameters based at least on at least two audio channels, wherein the determination of the processing parameters is controlled at least in part on at least one control parameter, the means may be further configured to generate a processing matrix based at least on at least two audio channels, wherein the generation of the processing matrix may be controlled at least in part on at least one control parameter.
[0033] Means configured to generate a processing matrix controlled at least partially on control parameters may be further configured to regularize the generation of the processing matrix at least on control parameters.
[0034] Means configured to generate an audio signal based on at least two audio channels and processing parameters may be further configured to process at least two audio channels using a regularized processing matrix to generate an audio signal.
[0035] Means configured to generate a processing matrix controlled at least partially on at least one control parameter may be further configured to generate components of a diagonal matrix based at least on at least two control parameter values.
[0036] Means configured to obtain at least one input characteristic parameter associated with an input audio signal may be further configured to deduce at least one input characteristic parameter from the input audio signal.
[0037] Means configured to acquire at least one input characteristic parameter associated with an input audio signal may be further configured to receive at least one input characteristic parameter as part of the input audio signal or as configuration information.
[0038] Means configured to receive at least one input characteristic parameter as part of an input audio signal may be further configured to receive an encoded input characteristic parameter as part of an input audio signal and to decode the encoded input characteristic parameter to obtain an input characteristic parameter.
[0039] Means configured to generate an audio signal based on at least two audio channels and processing parameters may be further configured to generate at least one of a binaural audio signal and a multi-channel audio signal.
[0040] According to a third aspect, an apparatus for generating an audio signal is provided, the apparatus comprising at least one processor and at least one memory for storing instructions that, when executed by the at least one processor, cause the system to determine an input audio signal including at least two audio channels, to determine an input characteristic parameter associated with the input audio signal, to determine at least one control parameter based at least on the at least one input characteristic parameter, and to generate an audio signal based at least on the at least two audio channels, wherein the determination of the processing parameter is controlled at least in part on the at least one control parameter.
[0041] The input audio signal may also be a spatial audio signal, which further includes at least one spatial parameter associated with at least two audio channels.
[0042] A device that can be configured to determine processing parameters based on at least two audio channels may be further configured to determine processing parameters based on at least one spatial parameter.
[0043] At least one input characteristic parameter may include at least one of the following: a bitrate associated with at least one input audio signal, a codec format that indicates the origin of at least one input audio signal, and the configuration of at least one input audio signal.
[0044] The configuration of at least one input audio signal may include at least one of the following: a source format that indicates the original format in which the input audio signal was created or the input format, a description of the transport channels, the number of channels, the channel distance, and the channel angle.
[0045] A codec format that indicates the origin of at least one input audio signal may include at least one of the following: a metadata-assisted spatial audio stream origin codec format value that indicates at least one input audio signal represents a metadata-assisted spatial audio signal; a multichannel audio stream origin codec format value that indicates at least one input audio signal represents a multichannel audio signal; an audio object codec format value that indicates the input audio signal represents an audio object; and an ambisonic format value that indicates the input audio signal represents an ambisonic audio signal.
[0046] At least one spatial parameter may include information describing the spatial composition of sound with respect to at least one input audio signal, including a direction parameter configured to indicate where the sound is coming from and a ratio parameter configured to indicate the portion of the sound coming from that direction; information describing the characteristics of the original multi-channel or multi-object sound scene, including at least one of channel level, object level, inter-channel correlation, inter-object correlation, and object direction; and at least one of processing coefficients related to obtaining a spatial audio format signal based at least on the at least one input audio signal.
[0047] A codec format that indicates the origin of at least one input audio signal may include an indication that the origin is unknown or undefined.
[0048] An apparatus that can be configured to determine at least one control parameter based on at least one input characteristic parameter may further be configured to determine a first set of control parameter values based on a first input characteristic parameter among the at least one input characteristic parameter, and to select one control parameter value from the first set of control parameter values based on a second input characteristic parameter among the at least one input characteristic parameter.
[0049] The first of at least one input characteristic parameter may be the codec format, and the second of at least one input characteristic parameter may be the bitrate.
[0050] An apparatus that can be configured to determine processing parameters based on at least two audio channels, wherein the determination of the processing parameters is controlled at least in part on at least one control parameter, the apparatus may be configured to further generate a processing matrix based on at least two audio channels, wherein the generation of the processing matrix may be controlled at least in part on at least one control parameter.
[0051] An apparatus capable of generating a processing matrix controlled at least partially on control parameters may further be capable of regularizing the generation of the processing matrix at least on control parameters.
[0052] A device that can be configured to generate an audio signal based on at least two audio channels and processing parameters may be configured to further generate an audio signal by processing at least two audio channels using a regularized processing matrix.
[0053] An apparatus that can be configured to generate a processing matrix controlled at least partially on at least one control parameter may further be configured to generate components of a diagonal matrix based at least on at least two control parameter values.
[0054] A device that can be configured to obtain at least one input characteristic parameter associated with an input audio signal may be configured to further deduce at least one input characteristic parameter from the input audio signal.
[0055] A device that can be configured to acquire at least one input characteristic parameter associated with an input audio signal may further be configured to receive the at least one input characteristic parameter as part of the input audio signal or as configuration information.
[0056] A device that can be configured to receive at least one input characteristic parameter as part of an input audio signal may further be configured to receive an encoded input characteristic parameter as part of an input audio signal and to decode the encoded input characteristic parameter to obtain the input characteristic parameter.
[0057] A device that can be configured to generate an audio signal based on at least two audio channels and processing parameters may be configured to further generate at least one of a binaural audio signal and a multi-channel audio signal.
[0058] According to a fourth aspect, an apparatus for generating an audio signal is provided, the apparatus comprising means for acquiring an input audio signal including at least two audio channels; means for acquiring at least one input characteristic parameter associated with the input audio signal; means for determining at least one control parameter at least based on the at least one input characteristic parameter; means for determining processing parameters at least based on at least two audio channels, wherein the determination of the processing parameters is controlled at least in part based on at least one control parameter; and generating an audio signal at least based on the at least two audio channels and the processing parameters.
[0059] According to a fifth aspect, an apparatus for generating an audio signal is provided, comprising: an acquisition circuit configured to acquire an input audio signal including at least two audio channels; an acquisition circuit configured to acquire at least one input characteristic parameter associated with the input audio signal; a determination circuit configured to determine at least one control parameter at least based on the at least one input characteristic parameter; a determination circuit configured to determine processing parameters at least based on at least two audio channels, wherein the determination of the processing parameters is controlled at least partially based on at least one control parameter; and generating an audio signal at least based on at least two audio channels and processing parameters.
[0060] According to the sixth aspect, a computer program (or a computer-readable medium including program instructions) is provided that includes instructions for causing an audio signal generating device to perform at least the following: acquiring an input audio signal including at least two audio channels; acquiring at least one input characteristic parameter associated with the input audio signal; determining at least one control parameter based at least on the at least one input characteristic parameter; determining processing parameters based at least on at least two audio channels, wherein the determination of the processing parameters is controlled at least in part on at least one control parameter; and generating an audio signal based at least on at least two audio channels and the processing parameter.
[0061] According to the seventh aspect, a non-temporary computer-readable medium is provided which includes program instructions for causing an audio signal generating apparatus to perform at least the following: acquiring an input audio signal including at least two audio channels; acquiring at least one input characteristic parameter associated with the input audio signal; determining at least one control parameter based at least on the at least one input characteristic parameter; and determining processing parameters based at least on at least two audio channels, wherein the determination of the processing parameters is controlled at least in part on at least one control parameter; and generating an audio signal based at least on at least two audio channels and the processing parameters.
[0062] According to the eighth aspect, a computer-readable medium is provided which includes program instructions for causing an apparatus for generating an output audio signal to perform at least the following: acquiring an input audio signal including at least two audio channels; acquiring at least one input characteristic parameter associated with the input audio signal; determining at least one control parameter based at least on the at least one input characteristic parameter; determining processing parameters based at least on at least two audio channels, wherein the determination of the processing parameters is controlled at least in part on at least one control parameter; and generating an audio signal based at least on at least two audio channels and the processing parameters.
[0063] A device equipped with means for performing actions in the manner described above.
[0064] A device configured to perform actions in the manner described above.
[0065] A computer program that contains program instructions to cause a computer to perform the methods described above.
[0066] Computer program products stored on a medium may be used to cause the device to perform actions as described herein.
[0067] The electronic device may include any apparatus as described herein.
[0068] The chipset may include devices such as those described herein.
[0069] The embodiments of this application aim to address problems related to cutting-edge technology.
[0070] For a better understanding of this application, references to the attached drawings will be made here as an example. [Brief explanation of the drawing]
[0071] [Figure 1] This diagram schematically illustrates an exemplary system for capturing or otherwise acquiring spatial audio signals in the form of transport audio signals and spatial metadata. [Figure 2] This diagram schematically illustrates an exemplary system for capturing or otherwise acquiring spatial audio signals in the form of transport audio signals and spatial metadata. [Figure 3] This diagram schematically illustrates an exemplary system for capturing or otherwise acquiring spatial audio signals in the form of transport audio signals and spatial metadata. [Figure 4] This figure schematically illustrates an exemplary system for encoding and reproducing spatial audio signals in the form of transport audio signals and spatial metadata, suitable for implementing several embodiments. [Figure 5]This figure schematically illustrates an exemplary system of multiple operating modes based on encoding a spatial audio signal in the form of a transport audio signal and spatial metadata, and reproducing the spatial audio signal, suitable for implementing several embodiments. [Figure 6] This figure schematically illustrates an exemplary regeneration device suitable for carrying out several embodiments. [Figure 7] This figure schematically illustrates an exemplary decoder, such as the one shown in Figure 5, according to several embodiments. [Figure 8] Figure 7 is a flowchart illustrating the operation of an exemplary decoder device according to several embodiments. [Figure 9] This figure schematically illustrates an exemplary spatial synthesizer, as shown in Figure 5, according to several embodiments. [Figure 10] Figure 9 is a flowchart illustrating the operation of an exemplary spatial synthesizer in several embodiments. [Figure 11] Figure 9 shows a flowchart illustrating the operation of an exemplary regularization coefficient determinationer in several embodiments. [Figure 12] This figure shows an exemplary processing output. [Modes for carrying out the invention]
[0072] The following further details appropriate devices and possible mechanisms for rendering a suitable output audio signal from a parametric spatial audio stream (or signal) obtained from captured or otherwise acquired audio signals.
[0073] As discussed above, Metadata-Assisted Spatial Audio (MASA) is an example of a suitable parametric spatial audio format and representation for use as input to IVAS.
[0074] This can be thought of as an audio representation consisting of "N channels + spatial metadata." This is a scene-based audio format particularly well-suited for spatial audio capture on practical devices such as smartphones. The idea is to describe the sound scene in terms of the direction of sound changes over time and frequency, and, for example, the energy ratio. Sound energy that is not defined (described) by direction is described as diffusion (coming from all directions).
[0075] As discussed above, spatial metadata associated with an audio signal may include multiple parameters for each time-frequency tile (multiple directions, and associated with each direction (or direction value), such as direct-to-whole ratio, diffuse coherence, distance, etc.). Spatial metadata may also include other parameters, or may be associated with other parameters that are considered non-directional (such as surround coherence, diffuse-to-whole energy ratio, residual-to-whole energy ratio, etc.), but when combined with directional parameters, they can be used to define the characteristics of the audio scene. For example, a reasonable design choice that can produce a good output is one in which the spatial metadata includes one or more directions for each time-frequency element (also known as a time-frequency portion or time-frequency tile). Furthermore, associated with each direction, the metadata may include other parameters such as direct-to-whole ratio, diffuse coherence, distance value, etc.
[0076] As described above, the parametric spatial metadata representation can use multiple simultaneous spatial directions. In MASA, the proposed maximum number of simultaneous directions is two. For each simultaneous direction, there may be associated parameters such as the direction index, direct-to-whole ratio, diffusion coherence, and distance. In some embodiments, other parameters such as the diffusion-to-whole energy ratio, ambient coherence, and residual-to-whole energy ratio are defined.
[0077] Parametric spatial metadata values are available for each time-frequency tile (the MASA format defines each frame as having 24 frequency bands and 4 time subframes). The frame size of IVAS is 20 milliseconds. Furthermore, here, MASA supports one or two directions per time-frequency tile.
[0078] Exemplary metadata parameters may include the following: Format descriptor. This parameter defines the MASA format for IVAS. This may be stored as a 64-bit value, or as eight 8-bit ASCII characters (01001001, 01010110, 01000001, 01010011, 01001101, 01000001, 01010011, 01000001), and the value is stored as eight consecutive 8-bit unsigned integers. Channel audio format. This parameter defines the following combined fields, stored in 2 bytes, and the value is stored as a single 16-bit unsigned integer. The number of directions. This parameter defines the number of directions described by the spatial metadata. Each direction may be associated with a set of direction-dependent spatial metadata, as will be described later. The number of directions parameter may be one of a range of values. Here, the possible values of the defined range are 1 or 2, but in some embodiments, two or more directions may be used. The number of channels. This parameter defines the number of transport channels used by the format. This channel number parameter may be one of a range of values. Here, the possible values in the defined range are 1 or 2, but in some embodiments, two or more channels may be used. Source format. This parameter describes the original format from which the MASA output was created or the input format (for example, the input format is 5.1 multichannel).
[0079] In some situations, there may be additional descriptive or parameter fields based on the values of the "Number of Channels" and "Source Format" parameter fields. In some embodiments, zero-padding may be applied when not all of the allocated bits are used.
[0080] Examples of spatial metadata parameters in the MASA format that depend on the number of directions may be as follows: Directional index. This parameter defines the direction of sound arrival at a time-frequency parameter interval. Typically, this direction of arrival is a spherical representation with an accuracy of approximately 1 degree. Direct versus total energy ratio. This parameter defines the energy ratio of the directional index (in other words, it defines the energy ratio associated with direction relative to the time-frequency subframe). Diffuse coherence. This parameter defines the diffusion of energy with respect to the directional index (in other words, it defines a measure of the "width" of direction with respect to the time-frequency subframe). Transport definition. This describes the configuration of two transport channels. Channel angle. This describes the symmetrical angular position relative to the transport signal having a directional pattern. Channel distance. This describes the distance between two transport channels. Channel layout. When the source format is multi-channel, this describes the channel layout of the multi-channel source format.
[0081] Examples of spatial metadata parameters in the MASA format that are independent of the number of directions may be as follows: Diffuse versus total energy ratio. This parameter defines the energy ratio of omnidirectional sound across the surrounding direction. Surround coherence. This parameter defines the coherence of omnidirectional sound in all directions. Residual vs. Total Energy Ratio. This parameter defines the energy ratio of residual sound energy (such as microphone noise) to satisfy the requirement that the sum of the energy ratios is 1.
[0082] Furthermore, the frequency bandwidth of exemplary spatial metadata may be as follows: [Table 1]
[0083] The MASA stream can be rendered to various outputs, such as multi-channel loudspeaker signals (e.g., 5.1) or binaural signals.
[0084] One example rendering method is described in Vilkamo, J., Backstrom, T., & Kuntz, A. (2013). Optimized covariance domain framework for time-frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411. The rendering method is based on multi-channel mixing. The method processes a given audio signal within a frequency band so that a desired covariance matrix is obtained for the output signal within the frequency band. The covariance matrix includes the channel energy of all channels and the inter-channel relationships between all channel pairs, i.e., cross-correlation and inter-channel phase difference. These features are known to convey the perception-related spatial features of multi-channel sound in various playback situations, such as headphones, surround loudspeakers, ambisonics, and binaural sound compared to crosstalk-canceling stereo.
[0085] As discussed above, the rendering method in Vilkamo, J., Backstrom, T., & Kuntz, A. (2013). Optimized covariance domain framework for time-frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411 attempts to process the input audio signal (e.g., a decoded transport audio signal) so that the generated output audio signal (e.g., a binaural audio signal) has a desired covariance matrix. This target covariance matrix is determined based on the energy of the input audio signal and the spatial metadata received by the renderer.
[0086] In practice, this rendering method first attempts to achieve the target covariance matrix by mixing the input audio signals. However, regularization is applied during processing to prevent the input audio signals from being amplified with very large gain values. For example, if the input audio signals are highly correlated (in other words, very similar signals), but the target covariance matrix is set to have low correlation, this will lead to a processing gain that will over-amplify the non-coherent signal portions. If regularization is not applied, this will lead to a significant amplification of noise and other unwanted sounds in the input audio signals, resulting in degraded sound quality.
[0087] The process of acquiring the mixing gain is regularized, so the target covariance matrix is not necessarily achieved by mixing alone. When regularization is applied, some signal portions are not amplified as much as necessary to reach the target, and therefore some of the expected or desired signal energy is lost. As a result, these sounds are substantially attenuated during regeneration, and the desired non-coherence between output channels may not be achieved. A proposed solution to achieve the desired non-coherence is to decorrelate the input signals to obtain non-coherent signals, and then process these signals with the mixing gain to obtain the covariance matrix of the lost signal portions. These means that the target covariance matrix is obtained for the output signals even when processing is limited by regularization.
[0088] Both methods, mixing and decorrelation, have their advantages and disadvantages. Decorrelation processing has the drawback of affecting sound quality, especially when there is a significant amount of decorrelated signal in the output. When sound is decorrelated, its phase spectrum is usually modified, at least to some extent, in order to preserve the timbre. The perceived quality of certain sounds, such as speech and applause, is known to deteriorate with this processing.
[0089] On the other hand, mixing with too much gain can lead to problems such as excessive amplification of small signal components, which in some situations can be almost entirely noise.
[0090] The amount of regularization can be used to control how much uncorrelated energy is mixed into the output. When only mild regularization is applied (i.e., a large maximum allowable mixing gain), the output contains little uncorrelated signal and largely encompasses the mixture of the input signals. In contrast, when a significant amount of regularization is applied (i.e., a small maximum allowable mixing gain), the output contains more uncorrelated signal because a large mixing gain is prevented.
[0091] Therefore, the amount of regularization is a compromise between avoiding the mixing problem and avoiding the decorrelation problem. The optimal amount of regularization is situation-dependent, but a few examples are discussed below.
[0092] The IVAS codec operates across a wide range of bitrates (13.2kbps to 512kbps). At the lowest bitrates, in particular, the transport audio signal received by the renderer contains significant coding artifacts (musical and other noise and distortion). Therefore, to avoid amplifying these coding artifacts, a significant amount of regularization is required, which will make the artifacts even more prominent and significantly degrade the perceived sound quality. On the other hand, having a large amount of regularization (more decorrelation) would not be an optimal choice at higher bitrates. The transport audio signal contains significantly fewer of these coding artifacts, and therefore this leads to more decorrelation than necessary, resulting in suboptimal sound quality.
[0093] Furthermore, the IVAS codec supports multiple input formats. The MASA format often originates from mobile devices that typically have inexpensive microphones, and therefore usually contains perceptible microphone noise in the MASA input signal. Consequently, a significant amount of regularization is required in the renderer to prevent amplification of microphone noise. On the other hand, IVAS also supports multi-channel input (such as 5.1), which is typically recorded and produced in a studio using low-noise professional microphones. Therefore, a fixed setting for regularization is not optimal in this respect either, as different formats will require different values to achieve a good experience within a certain range.
[0094] Known rendering methods may have fixed regularization values that are set for situations with high signal-to-noise (SNR) characteristics. Fixed regularization values based on high SNR conditions may not produce acceptable results for low SNR conditions. For example, spatial audio codecs operating at low bitrates, such as 32kbps or less, may not benefit from regularization values set for high-bitrate, low-noise input signals.
[0095] As a result, a fixed regularization value produces optimal rendering only for certain types of input signals and / or for a certain bitrate at which the transported audio signal is encoded. For other types of input signals and / or bitrates, it can result in either excessive regularization (leading to excessive decorrelation) or insufficient regularization (leading to noise amplification), both of which lead to a degradation of perceived audio quality.
[0096] Excessive decorrelation can produce sounds that are resonant and non-engaging, while noise amplification can produce regenerated sounds that are perceived as artifacts.
[0097] The embodiments and concepts discussed herein are one of rendering spatial audio from a parametric spatial audio stream (audio signals and associated spatial metadata), or more generally, from at least one audio signal containing two or more audio channels. In these embodiments, a rendering method (and therefore a renderer device) is proposed that enables the acquisition of improved audio quality by controlling regularization in the determination of rendering gain based on input characteristics (or input parameters) associated with at least one audio signal (such as bitrate and / or codec format). The controlled regularization aims to improve perceived audio quality by minimizing artifacts resulting from uncorrelatedness (increased reverberation and loss of appeal) and excessive amplification of noise (perception of loud music and other noises), and therefore mitigating the perception of these artifacts.
[0098] In some embodiments, this can be achieved by acquiring a (parametric spatial) audio stream, acquiring input characteristics, determining a regularization value based on the input characteristics, determining a regularized rendering gain based on the regularization value, the audio signal, and associated spatial metadata, and rendering a spatial audio signal (such as a binaural audio signal) from the audio signal using the regularized rendering gain.
[0099] In some embodiments, input characteristics can mean various things depending on the different embodiment. For example, input characteristics could be the bitrate used to code the parametric spatial audio stream (if encoded / decoded before rendering), if it is, for example, a MASA stream (i.e., typically from a mobile device), or if it is created from a multi-channel signal (e.g., 5.1), the codec format describing the origin of the stream, or the composition of the audio signal received by the renderer (e.g., if it is from a cardioid or omnidirectional microphone).
[0100] As discussed above, the embodiments discussed in further detail below present a system that utilizes a spatial sound renderer to control the regularization of the rendering gain based on at least one input characteristic. In some embodiments, the renderer is implemented within the decoder, and two exemplary input characteristics are possible. The input characteristics discussed herein are the codec format and the bitrate. In some other embodiments, the renderer may be implemented outside the decoder, and other input characteristics may be possible.
[0101] First, exemplary “codec format” input characteristics are described in more detail, explaining how signals of different codec formats can be acquired. Next, “bitrate” input characteristics are described in more detail. After this, the implementation of several embodiments within the decoder is described in more detail.
[0102] The input characteristics of a codec format can indicate the type of parametric spatial audio signal being transmitted. As mentioned earlier, a parametric spatial audio signal is a signal that includes one or more audio signals and associated spatial metadata. The spatial metadata contains information that indicates how the sound is spatially composed, and it is usually provided at a significantly lower sampling rate than the audio signal.
[0103] Examples of spatial metadata in different "codec formats" include the following: Information describing the composition of sound in space. For example, in the frequency band, the directional parameter may indicate where the sound is coming from, and the ratio parameter may indicate the portion of the sound that came from that direction. Information describing the characteristics of the original multi-channel or multi-object sound scene. For example, the levels of channels or objects, the correlations between channels or objects, or the orientation of objects. A processing coefficient related to obtaining a certain spatial audio format signal (such as an ambisonic audio signal) based on a transmitted audio signal.
[0104] The example shown above can be extended by using other types of spatial metadata. The input characteristics of the codec format can therefore indicate the type of spatial metadata being transmitted.
[0105] In some embodiments, the input characteristics of the codec format may also indicate (or alternatively indicate) the type or origin of the transport audio signal.
[0106] For example, the input characteristics of the codec format can, in some embodiments, indicate that the transport audio signals are a downmix of multi-channel signals, or that they were captured by a microphone array, or that they are a decomposition of ambisonic audio signals. In some embodiments, the codec input characteristics can have defined values that can indicate that the origin of the transport audio signals is unknown or undefined.
[0107] In some embodiments, two different codec formats may have the same type of spatial metadata or the same type of transport audio signal.
[0108] Any of the above audio formats may be rendered to a number of output spatial audio formats, such as multi-channel loudspeaker output (e.g., 5.1), binaural output with or without head tracking, or crosstalk-canceling stereo. The following example illustrates the synthesis of binaural audio signals. The example described herein can be easily extended and applied to the regeneration of other output types.
[0109] While the following embodiments describe the application of regularization in a renderer where the input is parametric spatial audio, it will be understood that some embodiments may be applicable to any suitable renderer and rendering operation in which regularization is applied to the input audio to be rendered and the regularization value is controlled based on input parameters such as bitrate and codec format.
[0110] Figure 1 shows an example of a device configured to determine a parametric spatial audio signal from a microphone array signal. This device could be implemented, for example, in a mobile phone equipped with a microphone array (integrated within the mobile phone).
[0111] In such embodiments, the microphone array signal 100 can be transferred to a microphone array front end 101. The microphone array front end may be implemented using, for example, a method presented in U.S. Patent No. 1,0873814, and is configured to output a transport audio signal 102 and spatial metadata 104. The transport audio signal 102 and spatial metadata 104 may be in the form of, for example, a MASA stream.
[0112] In some embodiments, the microphone array front end 101 is configured to measure the inter-microphone correlation at different delays between microphones within a frequency band, find the delay that maximizes the correlation, and determine the direction of the incoming sound based on that delay. In some embodiments, the microphone array front end can also determine a direct-to-overall energy ratio parameter within a frequency band based on the correlation value.
[0113] In some embodiments, the microphone array front end can also provide a transport audio signal 102. The process for determining the transport audio signal 102 depends on the type or format of the microphone array signal that has been input. For example, if the microphone array signal 100 is from a mobile device, the microphone array front end 101 may be configured to select a microphone signal from the left side of the device as the left transport signal and another microphone signal from the right side of the device as the right transport signal. In some embodiments, the microphone array front end 101 can also be subjected to any appropriate preprocessing steps such as equalization, microphone noise suppression, wind noise suppression, automatic gain control, beamforming and other spatial filtering, ambient noise suppression, and limiters.
[0114] With respect to Figure 2, further exemplary apparatus configured to determine a parametric spatial audio signal from a microphone array signal is shown. In some embodiments, the microphone array providing the microphone array signal may be a dedicated microphone array, for example, a microphone array providing a primary ambisonic (FOA) signal at its output. Thus, in some embodiments, the apparatus comprises an ambisonic decomposer 201 configured to receive the ambisonic signal 200 and generate a transport audio signal 102 and spatial metadata 104. In some embodiments, the ambisonic decomposer 201 is configured to determine the spatial metadata. The spatial metadata may be determined using a method similar to directional audio coding (DirAC), as described, for example, in Pulkki, V. (2007). Spatial sound reproduction with directional audio coding. Journal of the Audio Engineering Society, 55(6), 503-516, and the transport audio signal 104 may be a left- and right-directed cardioid pattern generated from the ambisonic (FOA) signal 200. Therefore, it should be noted that spatial metadata and transport audio signals can be of similar types, but their origins can be different.
[0115] In some embodiments, the ambisonic decomposer 201 may be configured to receive an ambisonic signal 200, for example, FOA or higher-order ambisonics (HOA), and determine spatial metadata 104 so that the ambisonic signal can be reconstructed in the decoder based on the transport audio signal. For example, the ambisonic signal may be transmitted as one or more channel transport audio signals 102, where the first channel is an omnidirectional W component and the remaining channels are residual signals. The ambisonic decomposer 201 determines these signals by determining prediction coefficients that can predict the remaining channels from the W component, and the residual signals become prediction errors. For those channels in which the residuals are not transmitted, parameters may be estimated that can process the omnidirectional component by decorrelation with its residual signal. All of these prediction and processing parameters may form spatial metadata 104.
[0116] With respect to Figure 3, further exemplary apparatus configured to determine a parametric spatial audio signal from a microphone array signal is shown. In some embodiments, the downmixer and metadata determiner 301 is configured to receive a channel-based audio signal 300 and generate a transport audio signal 102 and spatial metadata 104. The exemplary channel-based audio signal 300 is a surround 5.1 sound or an audio object sound. The downmixer and metadata determiner 301 is configured to generate spatial metadata 104 by converting the audio object signal and / or audio channels to FOA format and then determining the spatial metadata using DirAC or similar means. The downmixer and metadata determiner 301 is configured to generate the transport audio signal 104 using amplitude panning such that channels / objects (and corresponding cone of confusion) greater than ±30 degrees are fully panned to the left and right channels, and channels / objects between ±30 degrees are panned to both channels depending on their direction.
[0117] The codec format input parameter can therefore indicate various aspects previously described. Depending on the use case, it can indicate entirely different signal formats (e.g., definition of stereo transport versus master residual signal), and / or different characteristics of the same signal format (e.g., stereo transport signal from a downmix versus stereo transport signal from a microphone), as well as / or it can indicate metadata formats. For example, the codec format input parameter could be an index in a predefined list of options that define both the type of transport audio signal and spatial metadata being used.
[0118] The bitrate input parameter can be configured to indicate the number of bits used to transmit information within a given time frame. In network transmissions, where audio codecs such as IVAS are also used, the usual description is to measure the bits used per second (abbreviated as bps). As mentioned earlier, IVAS is assumed to operate at bitrates of 13.2kbps to 512kbps. The bitrate may also be constant or vary over time. Thus, each transmission frame may have an equal or variable number of bits. In IVAS, a transmission frame is expected to be 20 milliseconds, and the bitrate is assumed to be constant under steady-state conditions. Thus, the above bitrates would correspond to 264 bits / frame and 10240 bits / frame. It should be noted that this is typically the total bitrate, which is then further divided for signaling, audio channel coding, and metadata coding, if present. This further division may also be constant or variable.
[0119] In the following description, bitrate indicates the overall quality of the transmitted audio. Generally, with lossy audio codecs such as IVAS, increasing the bitrate improves the overall quality of the transmitted audio. Depending on the codec format, this improvement in quality may stem from an increase in bits used for audio channel coding, which directly reduces the presence of coding artifacts, or from an increase in bits in spatial metadata coding, which improves the accuracy of the reconstructed spatial scene. Thus, based on the bitrate, an indication of the expected overall quality is provided.
[0120] Figures 4 and 5 show an overview of an exemplary apparatus illustrating the signal flow following the generation of the transport audio signal 102 and spatial metadata 104.
[0121] Figure 4 shows an exemplary device comprising an encoder 401 configured to receive a transport audio signal 102 and spatial metadata 104, and encode them to form a bitstream 402. Furthermore, the device comprises a decoder 402 configured to receive the bitstream 402 and output a spatial audio output 404. In other words, the system can be thought to operate in a first mode in which processing software outside the encoder 401 determines and provides the transport audio signal 102 and spatial metadata 104. For example, this software could be microphone array front-end software optimized for processing the microphone signals of a particular device.
[0122] Figure 5 shows a further exemplary apparatus comprising an encoder 401 configured to receive a transport audio signal 102 and spatial metadata 104 and encode them to form a bitstream 402. Furthermore, the apparatus comprises a decoder 402 configured to receive the bitstream 402 and output a spatial audio output 404. Furthermore, it is shown that there is an encoder preprocessor 501 prior to the encoder 401. The encoder preprocessor 501 is configured to receive an audio signal 500 and generate the transport audio signal 102 and spatial metadata 104. In other words, the system can be thought of to operate in a second mode in which the encoding system 511 that performs encoding also performs encoder preprocessing for obtaining the transport audio signal 102 and spatial metadata 104. For example, an IVAS encoder (equipped with both the encoder preprocessor 501 and the encoder 401) may be configured to accept an 5.1 or ambisonics input and generate the transport audio signal 102 and spatial metadata 104 to be encoded thereafter.
[0123] The encoder 401 thus forms a bitstream 402, which can be stored or transmitted. The bitstream contains a transport audio signal 102 and spatial metadata 104 in an encoded form. The audio signal may be encoded using, for example, an IVAS core codec, EVS, or AAC encoder (or any other suitable encoder), and the metadata may be encoded using, for example, the methods presented in U.S. Patent Application Publication No. 20210295855, U.S. Patent Application Publication No. 20220343928, U.S. Patent Application Publication No. 20220036906, EP4091166 (and / or any other suitable method). In some embodiments, the encoder 401 can further multiplex the encoded audio and encoded spatial metadata to form the bitstream 402.
[0124] Furthermore, the encoder 401 is configured to write input parameters, such as the codec format input parameters being used, to or include them within the bitstream 402. For example, this can be directly signaled using signaling bits that define the codec format input parameters. For instance, the codec format may be signaled based on two bits, such as "00" for multi-channel format, "01" for MASA format, and "10" for ambisonic format (these are merely illustrative values). In addition, the encoder 401 may also write other input parameters, such as the bitrate input parameter, to the bitstream 402. The bitrate input parameter may be constant or may change over time.
[0125] In some embodiments, the bitrate input parameter may be inferred within the decoder from the size of the received bitstream frame.
[0126] In some embodiments, input parameter information, such as codec format input parameters and bitrate input parameter information, is provided through a different communication channel other than the bitstream 402 that transmits the transport audio signal and spatial metadata.
[0127] The bitstream 402 is transferred to the decoder 403, which may be, for example, an IVAS decoder (or any other suitable decoder). The decoder 403 decodes the audio signal and metadata and renders a spatial audio output 404, which may be, for example, a binaural audio signal. The decoder 403 may be a different device from or the same device as the encoder 401.
[0128] With respect to Figure 6, exemplary (decoder) devices for carrying out several embodiments are shown. In the example shown in Figure 6, a mobile phone 601 is shown coupled to headphones 619 worn by the user of the mobile phone 601 via a wired or wireless connector 613. The wired or wireless connector 613 may allow an audio signal 615 to be passed to the headphones 619 and a monaural audio signal (or multiple audio signals) 617 to be passed from the headphones 619 (which in some embodiments may further include head orientation / position metadata from the headphones).
[0129] In the following, the exemplary device or apparatus is a mobile phone as shown in Figure 6. However, the exemplary device or apparatus may also be any other suitable device, such as a tablet, laptop, computer, or any video conferencing device. The apparatus or apparatus may also be the headphones themselves, and therefore the operation of the exemplary mobile phone 601 may be performed by the headphones.
[0130] In this example, the mobile phone 601 includes a processor 603. The processor 603 may be configured to execute various program codes, such as those described herein. The processor 603 is configured to communicate with headphones 619 using a wired or wireless headphone connector 615. In some embodiments, the wired or wireless headphone connector 615 is a Bluetooth 5.3 or Bluetooth LE audio connector. The connector 615 provides the processor 603 with a (2-channel) audio signal that is regenerated for the user using headphones 619.
[0131] The headphones 619 may be over-ear headphones as shown in Figure 6, or any other suitable type of headphones such as in-ear or bone conduction headphones, or any other type of headphones. In some embodiments, the headphones 619 include a head orientation sensor that provides head orientation information to the processor 603 via a connector 613. In some embodiments, the head orientation sensor is separate from the headphones 619, and the data is provided separately to the processor 603. In further embodiments, the head orientation is tracked by a camera in device 601 and other means such as machine learning-based face orientation analysis.
[0132] In some embodiments, the processor 603 is coupled to a memory 605 having program code 607 that provides processing instructions according to the following embodiments. The program code 607 has instructions for processing transport audio signals and spatial metadata received by the transceiver 611 or retrieved from storage 609 into a rendering format suitable for effective output to headphones.
[0133] The transceiver 611 can communicate with further devices by any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable radio access architecture based on Long Term Evolution Advanced (LTE Advanced, LTE-A) or NR (which may be referred to as New Radio, or 5G), Universal Mobile Communications System (UMTS) radio access network (UTRAN or E-UTRAN), Long Term Evolution (LTE, same as E-UTRA), 2G network (legacy network technology), Wireless Local Area Network (WLAN or Wi-Fi), Worldwide Interoperability for Microwave Access (WiMAX), Bluetooth®, Personal Communication Services (PCS), ZigBee® Wideband Code Division Multiple Access (WCDMA), systems using ultra-wideband (UWB) technology, sensor networks, Mobile Ad Hoc Networks (MANET), Cellular Internet of Things (IoT) RAN and Internet Protocol Multimedia Subsystem (IMS), any other suitable options, and / or any combination thereof.
[0134] In some embodiments, the apparatus in Figure 6 may also be configured to further perform the functions of the microphone array front end 101, ambisonic decomposer 201, downmixer and metadata determiner 301, encoder 401, and encoder preprocessor 501, for example, when the use case is bidirectional spatial audio communication with a remote device.
[0135] Regarding Figure 7, a schematic diagram of an exemplary decoder 403, as shown in Figure 5, is provided.
[0136] The input bitstream 402 is transferred to the demultiplexer and decoder 701. The demultiplexer and decoder 701 is configured to multiplex, decouple, and decode the transport audio signal 706 from the bitstream 402, and to multiplex, decouple, and decode the spatial metadata 708 from the bitstream 402. The decoding corresponds to the encoding applied by the encoder 401. It should be noted that the transport audio signal 706 and spatial metadata 708 are, since they have been encoded and decoded, not necessarily identical to those previously presented. Nevertheless, for the sake of brevity, they are referred to using the same terminology (though with different reference numbers).
[0137] Furthermore, the demultiplexer and decoder 701 are configured to determine the parameters of the bitrate 702 and codec format 704 based on the information in the bitstream 402. In one example, these may be indicated by dedicated bits in the bitstream 402. In other examples, one or both may be detected from the bitstream, or they may be communicated through other channels, for example, via RTP or session control.
[0138] The spatial metadata 708, transport audio signal 706, codec format 704, and bitrate 702 are then provided to the spatial synthesizer 703.
[0139] The spatial synthesizer 703 is configured to synthesize the spatial audio output 404 in a desired format based on spatial metadata 708, transport audio signal 706, codec format 704, and bitrate 702. In some examples, only one of the parameters of codec format 704 and bitrate 702 is used, while in other examples, both are used. This information is used to balance the use of decorrelation and signal mixing in the execution of spatial synthesis. The output provided by the spatial synthesizer 703 may be, for example, a binaural audio signal.
[0140] With respect to Figure 8, the exemplary flowchart illustrates the operation of the exemplary decoder shown in Figure 7, according to several embodiments.
[0141] Therefore, the first operation could include obtaining a bitstream (encoded spatial audio), as indicated by 801.
[0142] Subsequently, as shown in 803, the (encoded spatial audio) bitstream is multiplexed, decoupled, and decoded to generate the transport audio signal, spatial metadata, and input parameters such as the codec format and bitrate.
[0143] Subsequently, as shown by 805, a spatial audio signal is synthesized from the transport audio signal based on spatial metadata and input parameters such as codec format and bitrate. In other words, the transport audio signal is spatially processed to generate a spatial audio signal based on the codec format, bitrate, and spatial metadata.
[0144] Subsequently, as indicated by 807, a spatial audio signal is output (for example, a binaural audio signal is output to headphones).
[0145] With respect to Figure 9, the spatial combiner 703 of Figure 7 is shown in more detail. In some embodiments, the spatial combiner 703 is configured to receive a transport audio signal 706, spatial metadata 708, and input parameters such as the bitrate 702 and codec format 704.
[0146] In some embodiments, the spatial combiner 703 includes a forward-filter bank 901. As shown in Figure 9, the transport audio signal 706 is supplied to the forward-filter bank 901, which converts the signal to a time-frequency representation. Any filter bank suitable for audio processing may be used, such as a complex-modulated quadrature mirror filter (QMF) bank, a low-latency version thereof, or a short-time Fourier transform (STFT). Similarly, the forward-filter bank 901 can be implemented by any suitable time-frequency converter.
[0147] In the examples described herein, the filter bank is configured to have 60 frequency bins and sufficient stop-band attenuation to prevent significant aliasing when the frequency bin signals are processed. In this configuration, all frequency bins can be processed independently of each other, except that some frequency bins share the same spatial metadata. For example, the spatial metadata may include spatial parameters in a limited number of frequency bands, e.g., 5, 12, or 24 bands, each of which corresponds to a set of one or more frequency bins provided by the forward filter bank 901. While this example shows a specific number of bands, any suitable number of bands is possible; for example, the number of frequency bands could be 5, 8, 12, 18, or 24.
[0148] The output of the forward filter bank 901 is a time-frequency transport signal 906, which is supplied to the decorrelator and mixer 907, the processing matrix decisioner 905, and the input and target covariance matrix decisioner 909. The time-frequency transport signal 906 S(b,t,i) is,
[0149]
number
[0150] In some embodiments, the spatial synthesizer 703 comprises an input and target covariance matrix determinator 909. The input and target covariance matrix determinator 909 is configured to receive spatial metadata 708 and a time-frequency transport signal 906 and to determine a covariance matrix 906. The covariance matrix 906 includes an input covariance matrix representing the time-frequency transport signal 906 and a target covariance matrix representing the desired time-frequency spatial audio signal (to be rendered). The input covariance matrix may be determined or measured from the time-frequency transport signal 906, indicated as a column vector x(b,t), where the rows indicate the transport signal channels. In some embodiments,
[0151]
number
[0152] The target covariance matrix can be determined based on spatial metadata and the total signal energy. The total signal energy E0(b,n) is C xIt is obtained as the average of the diagonal values of (b,n). Then, in some embodiments, the spatial metadata is the direction parameter azimuth angle θ(k,n) and elevation angle
[0153]
number
[0154]
number
number
[0155]
number
[0156] Input covariance matrix C x (b,n) and target covariance matrix C y (b,n) is then output to the processing matrix decisioner 905 as the covariance matrix 906.
[0157] The above example only considers direction and ratio. The procedure for generating the target covariance matrix is described in more detail in WO2019086757A1, which also describes the use of spatial coherence parameters in addition to direction and ratio, and further covers other output types in addition to binaural output.
[0158] In some embodiments, the spatial synthesizer 703 comprises a regularization coefficient determiner 903. The regularization coefficient determiner 903 is configured to receive input parameters such as codec format 704 and bit rate 702 and determine a regularization coefficient R(n) 904. The regularization coefficient determiner 903 is configured to output the regularization coefficient R(n) 904 to a processing matrix determiner 905.
[0159] In some embodiments, the spatial synthesizer 703 comprises a processing matrix determiner 905. The processing matrix determiner 905 receives covariance matrices C x (b,n) and C y (b,n) 906 and the regularization coefficient R(n) 904 and is configured to determine processing matrices 908 M(b,n) and M r (b,n). The determination of such a processing matrix based on the covariance matrix is based on a method as shown in Vilkamo, J., Backstrom, T., & Kuntz, A. (2013). Optimized covariance domain framework for time-frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411. This method determines a mixing matrix for processing an input audio signal having a measured covariance matrix C y (b,n) such that the output audio signal (i.e., the processed input audio signal) achieves the target covariance matrix C x (b,n).
[0160] This method is used in a variety of situations, including the generation of binaural and surround loudspeaker signals. In formulating the processing matrix, the method further uses a prototype matrix, which is a matrix that informs the optimization procedure what kind of signal is generally meant for each output (with the constraint that the output must achieve the target covariance matrix). When the transport audio signal is two-channel (left and right) and the output is a binaural signal, the prototype matrix Q is, for example, simply:
[0161]
number
number
[0162] Embodiments of this specification are C x (b,n), C y Processing matrices M(b,n) and M based on (b,n) and Q(n) rWhen determining (b,n), regularization based on R(n) is performed. The derivation of the equation for determining the processing matrix is thoroughly described in Vilkamo, J., Backstrom, T., & Kuntz, A. (2013). Optimized covariance domain framework for time-frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411. Note that this implementation is illustrative, and some operations may be performed in other ways to achieve the same or similar results. In some embodiments, the illustrative implementation provided in the appendix of Vilkamo, J., Backstrom, T., & Kuntz, A. (2013). Optimized covariance domain framework for time-frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411 can form the basis for the implementation for determining the processing matrix, the difference being that the regularization coefficient R(n) is used instead of the fixed number 0.2 in program code line 21 (of the illustrative implementation).
[0163] The use of regularization in determining the mixture matrix is explained here for completeness. Note that the matrix notation below is not exactly the same as that in Vilkamo, J., Backstrom, T., & Kuntz, A. (2013). Optimized covariance domain framework for time-frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411. Also, frequency and bin indices have been omitted for brevity.
[0164] First, let's define input covariance matrix and target covariance matrix.
number
[0165] after that,
number
[0166]
number
[0167]
number
[0168] In some embodiments, regularization is performed on the decomposed region
number
number
[0169]
number
[0170] Subsequently, the regularized inverse matrix is the mixture matrix.
number
[0171] In this case, we need to determine whether regularization is effective (i.e., whether the lower bound regularization is effective for matrix S) x,sq (depending on whether it acted on the condition MC) x M H =C y This may no longer be maintained. Non-zero missing covariance matrix C r =C y -MC x M H It is possible, This is sometimes called the residual covariance matrix. Therefore, to process the uncorrelated version of the transport audio signal and achieve the characteristics of its missing parts, in other words, the covariance matrix C r Another processing matrix M used to achieve this r To generate this, the same method presented in previously cited references is used. This is C r Set C as the target covariance matrix. x Remove the off-diagonal elements (because the sounds are uncorrelated) and use that as the input covariance matrix, otherwise use M r This may be done by performing the same operations described to obtain the result. Since the input is non-coherent, the function of regularization values is not so important at this stage. Therefore, a fixed value of 0.2, for example, may be used.
[0172] The processing matrix decisioner 905 uses the regularization coefficient R(n) 904 to determine the processing matrix M(b,n) and M r (b,n)908 can then be configured to output 908.
[0173] In some embodiments, the spatial combiner 703 comprises a decorrelator and a mixer 907. The decorrelator and mixer 907 take a time-frequency transport signal x(b,t) 906 and processing matrices M(b,n) and M r (b,n)908 is configured to receive the decorrelator and mixer 907 process the time-frequency transport signal 906 using the decorrelator to obtain the decorrelated signal x DIt is first constructed to generate (b,t).
[0174] Subsequently, the decorrelator and mixer 907 are configured to apply the following mixing procedure to generate a time-frequency space audio signal y(b,t)910, which can then be output by the decorrelator and mixer 907. y(b,t)=M(b,n)x(b,t)+M r (b,n)x D (b,t)
[0175] In the above process, although not explicitly stated in the formula, the processing matrix may be linearly interpolated between frames n such that the matrix steps from M(b,n-1) to M(b,n) at each time index of the time-frequency signal. The interpolation rate may be adjusted depending on whether an onset was detected (fast interpolation) or not (normal interpolation).
[0176] In some embodiments, the spatial combiner 703 includes an inverse filter-bank 911. The inverse filter-bank 911 converts the time-frequency spatial audio signal 910 into a spatial audio output 404, which is the output of the system, by applying an inverse transform corresponding to that used by the forward filter-bank 901.
[0177] With respect to Figure 10, the exemplary flowchart illustrates the operation of the spatial synthesizer shown in Figure 9 in several embodiments.
[0178] Therefore, the initial operation may include obtaining the transport audio signal, spatial metadata, and input parameters such as the codec format and bitrate, as shown by 1001.
[0179] Subsequently, as shown in 1003, the transport audio signal is converted to a time-frequency signal to generate a time-frequency transport audio signal.
[0180] The time-frequency transport audio signal can then be used to generate the input covariance matrix. Once the input covariance matrix is determined, it can then be used to determine the total energy. Subsequently, the total energy and spatial metadata can be used to generate the target covariance matrix. The generation of the input and target covariance matrices is shown in 1007.
[0181] Furthermore, as shown in 1005, the regularization coefficient is determined based on input parameters such as the codec format and bitrate.
[0182] Once the covariance matrix and regularization coefficients are determined, the processing matrix is determined based on the input covariance matrix, the target covariance matrix, and the regularization coefficients, as shown in 1009.
[0183] Subsequently, the time-frequency transport signal is decorrelated and mixed based on a processing matrix to generate a time-frequency spatial audio signal, as shown by 1011.
[0184] Subsequently, the time-frequency spatial audio signal is inverse time-frequency transformed to generate a spatial audio signal (e.g., using an inverse filter bank), as shown in 1013.
[0185] The spatial audio signal can then be output as a spatial audio output, as shown by 1015.
[0186] Regarding Figure 11, the operation of an exemplary regularization coefficient determinationer 903, as shown in Figure 9, is illustrated. In this example, the regularization coefficient determinationer is presented in the context of the exemplary spatial audio codec system described above. However, a similar process may be applied in any context where a regularization coefficient or similar mixing or amplification limiting coefficient is used to control the automatic mixing of input channels to output channels, and where constant values do not produce optimal quality.
[0187] Therefore, the initial operation is to obtain the input parameters that affect the regularization coefficients. For example, this could be obtaining the codec format, as shown in 1101, which can describe the spatial audio format of the codec passed to the renderer. For example, this could be the MASA format, the premixed multichannel format, the ambisonics format, etc.
[0188] In some embodiments, this also includes obtaining a bitrate input parameter that describes the bitrate used to encode the bitstream, as shown by 1105.
[0189] After obtaining the parameters, the process of defining or determining the regularization coefficients begins.
[0190] In some embodiments, the regularization coefficient itself can be defined as a decimal value in the range [0,1], where a value of 1 most restricts the rendering gain and a value of 0 does not restrict the rendering gain at all.
[0191] In some embodiments, the codec format parameters are used to select an initial set of available values for the regularization coefficients used in subsequent operation steps, as shown in 1103.
[0192] For example, this can be done in the form of several tables, each corresponding to a specific codec format. For instance, a table for the MASA format might be [1.0, 0.8, 0.5, 0.2], and a similar table for the Ambisonics format might be [1.0, 0.7, 0.4, 0.2]. For premixed multichannel formats, a similar exemplary table might be [0.8, 0.5, 0.3, 0.2], because this content is typically professionally generated and audio compression is generally less prone to compression artifacts. It should be noted that the use of tables is just one example of an implementation option, and other equivalent or even better methods for implementing the selection can be created.
[0193] Furthermore, as shown in 1107, a set of values is taken, and an initial value for the regularization coefficient is selected based on the bitrate of the acquired codec and the selection from the initial set of values. For example, taking the table in the aforementioned MASA format, a specific bitrate is associated with each table value. In this particular example, there are four values, so the exemplary rule set is: Use the value 1.0 if the bitrate is ≤ 64kbps. If the bitrate is >64kbps and the bitrate is ≤96kbps, use the value 0.8. If the bitrate is >96kbps and ≤160kbps, use the value 0.5. If the bitrate is >160kbps, use the value 0.2. It is possible.
[0194] In some embodiments, an alternative way to carry out this decision would be to have a value defined within a set of values for each possible bitrate.
[0195] Any suitable method can be used to obtain the initial regularization coefficient for a given bitrate and codec format combination. A general inference with the exemplary values presented herein is that as the bitrate increases, the overall audio signal quality can also be expected to improve. This supports the use of smaller regularization coefficients (in other words, larger rendering gains) because artifacts become less likely to appear.
[0196] Subsequently, the regularization coefficient determination unit 903 outputs the regularization coefficient R(n)904 and passes the coefficient to the processing matrix determination unit 905, which applies the regularization coefficient R(n)904 as part of a mixed matrix solution.
[0197] It should be noted that all the exemplary values shown above are so-called adjustable values for the system. In practice, there are many interactions between different parts of the codec and renderer, and the values ultimately used are usually obtained using a set of objective measurements and subjective listening tests. While given exemplary values are appropriate, they may not produce the best quality in all situations, and it is quite foreseeable that various different sets of values may be used within the context of these embodiments.
[0198] In the example above, a single regularization coefficient is defined. When considering the matrix solution described in the exemplary embodiment, regularization is performed by controlling the components of the diagonal matrix based on this regularization coefficient. In some other embodiments, finer adjustments of regularization may be applied, for example, so that different regularization coefficients or different regularization schemes are applied to different components of the diagonal matrix. In other words, regularization information may not be a single value, but rather a more sophisticated set of information.
[0199] The above description illustrates the present invention in the context of binaural rendering. However, the method presented in V Vilkamo, J., Backstrom, T., & Kuntz, A. (2013). Optimized covariance domain framework for time-frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411 provides a general mixing framework that can be applied to any input-output mixing case. This means that the target output may be a loudspeaker, ambisonics, or any other audio channel, in addition to the binaural target output. Similarly, exemplary embodiments illustrating the adaptation of regularization coefficients can be appropriately adapted to any other input-output mixing case. For example, mixing from two transport channels and metadata to the 7.1.4 output can use the method described in the above reference, and depending on the bitrate, different regularization coefficients should be used to obtain optimal quality. The output format may also be used to determine the selection of regularization coefficients. For example, some artifacts may be more noticeable when using loudspeaker rendering than when using binaural rendering. This, in turn, leads to the proposal of different regularization coefficient values for them.
[0200] In some alternative embodiments, after initial values are selected based on input parameters such as codec format and bitrate, additional adjustment steps may be performed on the regularization coefficient. These adjustments may only change the regularization coefficient by steps of absolute value (e.g., +0.1 or -0.1) or relative value (e.g., multiplying by 0.9). If such adjustments are performed, an additional step should also be introduced before outputting the final regularization coefficient. This step restricts the regularization coefficient to a specific range, e.g., 0.2 to 1.0, to avoid unintended values.
[0201] The source information for these adjustment steps can have many options. For example, the source information could be one or more of the following: Descriptive metadata for the MASA format or any similar data. For example, the source format may use different tables as described above, and can distinguish between channel-based audio capture and microphone capture. Also, for example, a cardioid pattern transport may be more susceptible to artifacts than an omnidirectional type transport, so a transport channel configuration may be used to make adjustments.
[0202] If objective measurements of the signal quality coefficient can be obtained on a frame-by-frame basis, these can be used to reduce or increase the regularization coefficient, depending on whether the quality is high or low, respectively.
[0203] In some embodiments, the order of selection of regularization coefficients in the presented method can be rearranged with little effort, since the resulting regularization coefficient is a simple combination of multiple coefficients. Other practical implementations are equally valid.
[0204] In some embodiments, the codec format may signal, instead of directly indicating the format, as a combination of bits that indicate the type of spatial metadata and / or transport audio signal being used.
[0205] The above explanation suggests that adjustments are controlled between mixing the original signals and adding their uncorrelated versions in order to achieve the target covariance matrix, but it should be noted that the presence of uncorrelated signals is not always necessary. The presented method can also be used in situations where only a solution that mixes the original signals is used. In this case, the target covariance matrix is not necessarily achieved.
[0206] In the embodiments described above, covariance matrix-based rendering, as presented in the references above, was used as an example. However, it should be noted that the presented method can be used with other types of renderers as long as there is some kind of regularization of the gains. For example, in an alternative embodiment, the “direct” sound and the “ambient” sound may be rendered separately. In that case, prototype signals are created separately for each of them (e.g., using decorrelation for the ambient part), and target energies are also created separately for the direct and ambient parts. Gains for rendering the direct and ambient parts are then calculated by comparing the energies of the prototype signals with the target energies. These gains typically require regularization to avoid excessive gain of noise, as presented above for the main embodiment. These gains can be limited using the present invention. In this embodiment, it should be noted that the regularization balances artifacts from excessive gain of noise with some signal attenuation, rather than balancing mixing and decorrelation.
[0207] With respect to Figure 12, several examples of processing according to embodiments are shown. Processing is performed using a system as described herein, in which the original 5.1 sound is transmitted across two transport audio channels, and in the decoder, the audio is rendered into binaural sound using the transmitted spatial metadata. The spatial metadata and audio signal are encoded at two different bitrates, 48kbps and 256kbps, represented as two columns, 48kbps 1200 and 256kbps 1202.
[0208] The top row 1201 shows binaurally processed sounds, while the middle row 1203 and bottom row 1205 show only uncorrelated path sounds, in other words, M(b,n) and M rAfter formulating (b,n), the former is set to zero before applying them to the signal. The middle line 1203 shows the conventional solution where the regularization coefficient is not adapted based on the bitrate, but instead it is set to the default value of 1.0, which is safer in that codec artifacts in rendering are amplified to a minimum. Instead it uses decorrelation to a greater degree.
[0209] The line 1205 below illustrates processing by several embodiments where regularization is applied based on the bitrate. When the bitrate is higher, the regularization coefficient drops to 0.2, which causes the system to generate more decoherence required in the mixing approach, and therefore uses a small amount of decorrelation. In this case, because the bitrate is high, the codec artifacts are at a lower level and their amplification is inaudible. On the other hand, a small amount of decorrelation is applied, so better quality is provided.
[0210] In general, various embodiments of the present invention may be implemented in hardware, dedicated circuits, software, logic, or any combination thereof. For example, some embodiments may be implemented in hardware, while others may be implemented in firmware or software which may be executed by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Various embodiments of the present invention may be illustrated and described using block diagrams, flowcharts, or any other graphical representation, but it should be understood that these blocks, devices, systems, techniques, or methods described herein may, in non-limiting examples, be implemented in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers or other computing devices, or any combination thereof.
[0211] Embodiments of the present invention may be implemented by computer software or hardware executable by a data processor of a mobile device, such as a processor entity, or by a combination of software and hardware. Furthermore, it should be noted that any block of the logic flow as shown in the figure may represent a program step, interconnected logic circuits, blocks, and functions, or a combination of a program step and logic circuits, blocks, and functions. The software may be stored on physical media such as memory chips or memory blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants, CDs, etc.
[0212] Memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technical environment and may include, in non-limiting examples, one or more of the following: general-purpose computers, dedicated computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits, and processors based on multi-core processor architectures.
[0213] Embodiments of the present invention may be put into practice in various components, such as integrated circuit modules. Integrated circuit design is generally a highly automated process. Complex and powerful software tools are available to translate logic-level designs into semiconductor circuit designs ready to be etched and formed on semiconductor substrates.
[0214] Programs such as those offered by Synopsys, Inc. in Mountain View, California, and Cadence Design in San Jose, California, automatically route conductors and place components on a semiconductor chip using well-established design rules and a library of pre-stored design modules. Once the design for the semiconductor circuit is complete, the resulting design in a standardized electronic format (e.g., Opus or GDSII) may be sent to a semiconductor manufacturing facility or "fab" for production.
[0215] As used in this application, the term “circuit” may mean one or more or all of the following: (a) Implementation of the circuit using only hardware (such as implementation using only analog and / or digital circuits), (b) (If applicable) Combinations of hardware circuits and software, such as the following: (i) combination of analog and / or digital hardware circuits with software / firmware, (ii) Any part of a hardware processor having software (including digital signal processors, software, and memory that work together to enable a device such as a mobile phone or server to perform various functions), Hardware circuits and processors, such as microprocessors or parts of microprocessors, that require software (e.g., firmware) for operation but do not need to be present when the software is not required for operation.
[0216] This definition of "circuit" applies to all uses of the term in this application, including in any claim. As a further example, as used in this application, the term "circuit" also covers the implementation of merely a hardware circuit or processor (or more processors), or a part of a hardware circuit or processor, and the software and / or firmware associated with it (or them). The term "circuit" also covers, for example, a baseband integrated circuit or processor integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or network device, as applicable to the elements of a particular claim.
[0217] When used herein, the term "non-transient" refers to a limitation of the medium itself (i.e., tangible and not signaling), rather than a limitation on the persistence of data storage (e.g., RAM vs. ROM).
[0218] When used herein, “i.e., at least one of the <list of two or more elements>” and “at least one of the <list of two or more elements>,” as well as similar phrases where lists of two or more elements are linked by “and” or “or,” mean at least one of the elements, or at least two or more of the elements, or at least all of the elements.
[0219] The foregoing description, by illustrative and non-limiting examples, has provided a sufficient and useful description of exemplary embodiments of the invention. However, when read in conjunction with the accompanying drawings and claims, various modifications and adaptations may become apparent to those skilled in the art in consideration of the foregoing description. Nevertheless, all such and similar modifications of the teachings of the invention would still fall within the scope of the invention as defined in the claims.
Claims
1. A method for generating an audio signal, To acquire an input audio signal that includes at least two audio channels, To obtain at least one input characteristic parameter associated with the input audio signal, Determining at least one control parameter based at least one of the aforementioned input characteristic parameters, Determining processing parameters based at least on the at least two audio channels, wherein the determination of the processing parameters is controlled at least in part on the at least one control parameter. The audio signal is generated based at least on the two audio channels and the processing parameters. Methods that include...
2. The method according to claim 1, wherein the input audio signal is a spatial audio signal, and the spatial audio signal further includes at least one spatial parameter associated with the at least two audio channels.
3. The method according to claim 2, wherein determining the processing parameters based on at least two audio channels further comprises determining the processing parameters based on at least one spatial parameter.
4. The at least one input characteristic parameter is, The bitrate associated with the at least one input audio signal, A codec format that indicates the origin of at least one input audio signal, and Configuration of the at least one input audio signal The method according to any one of claims 1 to 3, comprising at least one of the above.
5. The configuration of the at least one input audio signal is, The source format that indicates the original format in which the input audio signal was created or the input format, Description of transport channels, Number of channels, Channel distance, and Channel angle The method according to claim 4, comprising at least one of the following.
6. The codec format for indicating the origin of the at least one input audio signal is: A metadata-assisted spatial audio stream origin codec format value that indicates that the at least one input audio signal represents a metadata-assisted spatial audio signal, A multichannel audio stream origin codec format value that indicates that the at least one input audio signal represents a multichannel audio signal, An audio object codec format value that indicates that the input audio signal represents an audio object, and Ambisonic format value that indicates the input audio signal represents an ambisonic audio signal The method according to claim 4 or 5, comprising at least one of the following.
7. The aforementioned at least one spatial parameter is Information describing the spatial composition of sound with respect to at least one input audio signal, including a direction parameter configured to indicate where the sound is coming from, and a ratio parameter configured to indicate the portion of the sound coming from the direction, Information describing the characteristics of the original multi-channel or multi-object sound scene, including at least one of channel level, object level, inter-channel correlation, inter-object correlation, and object direction. Processing coefficients related to acquiring a spatial audio format signal based at least on the aforementioned at least one input audio signal. The method according to any one of claims 4 to 6, as dependent on claim 2, comprising at least one of the above.
8. The method according to any one of claims 4 to 7, wherein the codec format for indicating the origin of the at least one input audio signal includes an indication that the origin is unknown or undefined.
9. Determining the at least one control parameter based at least on the at least one input characteristic parameter means that A first set of control parameter values is determined based on a first input characteristic parameter among the at least one input characteristic parameter. Based on the second input characteristic parameter among the at least one input characteristic parameter, one control parameter value is selected from the first set of control parameter values. The method according to any one of claims 1 to 8, including the method described in any one of claims 1 to 8.
10. The method according to claim 9, as dependent on claim 4, wherein the first input characteristic parameter of the at least one input characteristic parameter is a codec format, and the second input characteristic parameter of the at least one input characteristic parameter is a bitrate.
11. Determining processing parameters based at least on the at least two audio channels, wherein the determination of the processing parameters is controlled at least in part on the at least one control parameter, A process matrix is generated based at least on the at least two audio channels, wherein the generation of the process matrix is controlled at least in part on the at least one control parameter. The method according to any one of claims 1 to 10, including the method described above.
12. The method according to claim 11, wherein generating the processing matrix which is controlled at least in part on the control parameters includes regularizing the generation of the processing matrix at least on the control parameters.
13. The method according to claim 12, wherein generating the audio signal based on at least the at least two audio channels and the processing parameters comprises processing the at least two audio channels using the regularized processing matrix to generate the audio signal.
14. The method according to claim 12 or 13, wherein generating the processing matrix controlled at least partially on the at least one control parameter comprises generating the components of a diagonal matrix based at least on the values of at least two control parameters.
15. The method according to any one of claims 1 to 14, wherein obtaining at least one input characteristic parameter associated with the input audio signal includes deduction of the at least one input characteristic parameter from the input audio signal.
16. The method according to any one of claims 1 to 14, wherein obtaining at least one input characteristic parameter associated with the input audio signal includes receiving the at least one input characteristic parameter as part of the input audio signal or as configuration information.
17. Receiving the aforementioned at least one input characteristic parameter as part of the input audio signal means Receiving the encoded input characteristic parameters as part of the input audio signal, Decoding the encoded input characteristic parameters to obtain the input characteristic parameters. The method according to claim 16, further comprising:
18. Generating the audio signal based at least on the two audio channels and the processing parameters is: Binaural audio signals, and Multichannel audio signal The method according to any one of claims 1 to 17, comprising generating at least one of the following.
19. An apparatus comprising means for performing the method described in any one of claims 1 to 18.
20. A computer program, when executed by a device, includes instructions that cause the device to perform the method according to any one of claims 1 to 18.
21. A device for generating audio signals, To acquire an input audio signal that includes at least two audio channels, To obtain at least one input characteristic parameter associated with the input audio signal, Determining at least one control parameter based at least one of the aforementioned input characteristic parameters, Determining processing parameters based at least on the at least two audio channels, wherein the determination of the processing parameters is controlled at least in part on the at least one control parameter. The audio signal is generated based at least on the two audio channels and the processing parameters. An apparatus comprising means configured to perform a certain action.
22. An apparatus for generating an audio signal, the apparatus comprising at least one processor and at least one memory for storing instructions, wherein the instructions are executed by the at least one processor. To acquire an input audio signal that includes at least two audio channels, To obtain at least one input characteristic parameter associated with the input audio signal, Determining at least one control parameter based at least one of the aforementioned input characteristic parameters, Determining processing parameters based at least on the at least two audio channels, wherein the determination of the processing parameters is controlled at least in part on the at least one control parameter. The audio signal is generated based at least on the two audio channels and the processing parameters. An apparatus that causes the aforementioned apparatus to perform at least the above.