Audio rendering of spatial audio
By obtaining the input characteristic parameters of the audio signal and dynamically controlling the regularization process of the rendering gain, a high-quality spatial audio signal is generated, which solves the problem of poor audio quality in the existing technology and improves the audio rendering effect.
Patent Information
- Application Number
- CN202480011718.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-08
- Filing Date
- 2024-01-16
- Publication Date
- 2025-09-19
AI Technical Summary
Existing spatial audio rendering methods cannot effectively control regularization under different bit rates and codec formats, resulting in poor audio quality. In particular, coding artifacts and microphone noise amplification affect the perceptual quality in low signal-to-noise ratio conditions.
By obtaining input characteristic parameters of the audio signal, such as bit rate and codec format, the regularization process of the rendering gain is dynamically controlled, and a processing matrix is generated to process the audio channels to generate a high-quality spatial audio signal.
Improved audio quality under different input signals and bit rates, reduced the impact of encoding artifacts and microphone noise, and enhanced the perceived audio experience.
Smart Images

Figure CN120677524A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to audio rendering for spatial audio and applying regularization in rendering, but is not limited to apparatus and methods for configuring hybrid solution regularization for rendering. Background Art
[0002] There are many ways to capture spatial audio. One option is to use a microphone array (e.g., as part of a mobile device) to capture spatial audio. Using the microphone signals, a spatial analysis of the sound scene can be performed to determine spatial metadata in the frequency bands. Furthermore, the microphone signals can be used to determine a transmitted audio signal. The spatial metadata and the transmitted audio signal can be combined to form a spatial audio stream.
[0003] Metadata Assisted Spatial Audio (MASA) is an example of a spatial audio stream. It is one of the input formats that the upcoming Immersive Voice and Audio Services (IVAS) codec will support. It uses an audio signal along with corresponding spatial metadata (containing, for example, directional and direct-to-total energy ratios in frequency bands) and descriptive metadata (containing additional information related to, for example, the original captured and (transmitted) audio signal). A MASA stream can be obtained, for example, by capturing spatial audio with a microphone of, for example, a mobile device, where the set of spatial metadata is estimated based on the microphone signal. A MASA stream can also be obtained from other sources, such as a specific spatial audio microphone (such as Ambisonics), a studio mix (e.g., a 5.1 mix), or other content converted to a suitable format. A MASA tool inside the codec can also be used to encode a multi-channel signal by converting the multi-channel signal into a MASA stream and encoding the stream. Summary of the Invention
[0004] According to a first aspect, a method for generating an audio signal is provided, the method comprising: obtaining an input audio signal comprising at least two audio channels; obtaining at least one input characteristic parameter associated with the input audio signal; determining at least one control parameter based at least on the at least one input characteristic parameter; determining a processing parameter based at least on the at least two audio channels, wherein the determination of the processing parameter is controlled at least in part based on the at least one control parameter; and generating the audio signal based at least on the at least two audio channels and the processing parameter.
[0005] The input audio signal may be a spatial audio signal further comprising at least one spatial parameter associated with at least two audio channels.
[0006] Determining the processing parameters based at least on the at least two audio channels may further include determining the processing parameters further based at least on at least one spatial parameter.
[0007] The at least one input characteristic parameter may include at least one of: a bit rate associated with the at least one input audio signal; a codec format indicating a source of the at least one input audio signal; and a configuration of the at least one input audio signal.
[0008] The configuration of the at least one input audio signal may include at least one of: a source format indicating an original format or input format from which the input audio signal was created; a transmission channel description; a number of channels; a channel distance; and a channel angle.
[0009] The codec format indicating the source of at least one input audio signal may include at least one of the following: a metadata-assisted spatial audio stream source codec format value indicating that the at least one input audio signal represents a metadata-assisted spatial audio signal; a multi-channel audio stream source codec format value indicating that the at least one input audio signal represents a multi-channel audio signal; an audio object codec format value indicating that the input audio signal represents an audio object; and a panoramic surround sound format value indicating that the input audio signal represents a panoramic surround sound audio signal.
[0010] At least one spatial parameter may include at least one of the following: information describing the organization of sound in space relative to at least one input audio signal, the information including: a direction parameter configured to indicate where the sound arrives from; and a ratio parameter configured to indicate the portion of the sound arriving from that direction; information describing properties of the original multi-channel or multi-object sound scene, the information including at least one of the following: channel level; object level; inter-channel correlation; inter-object correlation; and object direction; processing coefficients associated with obtaining a spatial audio format signal based on at least at least one input audio signal.
[0011] The codec format indicating the source of the at least one input audio signal may include an indication that the source is unknown or undefined.
[0012] Determining at least one control parameter based at least on at least one input characteristic parameter may include: determining a first set of control parameter values based on a first input characteristic parameter among the at least one input characteristic parameter; and selecting a control parameter value from the first set of control parameter values based on a second input characteristic parameter among the at least one input characteristic parameter.
[0013] A first input characteristic parameter of the at least one input characteristic parameter may be a codec format, and a second input characteristic parameter of the at least one input characteristic parameter may be a bit rate.
[0014] Determining processing parameters based at least on at least two audio channels, wherein the determination of the processing parameters is controlled at least in part based on at least one control parameter, may include generating a processing matrix based at least on the at least two audio channels, wherein the generation of the processing matrix may be controlled at least in part based on the at least one control parameter.
[0015] Generating the processing matrix to be controlled based at least in part on the control parameter may include regularizing generation of the processing matrix based at least on the control parameter.
[0016] Based at least on the at least two audio channels and the processing parameters, generating the audio signal may include processing the at least two audio channels using a regularized processing matrix to generate the audio signal.
[0017] Generating a processing matrix to be controlled based at least in part on the control parameter may include generating entries of a diagonal matrix based at least on at least two control parameter values.
[0018] Obtaining at least one input characteristic parameter associated with the input audio signal may include inferring the at least one input characteristic parameter from the input audio signal.
[0019] Obtaining at least one input characteristic parameter associated with the input audio signal may include receiving the at least one input characteristic parameter as part of the input audio signal or as configuration information.
[0020] Receiving the at least one input characteristic parameter as part of the input audio signal may further comprise: receiving an encoded input characteristic parameter as part of the input audio signal; and decoding the encoded input characteristic parameter to obtain the input characteristic parameter.
[0021] Based at least on the at least two audio channels and the processing parameters, generating the audio signal may include generating at least one of: a binaural audio signal; and a multi-channel audio signal.
[0022] According to a second aspect, a device for generating an audio signal is provided, the device comprising components configured to: obtain an input audio signal comprising at least two audio channels; obtain at least one input characteristic parameter associated with the input audio signal; determine at least one control parameter based at least on the at least one input characteristic parameter; determine a processing parameter based at least on the at least two audio channels, wherein the determination of the processing parameter is controlled at least in part based on the at least one control parameter; and generate an audio signal based at least on the at least two audio channels and the processing parameter.
[0023] The input audio signal may be a spatial audio signal further comprising at least one spatial parameter associated with at least two audio channels.
[0024] The component configured to determine the processing parameter based at least on at least two audio channels may further be configured to determine the processing parameter based at least on at least one spatial parameter.
[0025] The at least one input characteristic parameter may include at least one of: a bit rate associated with the at least one input audio signal; a codec format indicating a source of the at least one input audio signal; and a configuration of the at least one input audio signal.
[0026] The configuration of the at least one input audio signal may include at least one of: a source format indicating an original format or input format from which the input audio signal was created; a transmission channel description; a number of channels; a channel distance; and a channel angle.
[0027] The codec format indicating the source of at least one input audio signal may include at least one of the following: a metadata-assisted spatial audio stream source codec format value indicating that the at least one input audio signal represents a metadata-assisted spatial audio signal; a multi-channel audio stream source codec format value indicating that the at least one input audio signal represents a multi-channel audio signal; an audio object codec format value indicating that the input audio signal represents an audio object; and a panoramic surround sound format value indicating that the input audio signal represents a panoramic surround sound audio signal.
[0028] At least one spatial parameter may include at least one of the following: information describing the organization of sound in space relative to at least one input audio signal, the information including: a direction parameter configured to indicate where the sound arrives from; and a ratio parameter configured to indicate the portion of the sound arriving from that direction; information describing properties of the original multi-channel or multi-object sound scene, the information including at least one of the following: channel level; object level; inter-channel correlation; inter-object correlation; and object direction; processing coefficients associated with obtaining a spatial audio format signal based on at least at least one input audio signal.
[0029] The codec format indicating the source of the at least one input audio signal may include an indication that the source is unknown or undefined.
[0030] The component configured to determine at least one control parameter based at least on at least one input characteristic parameter can be further configured to: determine a first set of control parameter values based on a first input characteristic parameter among the at least one input characteristic parameter; and select a control parameter value from the first set of control parameter values based on a second input characteristic parameter among the at least one input characteristic parameter.
[0031] A first input characteristic parameter of the at least one input characteristic parameter may be a codec format, and a second input characteristic parameter of the at least one input characteristic parameter may be a bit rate.
[0032] The component configured to determine processing parameters based at least on at least two audio channels, wherein the determination of the processing parameters is controlled at least in part based on at least one control parameter, may be further configured to generate a processing matrix based at least on the at least two audio channels, wherein the generation of the processing matrix may be controlled at least in part based on the at least one control parameter.
[0033] The component configured to generate the processing matrix to be controlled at least in part based on the control parameter may be further configured to regularize the generation of the processing matrix based at least on the control parameter.
[0034] The component configured to generate the audio signal based at least on the at least two audio channels and the processing parameters may be further configured to process the at least two audio channels using a regularized processing matrix to generate the audio signal.
[0035] The component configured to generate the processing matrix to be controlled at least in part based on the control parameter may be further configured to generate entries of the diagonal matrix based at least on at least two control parameter values.
[0036] The means configured to obtain at least one input characteristic parameter associated with the input audio signal may be further configured to infer the at least one input characteristic parameter from the input audio signal.
[0037] The component configured to obtain at least one input characteristic parameter associated with the input audio signal may be further configured to receive the at least one input characteristic parameter as part of the input audio signal or as configuration information.
[0038] The component configured to receive at least one input characteristic parameter as part of the input audio signal may be further configured to: receive an encoded input characteristic parameter as part of the input audio signal; and decode the encoded input characteristic parameter to obtain the input characteristic parameter.
[0039] The component configured to generate the audio signal based at least on the at least two audio channels and the processing parameters may be further configured to generate at least one of: a binaural audio signal; and a multi-channel audio signal.
[0040] According to a third aspect, a device for generating an audio signal is provided, the device comprising at least one processor and at least one memory storing instructions, the instructions, when executed by the at least one processor, causing the system to at least perform: obtaining an input audio signal comprising at least two audio channels; obtaining at least one input characteristic parameter associated with the input audio signal; determining at least one control parameter based at least on the at least one input characteristic parameter; determining a processing parameter based at least on the at least two audio channels, wherein the determination of the processing parameter is controlled at least in part based on the at least one control parameter; and generating an audio signal based at least on the at least two audio channels and the processing parameter.
[0041] The input audio signal may be a spatial audio signal further comprising at least one spatial parameter associated with at least two audio channels.
[0042] The apparatus caused to determine the processing parameters based at least on at least two audio channels may be further caused to determine the processing parameters based further at least on at least one spatial parameter.
[0043] The at least one input characteristic parameter may include at least one of: a bit rate associated with the at least one input audio signal; a codec format indicating a source of the at least one input audio signal; and a configuration of the at least one input audio signal.
[0044] The configuration of the at least one input audio signal may include at least one of: a source format indicating an original format or input format from which the input audio signal was created; a transmission channel description; a number of channels; a channel distance; and a channel angle.
[0045] The codec format indicating the source of at least one input audio signal may include at least one of the following: a metadata-assisted spatial audio stream source codec format value indicating that the at least one input audio signal represents a metadata-assisted spatial audio signal; a multi-channel audio stream source codec format value indicating that the at least one input audio signal represents a multi-channel audio signal; an audio object codec format value indicating that the input audio signal represents an audio object; and a panoramic surround sound format value indicating that the input audio signal represents a panoramic surround sound audio signal.
[0046] At least one spatial parameter may include at least one of the following: information describing the organization of sound in space relative to at least one input audio signal, the information including: a direction parameter configured to indicate where the sound arrives from; and a ratio parameter configured to indicate the portion of the sound arriving from that direction; information describing properties of the original multi-channel or multi-object sound scene, the information including at least one of the following: channel level; object level; inter-channel correlation; inter-object correlation; and object direction; processing coefficients associated with obtaining a spatial audio format signal based on at least at least one input audio signal.
[0047] The codec format indicating the source of the at least one input audio signal may include an indication that the source is unknown or undefined.
[0048] The device that is caused to determine at least one control parameter based at least on at least one input characteristic parameter can be further caused to perform: determining a first set of control parameter values based on a first input characteristic parameter among the at least one input characteristic parameter; and selecting a control parameter value from the first set of control parameter values based on a second input characteristic parameter among the at least one input characteristic parameter.
[0049] A first input characteristic parameter of the at least one input characteristic parameter may be a codec format, and a second input characteristic parameter of the at least one input characteristic parameter may be a bit rate.
[0050] The apparatus may be further configured to determine, based at least on at least two audio channels, processing parameters, wherein the determination of the processing parameters is controlled at least in part based on at least one control parameter. The apparatus may be further configured to generate, based at least on at least two audio channels, a processing matrix, wherein the generation of the processing matrix may be controlled at least in part based on at least one control parameter.
[0051] The apparatus caused to perform generating the processing matrix is controlled based at least in part on the control parameter and may be further caused to perform, based at least on the control parameter, regularizing the generation of the processing matrix.
[0052] The apparatus caused to generate an audio signal based at least on the at least two audio channels and the processing parameters may be further caused to process the at least two audio channels using the regularized processing matrix to generate the audio signal.
[0053] The apparatus caused to generate the processing matrix is controlled at least in part based on the control parameter and may be further caused to generate entries of the diagonal matrix based at least on at least two control parameter values.
[0054] The apparatus being caused to perform obtaining at least one input characteristic parameter associated with the input audio signal may be further caused to perform inferring the at least one input characteristic parameter from the input audio signal.
[0055] The apparatus caused to perform obtaining at least one input characteristic parameter associated with the input audio signal may be further caused to perform receiving the at least one input characteristic parameter as part of the input audio signal or as configuration information.
[0056] The apparatus caused to perform receiving at least one input characteristic parameter as part of an input audio signal may be further caused to perform: receiving an encoded input characteristic parameter as part of the input audio signal; and decoding the encoded input characteristic parameter to obtain the input characteristic parameter.
[0057] The apparatus caused to generate an audio signal based at least on the at least two audio channels and the processing parameters may be further caused to generate at least one of: a binaural audio signal; and a multi-channel audio signal.
[0058] According to a fourth aspect, a device for generating an audio signal is provided, the device comprising: a component for obtaining an input audio signal comprising at least two audio channels; a component for obtaining at least one input characteristic parameter associated with the input audio signal; a component for determining at least one control parameter based at least on the at least one input characteristic parameter; a component for determining a processing parameter based at least on the at least two audio channels, wherein the determination of the processing parameter is controlled at least in part based on the at least one control parameter; and a component for generating the audio signal based at least on the at least two audio channels and the processing parameter.
[0059] According to a fifth aspect, a device for generating an audio signal is provided, the device comprising: an acquisition circuit configured to obtain an input audio signal comprising at least two audio channels; an acquisition circuit configured to obtain at least one input characteristic parameter associated with the input audio signal; a determination circuit configured to determine at least one control parameter based at least on the at least one input characteristic parameter; a determination circuit configured to determine a processing parameter based at least on the at least two audio channels, wherein the determination of the processing parameter is controlled at least in part based on the at least one control parameter; and generating the audio signal based at least on the at least two audio channels and the processing parameter.
[0060] According to a sixth aspect, a computer program comprising instructions [or a computer-readable medium comprising program instructions] is provided, wherein the instructions or program instructions are used to cause an apparatus for generating an audio signal to perform at least the following operations: obtain an input audio signal comprising at least two audio channels; obtain at least one input characteristic parameter associated with the input audio signal; determine at least one control parameter based at least on the at least one input characteristic parameter; determine a processing parameter based at least on the at least two audio channels, wherein the determination of the processing parameter is controlled at least in part based on the at least one control parameter; and generate an audio signal based at least on the at least two audio channels and the processing parameter.
[0061] According to a seventh aspect, a non-transitory computer-readable medium comprising program instructions is provided, the program instructions being used to cause an apparatus for generating an audio signal to perform at least the following operations: obtain an input audio signal comprising at least two audio channels; obtain at least one input characteristic parameter associated with the input audio signal; determine at least one control parameter based at least on the at least one input characteristic parameter; determine a processing parameter based at least on the at least two audio channels, wherein the determination of the processing parameter is controlled at least in part based on the at least one control parameter; and generate an audio signal based at least on the at least two audio channels and the processing parameter.
[0062] According to an eighth aspect, a computer-readable medium comprising program instructions is provided, the program instructions being used to cause an apparatus for generating an output audio signal to perform at least the following operations: obtain an input audio signal comprising at least two audio channels; obtain at least one input characteristic parameter associated with the input audio signal; determine at least one control parameter based at least on the at least one input characteristic parameter; determine a processing parameter based at least on the at least two audio channels, wherein the determination of the processing parameter is controlled at least in part based on the at least one control parameter; and generate an audio signal based at least on the at least two audio channels and the processing parameter.
[0063] An apparatus comprises means for performing the actions of the method as described above.
[0064] An apparatus is configured to perform the actions of the method described above.
[0065] A computer program comprising program instructions for causing a computer to execute the method described above.
[0066] A computer program product stored on a medium may cause an apparatus to perform the method as described herein.
[0067] An electronic device may include an apparatus as described herein.
[0068] A chipset may include the apparatus as described herein.
[0069] The embodiments of the present application are intended to solve the problems associated with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings, in which:
[0071] Figures 1 to 3 schematically illustrates an example system for capturing or otherwise obtaining a spatial audio signal in the form of a transmitted audio signal and spatial metadata;
[0072] Figure 4 schematically illustrates an example system for encoding a spatial audio signal in the form of a transmission audio signal and spatial metadata, and for playing back the spatial audio signal, suitable for implementing some embodiments;
[0073] Figure 5 Schematically illustrating an example system for encoding a spatial audio signal in the form of a transmission audio signal and spatial metadata based on multiple operation modes and for playing back the spatial audio signal, suitable for implementing some embodiments;
[0074] Figure 6 schematically illustrates an example playback apparatus suitable for implementing some embodiments;
[0075] Figure 7 Schematically illustrating how Figure 5 The example decoder shown in ;
[0076] Figure 8 Shown according to some embodiments Figure 7 A flowchart of the operation of the example decoder apparatus shown in FIG;
[0077] Figure 9 Schematically illustrating how Figure 5 The example spatial synthesizer shown in ;
[0078] Figure 10 Shown according to some embodiments Figure 9 A flowchart of the operation of an example spatial synthesizer shown in ;
[0079] Figure 11 Showing how according to some embodiments Figure 9 A flowchart of the operation of the example regularization factor determiner shown in ; and
[0080] Figure 12 Example processing output is shown. DETAILED DESCRIPTION
[0081] Suitable apparatus and possible mechanisms for rendering a suitable output audio signal from a parametric spatial audio stream (or signal) from a captured or otherwise obtained audio signal are described in more detail below.
[0082] As mentioned above, Metadata Assisted Spatial Audio (MASA) is an example of a parametric spatial audio format and representation suitable as an input format for IVAS.
[0083] An audio representation can be thought of as consisting of "N channels + spatial metadata." It is a scene-based audio format particularly well-suited for spatial audio capture on real-world devices such as smartphones. The idea is to describe the sound scene in terms of sound directions and, for example, energy ratios that vary over time and frequency. Sound energy that is not defined (described) by direction is described as diffuse (coming from all directions).
[0084] As described above, the spatial metadata associated with the audio signal may include multiple parameters per time-frequency tile (such as multiple directions and a direct-to-total ratio, spread coherence, distance, etc. associated with each direction (or directional value). The spatial metadata may also include other parameters, or may be associated with other parameters that are considered non-directional (such as surround coherence, diffuse-to-total energy ratio, remainder-to-total energy ratio), but when combined with directional parameters can be used to define the characteristics of the audio scene. For example, a reasonable design choice that can produce good quality output is one in which the spatial metadata includes one or more directions for each time-frequency element (also referred to as a time-frequency portion or time-frequency tile). In addition, associated with each direction, the metadata includes other parameters such as direct-to-total ratio, spread coherence, distance value.
[0085] As described above, parameterized spatial metadata representations can use multiple concurrent spatial directions. For MASA, the recommended maximum number of concurrent directions is two. For each concurrent direction, there can be associated parameters such as: direction index; direct to total ratio; extended coherence; and distance. In some embodiments, other parameters are defined, such as diffuse to total energy ratio; surround coherence; and residual to total energy ratio.
[0086] Parameterized spatial metadata values are available for each time-frequency tile (the MASA format defines 24 frequency bands and 4 time subframes in each frame). The frame size in IVAS is 20ms. In addition, current MASA supports 1 or 2 directions for each time-frequency tile.
[0087] Example metadata parameters could be:
[0088] Format descriptor, which defines the MASA format used for IVAS. This can be stored as a 64-bit value and can be eight 8-bit ASCII characters: 01001001, 01010110, 01000001, 01010011, 01001101, 01000001, 01010011, 01000001, where these values are stored as eight consecutive 8-bit unsigned integers;
[0089] Channel audio format, which defines the subsequent fields of the combination stored in two bytes, where the value is stored as a single 16-bit unsigned integer;
[0090] The number of directions parameter defines the number of directions described by the spatial metadata. Each direction can be associated with a set of direction-related spatial metadata, as described later. The number of directions parameter can be one of a range of values. Currently, the possible value range defined is 1 or 2, but in some embodiments more than 2 directions can be used;
[0091] Number of channels, which defines the number of transmission channels used by the format. The number of channels parameter can be a value in a range of values. Currently, the possible value ranges defined are 1 or 2, but in some embodiments more than 2 channels can be used;
[0092] SourceFormat: This parameter describes the original or input format from which the MASA output is created (e.g., the input format is 5.1 multichannel).
[0093] In some cases, there may be other description or parameter fields based on the values of the "Number of Channels" and "Source Format" parameter fields. In some embodiments, zero padding may be applied when all allocated bits are unused.
[0094] An example of a MASA format spatial metadata parameter related to the number of directions could be:
[0095] Direction index, which defines the direction of arrival of the sound in the time-frequency parameter interval. Usually, the direction of arrival is expressed spherically with an accuracy of about 1 degree;
[0096] Directly for the total energy ratio, this parameter defines the energy ratio for the direction index (in other words, defines the energy ratio associated with the direction for the time-frequency subframe);
[0097] Spread coherence, which defines the energy spread for the directional index (in other words, a measure of the “width” of the directions for the time-frequency subframe);
[0098] Transport definitions, which describe the configuration of two transport channels;
[0099] channel angle, which describes the symmetrical angular position of the transmitted signal with a directional pattern;
[0100] Channel distance, which describes the distance between two transmission channels; and
[0101] Channel layout, when the source format is multi-channel, it describes the channel layout of the multi-channel source format.
[0102] Examples of MASA format spatial metadata parameters that are independent of the number of directions could be:
[0103] Diffuse to Total Energy Ratio, which defines the ratio of energy in non-directional sound to that in surrounding directions;
[0104] Surround coherence, which defines the coherence of non-directional sounds with surrounding directions;
[0105] Residual to Total Energy Ratio: This parameter defines the energy ratio of the residual (such as microphone noise) acoustic energy to satisfy the requirement that the sum of the energy ratios is 1.
[0106] Additionally, example spatial metadata bands may be:
[0107]
[0108] The MASA stream can be rendered to various outputs, such as a multi-channel speaker signal (eg, 5.1) or a binaural signal.
[0109] In “Vilkamo, J., An example rendering method is described in "Optimized covariance domain framework for time-frequency processing of spatial audio", Journal of the Audio Engineering Society, 61(6), 403-411. The rendering method is based on multi-channel mixing. The method processes a given audio signal in a frequency band to obtain a desired covariance matrix for the output signal in the frequency band. The covariance matrix contains the channel energies of all channels and the inter-channel relationships between all pairs of channels, i.e., cross-correlations and inter-channel phase differences. These features are known to convey perceptually relevant spatial characteristics of multi-channel sounds in various playback situations, such as binaural for headphones, surround speakers, panoramic surround sound, and crosstalk-canceling stereo.
[0110] As mentioned above, in “Vilkamo, J., T. and Kuntz, A. (2013), Optimized covariance domain framework for time-frequency processing of spatial audio, Journal of the Audio Engineering Society, 61(6), 403-411". The rendering method attempts to process an input audio signal (e.g., a decoded transmitted audio signal) so that the generated output audio signal (e.g., a binaural audio signal) has a desired covariance matrix. The target covariance matrix is determined based on the energy of the input audio signal and spatial metadata received by the renderer.
[0111] In practice, the rendering method first attempts to achieve the target covariance matrix by mixing the input audio signals. However, in order to prevent amplification of input audio signals with very large gain values, regularization is applied in the processing. For example, if the input audio signals are highly correlated (in other words, have highly similar signals), but the target covariance matrix is set to have low correlation, this will result in a processing gain that will over-amplify the irrelevant signal parts. If regularization is not applied, this will result in significant amplification of noise and other unwanted sounds in the input audio signal, resulting in poor audio quality.
[0112] Since the operation of obtaining the mixing gain is regularized, the target covariance matrix is not always achieved by mixing alone. When regularization is applied, some signal parts are not amplified as much as required to achieve the target, and therefore, some expected or desired signal energy is lost. As a result, these sounds will be effectively attenuated in the reproduction, and the desired incoherence between the output channels may not be achieved. The proposed solution to achieve the desired incoherence is to decorrelate the input signals to obtain incoherent signals, and process these signals with the mixing gain to obtain the covariance matrix of the lost signal parts. By these means, the target covariance matrix for the output signal is also obtained when the regularization limits the processing.
[0113] Both mixing and decorrelation have their advantages and disadvantages. Decorrelation has the disadvantage of affecting sound quality, especially when a large amount of decorrelated signals are present in the output. When a sound is decorrelated, its phase spectrum is modified to at least some degree, usually in a time-invariant manner to preserve pitch. This process is known to degrade the perceived quality of certain sounds, such as speech or applause.
[0114] On the other hand, mixing with too large a gain leads to problems such as over-amplification of small signal components, which in some cases may be primarily noise.
[0115] The amount of regularization can be used to control how much decorrelated energy is mixed into the output. If only light regularization is applied (i.e., the maximum allowed mixing gain is large), the output contains mostly a mixture of the input signals without much decorrelated signal. Conversely, if a lot of regularization is applied (i.e., the maximum allowed mixing gain is small), the output contains more decorrelated signal because large mixing gains are prevented.
[0116] The amount of regularization is therefore a trade-off between avoiding problems of mixing and decorrelation. The optimal amount of regularization depends on the situation, and several examples are discussed below.
[0117] The IVAS codec operates over a wide range of bit rates (from 13.2 kbps to 512 kbps). Especially at the lowest end, there are significant coding artifacts (such as music and other noise and distortion) in the transmitted audio signal received by the renderer. Therefore, a large amount of regularization is required to avoid amplification of the coding artifacts, which will make them more prominent and significantly reduce the perceived audio quality. On the other hand, at higher bit rates, having a large amount of regularization (more decorrelation) would be a suboptimal option. There are significantly fewer of those coding artifacts in the transmitted audio signal, so this will result in more decorrelation than required, resulting in suboptimal audio quality.
[0118] In addition, the IVAS codec supports multiple input formats. The MASA format often originates from mobile devices that typically have inexpensive microphones, and therefore there is often perceptible microphone noise in the MASA input signal. Therefore, a large amount of regularization is required in the renderer to avoid amplification of microphone noise. On the other hand, IVAS also supports multi-channel input (such as 5.1), which is typically recorded and produced in a recording studio with a professional microphone with less noise. Therefore, since different formats require different values to achieve a good experience range, the fixed setting for regularization is not optimal in this regard.
[0119] Known rendering methods may have fixed regularization with a value set for situations where the signal SNR characteristic is large. Fixed regularization values based on high SNR situations may not produce acceptable results for low SNR situations. For example, a spatial audio codec operating at a low bitrate, such as 32 kbps or lower, may not benefit from a regularization value set for a high bitrate, low-noise input signal.
[0120] As a result, a fixed regularization value produces a rendering that is optimal only for certain types of input signals and / or for certain bit rates at which the transmitted audio signal is encoded. For other types of input signals and / or bit rates, there may be too much regularization (leading to too much decorrelation) or too little regularization (leading to amplification of noise), both of which lead to a degradation of the perceived audio quality.
[0121] Too much decorrelation produces a sound that can be perceived as reverberant and lacking in appeal / involvement, while amplification of noise produces a reproduced sound with perceptual artifacts.
[0122] The following embodiments and concepts as discussed in the application herein are to render spatial audio from a parameterized spatial audio stream (audio signal and associated spatial metadata), or more generally, at least one audio signal comprising two or more audio channels. In these embodiments, a rendering method (and therefore a renderer device) is proposed that enables improved audio quality by controlling regularization in the determination of rendering gains based on input characteristics (or input parameters) associated with at least one audio signal (such as bitrate and / or codec format). The purpose of the controlled regularization is to minimize artifacts caused by decorrelation (increased reverberation and loss of engagement) and over-amplification of noise (perception of loud music and other noise), thereby improving the perceived audio quality by alleviating the perception of those artifacts.
[0123] In some embodiments, this may be achieved by obtaining a (parameterized spatial) audio stream; obtaining input characteristics; determining a regularization value based on the input characteristics; determining a regularized rendering gain based on the regularization value, one or more audio signals, and associated spatial metadata; and rendering a spatial audio signal (such as a binaural audio signal) from the one or more audio signals using the adjusted rendering gain.
[0124] In some embodiments, input characteristics can mean various things in different embodiments. For example, input characteristics can be: the bitrate used to encode the parametric spatial audio stream (in the case where the parametric spatial audio stream is encoded / decoded before rendering); the codec format, which declares the source of the stream, for example, whether it is a MASA stream (i.e., typically from a mobile device) or whether it is created from a multi-channel signal (such as 5.1); the configuration of the one or more audio signals received by the renderer (e.g., whether they come from a cardioid microphone or an omnidirectional microphone).
[0125] As described above, embodiments discussed in greater detail below propose systems that utilize a spatial sound renderer and control regularization of rendering gain based on at least one input characteristic. In some embodiments, the renderer is implemented in a decoder, and two example input characteristics are considered. The input characteristics discussed here are codec format and bit rate. In other embodiments, the renderer may be implemented elsewhere than in the decoder, and other input characteristics may be present.
[0126] First, an example of the "Codec Format" input characteristic is described in further detail, along with how signals of different codec formats can be obtained. Next, the "Bitrate" input characteristic is described in further detail. Thereafter, the implementation of some embodiments within a decoder is described in further detail.
[0127] The codec format input characteristics can indicate the type of parametric spatial audio signal being sent. As previously mentioned, a parametric spatial audio signal is a signal that includes one or more audio signals and associated spatial metadata. The spatial metadata contains information indicating how the sound is spatially organized and is typically provided at a much lower sampling rate than the audio signals.
[0128] Examples of spatial metadata for different "codec formats" include:
[0129] Information describing the organization of sound in space. For example, in a frequency band, a direction parameter may indicate where a sound is arriving from, and a ratio parameter may indicate the portion of the sound arriving from that direction;
[0130] Information describing the properties of the original multi-channel or multi-object sound scene, for example, channel or object levels and inter-channel or inter-object correlations, or object orientations;
[0131] Processing coefficients related to obtaining certain spatial audio format signals (such as panoramic surround sound audio signals) based on the transmitted audio signal.
[0132] The examples shown above can be extended with other types of spatial metadata also being used.The codec format input characteristics may thus indicate the type of spatial metadata being transmitted.
[0133] In some embodiments, the codec format input characteristic may also (or alternatively) indicate the type or source of the transmitted audio signal.
[0134] For example, in some embodiments, the codec format input characteristic may indicate that the transmitted audio signal is a downmix of a multi-channel signal, or that the transmitted audio signal has been captured using a microphone array, or that the transmitted audio signal is a decomposition of a immersive surround sound audio signal. In some embodiments, the codec input characteristic may have a defined value that may indicate that the source of the transmitted audio signal is unknown or undefined.
[0135] In some embodiments, two different codec formats may have the same type of spatial metadata or the same type of transmitted audio signal as each other.
[0136] Any of the above audio formats can be rendered into a variety of output spatial audio formats, such as multi-channel speaker output (e.g., 5.1), head-tracked or non-head-tracked binaural output, or crosstalk-cancelled stereo. The following examples describe the synthesis of binaural audio signals. The examples described herein can be directly extended and applied to the reproduction of other output types.
[0137] Although the following embodiments are described with respect to applying regularization in a renderer whose input is parameterized spatial audio, it will be understood that some embodiments may be applied to any suitable renderer and rendering operation where regularization is applied to the input audio to be rendered and the regularization value is controlled based on input parameters such as bitrate and codec format.
[0138] about Figure 1 , shows an example of an apparatus configured to determine a parameterized spatial audio signal from a microphone array signal. The apparatus can be implemented on a mobile phone comprising a microphone array (integrated in the mobile phone), for example.
[0139] In such an embodiment, the microphone array signal 100 may be forwarded to a microphone array front end 101. The microphone array front end may be implemented using, for example, a method such as that proposed in US Pat. No. 1,087,3814, and is configured to output a transmission audio signal 102 and spatial metadata 104. The transmission audio signal 102 and the spatial metadata 104 may, for example, take the form of a MASA stream.
[0140] In some embodiments, the microphone array front end 101 is configured to measure the inter-microphone correlation at different delays between microphones in a frequency band, find the delay that maximizes the correlation, and determine the direction of the arriving sound based on the delay. In some embodiments, the microphone array front end can also determine a direct-to-total energy ratio parameter in the frequency band based on the correlation value.
[0141] In some embodiments, the microphone array front end can also provide a transmission audio signal 102. The process of determining the transmission audio signal 102 depends on what type or format of the microphone array signal is input. For example, if the microphone array signal 100 comes from a mobile device, the microphone array front end 101 can be configured to select the microphone signal from the left side of the device as the left transmission signal, and select another microphone signal from the right side of the device as the right transmission signal. In addition, in some embodiments, the microphone array front end 101 can also apply any suitable pre-processing steps, such as equalization, microphone noise suppression, wind noise suppression, automatic gain control, beamforming and other spatial filtering, ambient noise suppression, and limiters.
[0142] about Figure 2, shows another example apparatus configured to determine a parameterized spatial audio signal from a microphone array signal. In some embodiments, the microphone array providing the microphone array signal may be a dedicated microphone array, such as a microphone array providing a first-order panoramic surround (FOA) signal at its output. Thus, in some embodiments, the apparatus includes a panoramic surround decomposer 201 configured to receive a panoramic surround signal 200 and generate a transmission audio signal 102 and spatial metadata 104. In some embodiments, the panoramic surround decomposer 201 is configured to determine the spatial metadata. The spatial metadata may be determined, for example, using a method similar to directional audio coding (DirAC) such as described in "Pulkki, V. (2007), Spatial sound reproduction with directional audio coding, Journal of the Audio Engineering Society, 55(6), 503-516", and the transmission audio signal 104 may be a cardioid pattern generated from the panoramic surround (FOA) signal 200 in the left and right directions. Therefore, it should be noted that although the spatial metadata and the transmitted audio signal may be of similar type, the origin of the signals may be different.
[0143] In some embodiments, the panoramic surround sound decomposer 201 can be configured to receive a panoramic surround sound signal 200, such as FOA or high-order panoramic surround sound (HOA), and determine spatial metadata 104 so that it can reconstruct the panoramic surround sound signal in the decoder based on the transmission audio signal. For example, the panoramic surround sound signal can be transmitted as one or more channel transmission audio signals 102, where the first channel is the omnidirectional W component and the remaining channels are residual signals. The panoramic surround sound decomposer 201 determines these signals by determining prediction coefficients that enable the remaining channels to be predicted from the W component, and the residual signal is the prediction error. For those channels for which no residual is sent, parameters that enable the omnidirectional component for the residual signal to be processed by means of decorrelation can be estimated. All these prediction and processing parameters can form spatial metadata 104.
[0144] about Figure 3, shows another example apparatus configured to determine a parameterized spatial audio signal from a microphone array signal. In some embodiments, a downmixer and metadata determiner 301 is configured to receive a channel-based audio signal 300 and generate a transmission audio signal 102 and spatial metadata 104. The example channel-based audio signal 300 is a surround 5.1 sound or an audio object sound. The downmixer and metadata determiner 301 is configured to generate the spatial metadata 104 by converting the audio object signal and / or the audio channel into a FOA format and then determining the spatial metadata using DirAC or similar means. The downmixer and metadata determiner 301 is configured to generate the transmission audio signal 104 using amplitude panning such that channels / objects (and corresponding confusion cones) exceeding ±30 degrees are fully panned to the left and right channels, and channels / objects between ±30 degrees are panned to both channels according to their direction.
[0145] The codec format input parameter can thus indicate various aspects described above. Depending on the use case, it can indicate a completely different signal format (e.g., stereo transmission versus main residual signal definition), and / or different characteristics of the same signal format (e.g., stereo transmission signal from a downmix versus stereo transmission signal from a microphone), and / or it can indicate a metadata format. For example, the codec format input parameter can be an index into a predefined list of options that defines both the kind of transmitted audio signal and the spatial metadata used.
[0146] The bit rate input parameter can be configured to indicate the number of bits used to send information within a specific time frame. In network transmission, which also uses audio codecs such as IVAS, a common description is to measure the bits used per second (which is abbreviated as bps). As mentioned above, IVAS is expected to operate between bit rates of 13.2kbps and 512kbps. The bit rate can also be constant or change over time. Therefore, each transmission frame can have an equal number of bits or a variable number of bits. In IVAS, the transmission frame is expected to be 20ms, and the bit rate is expected to be constant in a steady-state situation. Therefore, the above bit rate will be converted to 264 bits / frame and 10240 bits / frame. It should be noted that this is usually the total bit rate, which is then further allocated to be used for signaling, audio channel coding, and metadata coding (if any). This further division can also be constant or variable.
[0147] In the following description, bitrate indicates the overall quality of the transmitted audio. Generally, as the bitrate increases, the overall quality of the transmitted audio improves in lossy audio codecs such as IVAS. Depending on the codec format, this quality improvement can come from the increased bits used to encode the audio channels, which directly reduces the presence of coding artifacts, or from the increased bits used to encode the spatial metadata, which improves the accuracy of the reproduced spatial scene. Thus, based on the bitrate, an indication of the expected overall quality is provided.
[0148] about Figure 4 and Figure 5 , shows an overview of an example apparatus illustrating the signal flow after the transmission audio signal 102 and the spatial metadata 104 are generated.
[0149] Figure 4 An example apparatus is shown that includes an encoder 401 configured to receive a transmit audio signal 102 and spatial metadata 104 and encode them to form a bitstream 402. Furthermore, the apparatus includes a decoder 402 configured to receive the bitstream 402 and output a spatial audio output 404. In other words, the system can be considered to operate in a first mode, in which processing software external to the encoder 401 determines and provides the transmit audio signal 102 and spatial metadata 104. For example, the software may be microphone array front-end software optimized to process microphone signals of a particular device.
[0150] Figure 5 Another example apparatus is shown including an encoder 401 configured to receive a transmission audio signal 102 and spatial metadata 104 and encode them to form a bitstream 402. Furthermore, the apparatus includes a decoder 402 configured to receive the bitstream 402 and output a spatial audio output 404. Furthermore, an encoder preprocessor 501 is shown preceding the encoder 401. The encoder preprocessor 501 is configured to receive the audio signal 500 and generate the transmission audio signal 102 and the spatial metadata 104. In other words, the system can be considered to operate in a second mode, wherein the encoding system 511 that performs the encoding also performs encoder preprocessing to obtain the transmission audio signal 102 and the spatial metadata 104. For example, an IVAS encoder (which includes both the encoder preprocessor 501 and the encoder 401) can be configured to accept a 5.1 or panoramic surround sound input to generate the transmission audio signal 102 and the spatial metadata 104 that are then encoded.
[0151] The encoder 401 thus forms a bitstream 402 that can be stored or transmitted. The bitstream contains the transmitted audio signal 102 and spatial metadata 104 in coded form. The audio signal can be encoded, for example, using the IVAS core codec, EVS or AAC encoder (or any other suitable encoder), and the metadata can be encoded, for example, using the methods proposed in US20210295855, US20220343928, US20220036906, EP4091166 (and / or any other suitable method). In addition, in some embodiments, the encoder 401 can multiplex the encoded audio and the encoded spatial metadata to form the bitstream 402.
[0152] In addition, the encoder 401 is configured to write or include input parameters such as the codec format input parameter used in the bitstream 402. For example, this can be directly signaled using signaling bits that define the codec format input parameter. For example, the codec format can be signaled based on two bits so that the value "00" is a multi-channel format, "01" is a MASA format, and "10" is a panoramic surround sound format (these are just example values). In addition, the encoder 401 can also write other input parameters into the bitstream 402, such as a bitrate input parameter. The bitrate input parameter can be constant or change over time.
[0153] In some embodiments, the bitrate input parameter may be inferred in the decoder from the size of the received bitstream frames.
[0154] In some embodiments, input parameter information such as codec format input parameter and bit rate input parameter information is provided over a different communication channel than the bitstream 402 that transmits the audio signal and spatial metadata.
[0155] The bitstream 402 may be forwarded to a decoder 403, which may be, for example, an IVAS decoder (or any other suitable decoder). The decoder 403 decodes the audio signal and metadata and renders a spatial audio output 404, which may be, for example, a binaural audio signal. The decoder 403 may be located on a different device than the encoder 401 or on the same device.
[0156] about Figure 6 , shows an example (decoder) apparatus for implementing some embodiments. Figure 6In the example shown in FIG, a mobile phone 601 is shown coupled to a headset 619 worn by a user of the mobile phone 601 via a wired or wireless connection 613. The wired or wireless connection 613 can enable an audio signal 615 to be passed to the headset 619 and a mono audio signal (or more than one audio signal) 617 to be passed from the headset 619 (which in some embodiments can also include head orientation / position information metadata from the headset).
[0157] In the following, example devices or apparatuses are as follows Figure 6 6. However, the example apparatus or device may also be any other suitable device, such as a tablet, a laptop, a computer, or any teleconferencing device. In addition, the apparatus or device may be the headset itself, thereby performing the operations of the illustrated mobile phone 601 through the headset.
[0158] In this example, mobile phone 601 includes processor 603. Processor 603 can be configured to execute various program codes, such as the methods described herein. Processor 603 is configured to communicate with headset 619 using a wired or wireless headset connection 615. In some embodiments, wired or wireless headset connection 615 is a Bluetooth 5.3 or Bluetooth LE audio connection. Connection 615 provides a (two-channel) audio signal from processor 603 to be reproduced to the user using headset 619.
[0159] The earphone 619 may be Figure 6 , or any other suitable type, such as in-ear or bone conduction headphones, or any other type of headphones. In some embodiments, headphones 619 have a head orientation sensor that provides head orientation information to processor 603 via connection 613. In some embodiments, the head orientation sensor is separate from headphones 619, and the data is provided separately to processor 603. In other embodiments, head orientation is tracked by other means, such as using a device 601 camera and machine learning-based facial orientation analysis.
[0160] In some embodiments, the processor 603 is coupled to a memory 605 having program code 607 that provides processing instructions according to the following embodiments. The program code 607 has instructions for processing the transmitted audio signal and spatial metadata received by the transceiver 611 or retrieved from the storage device 609 into a rendering form suitable for efficient output to headphones.
[0161] The transceiver 611 may communicate with other devices via any suitable known communication protocol. For example, in some embodiments, the transceiver may use a wireless network based on Long Term Evolution Advanced (LTE-Advanced, LTE-A) or New Radio (NR) (or may be referred to as 5G), Universal Mobile Telecommunications System (UMTS) Radio Access Network (UTRAN or E-UTRAN), Long Term Evolution (LTE, the same as E-UTRA), 2G network (legacy network technology), Wireless Local Area Network (WLAN or Wi-Fi), Worldwide Interoperability for Microwave Access (WiMAX), Personal Communications Service (PCS), Suitable radio access architectures for Wideband Code Division Multiple Access (WCDMA), systems using Ultra-Wideband (UWB) technology, sensor networks, Mobile Ad Hoc Networks (MANET), Cellular Internet of Things (IoT) RAN, and Internet Protocol Multimedia Subsystem (IMS), any other suitable options, and / or any combination thereof.
[0162] In some embodiments, Figure 6 The device may also be configured to also implement the functionality of the microphone array front end 101, the panoramic surround sound decomposer 201, the downmixer and metadata determiner 301, the encoder 401 and the encoder pre-processor 501, for example when the use case is two-way spatial audio communication with a remote device.
[0163] about Figure 7 , showing that Figure 5 A schematic diagram of an example decoder 403 is shown in FIG.
[0164] The input bitstream 402 is forwarded to the demultiplexer and decoder 701. The demultiplexer and decoder 701 is configured to demultiplex and decode the transport audio signal 706 and the spatial metadata 708 therefrom from the bitstream 402. The decoding corresponds to the encoding applied in the encoder 401. It should be noted that the transport audio signal 706 and the spatial metadata 708 are generally different from those presented previously, as they have been encoded and decoded. However, for simplicity, the same terminology (but with different reference numerals) is used to refer to them.
[0165] Furthermore, the demultiplexer and decoder 701 is configured to determine the bit rate 702 and codec format 704 parameters based on information within the bitstream 402. In one example, they may be indicated by dedicated bits in the bitstream 402. In other examples, one or both may be detected from the bitstream, or they may be conveyed over other channels, such as via RTP or session control.
[0166] The spatial metadata 708 , the transport audio signal 706 , the codec format 704 and the bit rate 702 are then provided to a spatial synthesizer 703 .
[0167] Spatial synthesizer 703 is configured to synthesize spatial audio output 404 in a desired format based on spatial metadata 708, transmitted audio signal 706, codec format 704, and bit rate 702. In some examples, only one of codec format 704 and bit rate 702 parameters is used, while in other examples, both are used. This information is used to balance the use of decorrelation and signal mixing when performing spatial synthesis. The output provided by spatial synthesizer 703 can be, for example, a binaural audio signal.
[0168] about Figure 8 , an example flow chart illustrating a method according to some embodiments Figure 7 The operation of the example decoder is shown in .
[0169] Thus, a first operation may include obtaining a (coded spatial audio) bitstream as shown by 801 .
[0170] The (encoded spatial audio) bitstream is then demultiplexed and decoded as indicated by 803 to generate a transport audio signal, spatial metadata and input parameters such as codec format and bitrate.
[0171] Thereafter, a spatial audio signal is synthesized from the transmission audio signal based on the spatial metadata and input parameters such as codec format and bit rate, as indicated by 805. In other words, the transmission audio signal is spatially processed based on the codec format, bit rate, and spatial metadata to generate the spatial audio signal.
[0172] Furthermore, as indicated by 807 , a spatial audio signal is output (eg, a binaural audio signal is output to headphones).
[0173] about Figure 9 , showing in more detail Figure 7 In some embodiments, the spatial synthesizer 703 is configured to receive the transmission audio signal 706 , the spatial metadata 708 , and input parameters such as the bit rate 702 and the codec format 704 .
[0174] In some embodiments, the spatial synthesizer 703 includes a forward filter bank 901. Figure 9 As shown in , the transmission audio signal 706 is provided to a forward filter bank 901, which transforms these signals into time-frequency representations. Any filter bank suitable for audio processing can be utilized, such as a complex modulation quadrature mirror filter (QMF) group or its low-delay variant or short-time Fourier transform (STFT). Similarly, the forward filter bank 901 can be realized by any suitable time-frequency converter.
[0175] In the example described herein, the filter bank is configured to have 60 frequency bins and sufficient stopband attenuation to avoid significant aliasing when the frequency bin signal is processed. In this configuration, all frequency bins can be processed independently of each other, except that some frequency bins share the same spatial metadata. For example, the spatial metadata can include spatial parameters in a limited number of frequency bands (e.g., 5, 12, or 24 frequency bands), and each of these frequency bands corresponds to a set of one or more frequency bins provided by the forward filter bank 901. Although this example shows a specific number of frequency bands, there can be any suitable number of frequency bands, for example, the number of frequency bands can be 5, 8, 12, 18, or 24 frequency bands.
[0176] The output of the forward filter bank 901 is a time-frequency transmission signal 906, which is provided to a decorrelator and mixer 907, a processing matrix determiner 905, and an input and target covariance matrix determiner 909. The time-frequency transmission signal 906S(b, t, i) can be represented as:
[0177]
[0178] Wherein, b is the frequency bin index, t is the time-frequency signal time index, and i is the channel index. In this example, the transmitted audio signal has exactly two channels, but in other examples there may be more than two.
[0179] In some embodiments, the spatial synthesizer 703 includes an input and target covariance matrix determiner 909. The input and target covariance matrix determiner 909 is configured to receive the spatial metadata 708 and the time-frequency transmission signal 906 and is configured to determine the covariance matrix 906. The covariance matrix 906 includes an input covariance matrix representing the time-frequency transmission signal 906 and a target covariance matrix representing the desired time-frequency spatial audio signal (to be rendered). The input covariance matrix can be determined or measured from the time-frequency transmission signal 906, which is labeled as a column vector x(b,t), where the rows indicate the transmission signal channels. In some embodiments, this is achieved by the following formula:
[0180]
[0181] Wherein the superscript H indicates the conjugate transpose, and t1(n) and t2(n) are the first and last time-frequency signal time indices corresponding to frame n (or in some embodiments, subframe n). In this example, there are four time indices t at each frame n. As described, a covariance matrix is determined for each bin. In other embodiments, the covariance matrix can also be averaged (or summed) over multiple frequency bins at a resolution close to the resolution of human hearing, or at the resolution of the determined spatial metadata parameters, or at any suitable resolution.
[0182] The target covariance matrix can be determined based on the spatial metadata and the total signal energy. o (b,n) can be obtained as C x Furthermore, in some embodiments, the spatial metadata includes the directional parameters azimuth θ(k,n) and elevation and directly to the total ratio parameter r(k,n). Note that the band index k is the band in which bin b is located. Furthermore, assuming that the output is a binaural signal, the target covariance matrix is:
[0183]
[0184] in, is for bin b, azimuth angle θ(k,n) and elevation angle The head-related transfer function column vector is a column vector of length two with complex values corresponding to the HRTF magnitude and phase for the left and right ears. In high frequencies, the HRTF values can also be real numbers, since phase differences are not required at high frequencies for perceptual reasons. Obtaining HRTFs for a given direction and frequency is known, and any suitable method can be used to obtain these HRTFs. d (b) is the diffuse-field binaural covariance matrix, which can be determined in an offline stage, for example, by taking a spatially uniform collection of HRTFs, formulating the covariance matrices of these HRTFs independently, and averaging the results.
[0185] Then, input the covariance matrix C x (b,n) and the target covariance matrix C y (b,n) is output as the covariance matrix 906 to the processing matrix determiner 905 .
[0186] The above examples only consider direction and ratio. The process of generating the target covariance matrix is also described in more detail in WO2019086757A1, where, in addition to direction and ratio, the use of spatial coherence parameters is also described, and other output types besides binaural output are also covered.
[0187] In some embodiments, the spatial synthesizer 703 includes a regularization factor determiner 903. The regularization factor determiner 903 is configured to receive input parameters such as the codec format 704 and the bit rate 702, and determine a regularization factor R(n) 904. The regularization factor determiner 903 is configured to output the regularization factor R(n) 904 to the processing matrix determiner 905.
[0188] In some embodiments, the spatial synthesizer 703 includes a processing matrix determiner 905. The processing matrix determiner 905 is configured to receive the covariance matrix C x (b,n) and C y (b,n) 906 and regularization factor R(n) 904, and determine the processing matrix 908 M(b,n) and M r (b,n). Determination of this processing matrix based on the covariance matrix is based on the work of “Vilkamo, J., T. and Kuntz, A. (2013), Optimized covariance domain framework for time-frequency processing of spatial audio, Journal of the Audio Engineering Society, 61(6), 403-411". The method determines the method for processing a sample with a measured covariance matrix C x (b, n) so that the output audio signal (ie, the processed input audio signal) obtains the determined target covariance matrix C y (b,n).
[0189] This approach has been used in various situations, including generating binaural and surround speaker signals. In formulating the processing matrix, the method further uses a prototype matrix, which is a matrix that informs the optimization process what kind of signal is generally meant for each output (with the constraint that the output must obtain the target covariance matrix). When the transmitted audio signal is two channels (left and right) and the output is a binaural signal, the prototype matrix Q can be, for example, simply or When the processing is binaural with head tracking, then when the user is facing in a rearward direction, the transmitted audio signals may be processed (e.g., just after or before the forward filter bank 901) so that they replace each other when the user is facing in a rearward direction.
[0190] The embodiment of this article is based on C x (b,n),C y (b,n) and Q(n) determine the processing matrices M(b,n) and M r (b,n) is regularized based on R(n). The derivation of the equations used to determine the processing matrix is thoroughly explained in “Vilkamo, J., and Kuntz, A. (2013), Optimized covariance domain framework for time-frequency processing of spatial audio, Journal of the Audio Engineering Society, 61(6), 403-411”. Note that this implementation is an example implementation and some operations may be performed in other ways to achieve the same or similar results. In some embodiments, the derivation of the equations used to determine the processing matrix is thoroughly explained in “Vilkamo, J., The example implementation provided in the appendix of "[I]n, T. and Kuntz, A. (2013), Optimized covariance domain framework for time-frequency processing of spatial audio, Journal of the Audio Engineering Society, 61(6), 403-411" may form the basis for an implementation for determining the processing matrix, except that a regularization factor R(n) is used instead of the fixed number 0.2 in line 21 of the program code (of that example implementation).
[0191] The use of regularization in determining the mixing matrix is explained here for completeness. Note that the notation used for the following matrices is the same as in “Vilkamo, J., T. and Kuntz, A. (2013), Optimized covariance domain framework for time-frequency processing of spatial audio, Journal of the Audio Engineering Society, 61(6), 403-411". The notation is not exactly the same. Also, frequency and bin indices are omitted for brevity.
[0192] First, the input and target covariance matrices are decomposed into And similarly for C y Using the Singular Value Decomposition (SVD) operation,
[0193] [U x ,S x ,V x ]=SVD(C x )
[0194] [U y ,S y ,V y ]=SVD(Cy )
[0195] and then, Wherein, the square root is an entry-by-entry operation, and the symbols x, y indicate the input covariance matrix or the target covariance matrix. As in “Vilkamo, J., T. and Kuntz, A. (2013), Optimized covariance domain framework for time-frequency processing of spatial audio, Journal of the Audio Engineering Society, 61(6), 403-411", without regularization, the processing matrix that provides the covariance matrix of the output signal when applied to the input transmission signal is obtained by the following formula:
[0196]
[0197] where the matrix inverse is unregularized, and where p is a unitary matrix formulated so that the similarity between the output (i.e., the input signal processed with M) and the prototype signal (i.e., the input signal processed with Q, with potential gain normalization) is maximized, where the constraint is that the output signal must obtain the target covariance matrix, i.e., MC x M H =C y The formulation of the matrix P is described in detail in the above references and is not the key focus of the current embodiments. However, the focus of the embodiments herein is on the regularized matrix inverse If the matrix is not regularized, there may be infinite processing gain in the matrix M.
[0198] In some embodiments, regularization can be implemented in the decomposition domain Among them, the diagonal matrix is normalized so that the lower limit of its diagonal value is limited to S x,sq The value of R times the maximum value of , thus obtaining the diagonal matrix S x,sq,reg This means that when R approaches zero, then S x,sq,reg Close to S x_sq , and thus the regularization approaches the state of no regularization. On the contrary, when R is close to 1, then S x,sq,reg Close to all diagonal values and S x,sq The matrix with the same maximum diagonal value is . This is the maximum regularization. The regularized inverse is formulated as:
[0199]
[0200] Furthermore, the regularized inverse is applied when formulating the mixing matrix:
[0201]
[0202] In this case, it depends on whether the regularization is active (i.e., on the matrix S x,sq The influence of the lower limit regularization), conditional MC x M H =C y may no longer hold. There may be a non-zero missing covariance matrix:
[0203] C r =C y -MC x M H
[0204] It can be called the residual covariance matrix. Therefore, the same method proposed in the previously cited reference can be used to generate another processing matrix M r , which is used to process the decorrelated version of the transmitted audio signal to obtain the characteristics of the lost part, in other words, to obtain the covariance matrix C r This can be done by adding C r Set as the target covariance matrix, remove C x off-diagonal elements of (since the sound is decorrelated) and use this as the input covariance matrix, and otherwise perform the same as described to obtain M r The same operation is performed. At this stage, the function of the regularization value is not very important because the input is irrelevant. So, for example, a fixed value of 0.2 can be used.
[0205] The processing matrix determiner 905 may then be configured to output the processing matrices M(b,n) and M(n) that have been regularized based on the regularization factor R(n) 904. r (b,n)908.
[0206] In some embodiments, the spatial synthesizer 703 includes a decorrelator and mixer 907. The decorrelator and mixer 907 is configured to receive the time-frequency transmission signal x(b,t) 906 and process the matrices M(b,n) and M r (b, n) 908. The decorrelator and mixer 907 is first configured to process the time-frequency transmission signal 906 with a decorrelator to generate a decorrelated signal x D (b,t).
[0207] The decorrelator and mixer 907 is then configured to apply the following mixing process to generate a time-frequency spatial audio signal y(b,t) 910 , which in turn may be output by the decorrelator and mixer 907 .
[0208] y(b,t)=M(b,n)x(b,t)+M r (b,n)x D (b,t)
[0209] In the above process, although not explicitly stated in the equation, the processing matrix can be linearly interpolated between frames n so that at each time index of the time-frequency signal, the matrix takes a step from M(b,n-1) toward M(b,n). The interpolation rate can be adjusted (fast interpolation) or not (normal interpolation) when the start is detected.
[0210] In some embodiments, the spatial synthesizer 703 comprises an inverse filterbank 911. The inverse filterbank 911 applies a corresponding inverse transform to that used by the forward filterbank 901 to convert the time-frequency spatial audio signal 910 into the spatial audio output 404, which is the output of the system.
[0211] about Figure 10 , showing a schematic diagram according to some embodiments Figure 9 An example flow chart of the operation of the spatial synthesizer shown in .
[0212] Thus, a first operation may include obtaining a transmission audio signal, spatial metadata and input parameters such as codec format and bit rate as indicated by 1001 .
[0213] Then, as indicated by 1003 , the transmission audio signal is time-frequency transformed to generate a time-frequency transmission audio signal.
[0214] Furthermore, the time-frequency transmitted audio signal can be used to generate an input covariance matrix. After the input covariance matrix has been determined, this can be used to determine the total energy. Then, using the total energy and spatial metadata, a target covariance matrix can be generated. Generating the input and target covariance matrices is shown by 1007.
[0215] Furthermore, as indicated by 1005, based on input parameters such as codec format and bit rate, a regularization factor is determined.
[0216] As indicated by 1009 , after the covariance matrix and regularization factor have been determined, a processing matrix is determined based on the input covariance matrix, the target covariance matrix, and the regularization factor.
[0217] Then, as indicated by 1011 , based on the processing matrix, the time-frequency transmission signals are decorrelated and mixed to generate a time-frequency spatial audio signal.
[0218] Furthermore, as indicated by 1013 , the time-frequency spatial audio signal is inversely time-frequency transformed to generate a spatial audio signal (eg, using an inverse filter bank).
[0219] Then, as indicated by 1015 , the spatial audio signal may be output as a spatial audio output.
[0220] about Figure 11 , showing that Figure 9 . In this example, the regularization factor determiner 903 is presented in the context of the example spatial audio codec system described above. However, a similar process can be applied in any context where a regularization factor or similar mixing or amplification limiting factor is used to control the automatic mixing of input channels to output channels and a constant value does not produce optimal quality.
[0221] Therefore, the initial operation is to obtain the input parameters that affect the regularization factor. For example, as shown by 1101, this can be to obtain the codec format, which can describe the spatial audio format of the codec passed to the renderer. For example, this can be the MASA format, the premixed multi-channel format, the panoramic surround sound format, etc.
[0222] Additionally, in some embodiments, as shown by 1105, this may also include obtaining a bitrate input parameter describing the bitrate used to encode the bitstream.
[0223] After obtaining these parameters, the process of defining or determining the regularization factor begins.
[0224] In some embodiments, the regularization factor itself may be defined as a decimal value in the range [0, 1], where a value of 1 limits the rendering gain the most and a value of 0 does not limit the rendering gain at all.
[0225] In some embodiments, as indicated by 1103 , the codec format parameters are used to select an initial set of available values for the regularization factors to be used in subsequent operation steps.
[0226] For example, this can be implemented in the form of multiple tables, where each table corresponds to a specific codec format. For example, a table for the MASA format might be [1.0, 0.8, 0.5, 0.2], and a similar table for the Atmos format might be [1.0, 0.7, 0.4, 0.2]. For a premixed multi-channel format, a similar example table might be [0.8, 0.5, 0.3, 0.2], since this content is typically professionally produced and the compression of the audio is typically less prone to compression artifacts. It should be noted that the use of tables is only one example implementation option, and other equivalent or even more suitable ways of implementing the options can be created.
[0227] Furthermore, as indicated by 1107, a set of values is taken and, based on the obtained codec bit rate, an initial value for the regularization factor is selected from the initial set of values. For example, using the MASA format table described above, a specific bit rate is associated with each table value. In this particular example, since there are four values, an example rule set may be:
[0228] - If bitrate ≤ 64kbps, use value 1.0
[0229] - If bitrate > 64kbps and bitrate ≤ 96kbps, use value 0.8
[0230] - If bitrate > 96kbps and bitrate ≤ 160kbps, use value 0.5
[0231] - If bitrate > 160kbps, use value 0.2
[0232] In some embodiments, an alternative way of implementing this determination would be to have a defined value in a set of values for each possible bit rate.
[0233] Any suitable method that allows obtaining an initial regularization factor for a given bit rate and codec format combination can be used. The general reasoning with the example values presented herein is that as the bit rate increases, the overall audio signal quality can be expected to improve. This in turn supports the use of smaller regularization factors (in other words, the rendering gain can be greater) because artifacts are less likely to appear.
[0234] A further operation is then to output a regularization factor R(n) 904 from the regularization factor determiner 903 and pass this factor to the processing matrix determiner 905 which applies the regularization factor R(n) 904 as part of the mixing matrix solution.
[0235] It should be noted that all of the example values described above are so-called tunable values for the system. In practice, there are many interactions between different parts of the codec and renderer, and a set of objective measurements and / or subjective listening tests are typically used to arrive at the values ultimately used. The example values given are suitable, but they may not produce optimal quality in all cases, and it is highly expected that a variety of different value sets can be used within the context of these embodiments.
[0236] In the above example, a single regularization factor is defined. When considering the matrix solution described in the example embodiment, regularization is achieved by controlling the entries of the diagonal matrix based on the regularization factor. In some other embodiments, for example, a finer granularity tuning of the regularization can be applied so that different regularization factors or different regularization schemes are used for different entries of the diagonal matrix. In other words, the regularization information may not be a single value, but a more refined set of information.
[0237] The above description shows the invention in the context of binaural rendering. However, in “Vilkamo, J., T. and Kuntz, A. (2013), Optimized covariance domain framework for time-frequency processing of spatial audio, Journal of the Audio Engineering Society, 61(6), 403-411". The method proposed in this paper provides a general mixing framework and can be applied to any input-output mixing case. This means that the target output can be loudspeakers, panoramic surround sound, or any other audio channel in addition to the binaural target output. Likewise, the example embodiments describing the adaptation of the regularization factor can be appropriately adapted to any other input-output mixing case. For example, a mix from two transport channels and metadata to a 7.1.4 output can use the method described in the above reference, and depending on the bitrate, different regularization factors should be used to obtain the best quality. Additionally, the output format can also be used to determine the choice of the regularization factor. For example, some artifacts may be more noticeable when using loudspeaker rendering than when using binaural rendering. This in turn suggests different regularization factor values for them.
[0238] In some alternative embodiments, after an initial value has been selected based on input parameters such as codec format and bit rate, additional adjustment steps can be performed on the regularization factor. These adjustments can change the regularization factor in absolute (e.g., +0.1 or -0.1) or relative (e.g., multiplied by 0.9) steps. If such adjustments are performed, an additional step should be introduced before outputting the final regularization factor. This step limits the regularization factor to a specific range (e.g., 0.2-1.0) to avoid unexpected values.
[0239] The source information used for these adjustment steps can have many options. For example, the source information can be one or more of the following:
[0240] Descriptive metadata in MASA format or any similar data. For example, the source format can be distinguished between channel-based audio and microphone capture, which can use different tables as described above. In addition, the transmission channel configuration can be used to perform adjustments, because, for example, cardioid pattern transmission may be more sensitive to artifacts than omnidirectional type transmission.
[0241] If an objective measure for the signal quality factor is available on a per-frame basis, this can be used to reduce or increase the regularization factor if the quality is known to be high or low, respectively.
[0242] In some embodiments, the order of selecting regularization factors in the proposed method can be easily reordered, since the resulting regularization factor is simply a combination of multiple factors. Other practical implementations are also valid.
[0243] In some embodiments, rather than directly signaling the format, the codec format may be signaled as a combination of bits indicating the spatial metadata used and / or the kind of transmitted audio signal.
[0244] It should be noted that while the above description shows that regularization controls the relationship between the original signal and the addition of a decorrelated version of the original signal to achieve a target covariance matrix, the presence of a decorrelated signal is not mandatory at all. The proposed method can also be used in the case of a solution using only the original signal. In this case, the target covariance matrix may not always be achieved.
[0245] In the above embodiments, covariance matrix based rendering as proposed in the above references is used as an example. However, it should be noted that the proposed method can also be used with other kinds of renderers, as long as there is some regularization for the gains. For example, in an alternative embodiment, "direct" and "ambient" sounds can be rendered separately. In this case, a prototype signal can be created separately for each of them (for example, using decorrelation for the ambient part), and target energies can also be created separately for the direct and ambient parts. Then, by comparing the energy of the prototype signal with the target energy, the gains for the rendering of the direct and ambient parts are calculated. These gains usually need regularization to avoid excessive noise gain, as proposed above for the main embodiment. These gains can be limited using the present invention. It should be noted that in this embodiment, the regularization is not a balance between mixing and decorrelation, but a balance between having artifacts from excessive noise gain and the attenuation of some signals.
[0246] about Figure 12, shows an example of processing according to some embodiments. This processing is performed using a system such as that described herein, where the original 5.1 sound is transmitted on two transport audio channels, and at the decoder, the audio is rendered as binaural sound using the transmitted spatial metadata. The spatial metadata and audio signal are encoded at two different bit rates, 48 kbps and 256 kbps, represented by the two columns 48 kbps 1200 and 256 kbps 1202.
[0247] The top row 1201 shows the binaurally processed sounds, wherein the middle row 1203 and the bottom row 1205 show only the decorrelated path sounds, in other words, wherein the M(b,n) and M(b,n) are formulated as r (b,n), the first row is set to zero before applying them to the signal. The middle row 1203 shows a prior art solution where the regularization factor is not adapted based on the bitrate, but is set to a default value of 1.0, which is safer in that it minimally amplifies codec artifacts when rendering. Instead, it uses decorrelation to a greater extent.
[0248] The bottom row 1205 illustrates processing according to some embodiments, where regularization is adapted based on the bitrate. When the bitrate is higher, the regularization factor is reduced to 0.2, which causes the system to make greater use of the hybrid approach to generate the desired incoherence and, therefore, use less decorrelation. In this case, due to the high bitrate, codec artifacts are at a lower level, and their amplification is inaudible. On the other hand, since less decorrelation is applied, better quality is provided.
[0249] In general, various embodiments of the present invention may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software, which may be executed by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Although various aspects of the present invention may be illustrated and described as block diagrams, flow charts, or using some other graphical representation, it is well understood that, as non-limiting examples, the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0250] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device, for example in a processor entity, or by hardware, or by a combination of software and hardware. Furthermore, it should be noted that any block of the logic flow, such as in the accompanying drawings, may represent program steps, or interconnected logic circuits, blocks, and functions, or a combination of program steps and logic circuits, blocks, and functions. The software may be stored on physical media such as memory chips, or storage blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants, CDs.
[0251] The memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technical environment and may include, by way of non-limiting example, one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processor architecture.
[0252] Embodiments of the present invention may be implemented in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0253] Programs such as those offered by Synopsys, Inc. of Mountain View, Calif., and Cadence Design, of San Jose, Calif., automatically route conductors and position components on semiconductor chips using established design rules and a library of pre-stored design modules. Once the design of a semiconductor circuit has been completed, the resulting design, in a standardized electronic format (e.g., Opus, GDSII, etc.), can be transferred to a semiconductor fabrication facility or "fab" for manufacturing.
[0254] As used in this application, the term "circuitry" may refer to one or more or all of the following:
[0255] (a) pure hardware circuit implementation (such as implementation in pure analog and / or digital circuits), and
[0256] (b) a combination of hardware circuitry and software such as (as applicable):
[0257] (i) a combination of analog and / or digital hardware circuitry and software / firmware, and
[0258] (ii) any portion of a hardware processor (including a digital signal processor) with software, software, and memory that work together to enable a device such as a mobile phone or server to perform various functions, and
[0259] A hardware circuit and / or processor, such as a microprocessor or portion of a microprocessor, that requires software (eg, firmware) for operation, but where software is not required for operation, the software may not be present.
[0260] This definition of circuitry applies to all uses of the term in this application, including in any claims. As a further example, as used in this application, the term "circuitry" also covers an implementation of a pure hardware circuit or processor (or multiple processors) or a portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term "circuitry" also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in a server, cellular network device or other computing or networking device. The term "non-transitory" as used herein is a restriction on the medium itself (i.e., tangible, non-signal), as opposed to a restriction on the persistence of data storage (e.g., RAM versus ROM).
[0261] As used herein, “at least one of: ” and “at least one of ” and similar expressions (wherein a list of two or more elements is connected by “and” or “or”) refer to at least any one element, or at least any two or more elements, or at least all elements.
[0262] The foregoing description has provided by way of exemplary and non-limiting examples a complete and informative description of the exemplary embodiments of the present invention. However, various modifications and variations may become apparent to those skilled in the relevant arts in view of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. Nevertheless, all such and similar modifications of the teachings of this invention will still fall within the scope of the invention as defined in the appended claims.
Claims
1. A method for generating an audio signal, the method comprising: obtaining an input audio signal comprising at least two audio channels; obtaining at least one input characteristic parameter associated with the input audio signal; determining at least one control parameter based at least on the at least one input characteristic parameter; determining a processing parameter based at least on the at least two audio channels, wherein the determination of the processing parameter is controlled at least in part based on the at least one control parameter; as well as The audio signal is generated based at least on the at least two audio channels and the processing parameters.
2. The method according to claim 1, wherein The input audio signal is a spatial audio signal further comprising at least one spatial parameter associated with the at least two audio channels.
3. The method according to claim 2, wherein: Determining the processing parameters based at least on the at least two audio channels further comprises determining the processing parameters further based at least on the at least one spatial parameter.
4. The method according to any one of claims 1 to 3, wherein The at least one input characteristic parameter includes at least one of the following: a bit rate associated with the at least one input audio signal; a codec format indicating a source of the at least one input audio signal; and configuration of the at least one input audio signal.
5. The method according to claim 4, wherein The configuration of the at least one input audio signal comprises at least one of the following: a source format indicating the original format or input format from which said input audio signal was created; Transmission channel description; Number of channels; Channel distance; as well as Channel angle.
6. The method according to any one of claims 4 or 5, wherein The codec format indicating the source of the at least one input audio signal includes at least one of the following: a metadata-assisted spatial audio stream source codec format value indicating that the at least one input audio signal represents a metadata-assisted spatial audio signal; a multi-channel audio stream source codec format value indicating that the at least one input audio signal represents a multi-channel audio signal; an audio object codec format value indicating that the input audio signal represents an audio object; as well as A panoramic surround sound format value indicating that the input audio signal represents a panoramic surround sound audio signal.
7. A method according to any one of claims 4 to 6 when appended to claim 2, wherein The at least one spatial parameter includes at least one of the following: information describing an organization of sound in space relative to the at least one input audio signal, the information comprising: a direction parameter configured to indicate from where the sound arrives; and a ratio parameter configured to indicate a portion of the sound arriving from that direction; Information describing characteristics of an original multi-channel or multi-object sound scene, the information comprising at least one of: channel level; object level; inter-channel correlation; inter-object correlation; and object direction; Processing coefficients related to obtaining a spatial audio format signal based at least on the at least one input audio signal.
8. The method according to any one of claims 4 to 7, wherein The codec format indicating the source of the at least one input audio signal includes an indication that the source is unknown or undefined.
9. The method according to any one of claims 1 to 8, wherein Determining the at least one control parameter based at least on the at least one input characteristic parameter includes: determining a first set of control parameter values based on a first input characteristic parameter of the at least one input characteristic parameter; and A control parameter value is selected from the first set of control parameter values based on a second input characteristic parameter of the at least one input characteristic parameter.
10. The method of claim 9 as appended to claim 4, wherein: The first input characteristic parameter of the at least one input characteristic parameter is the codec format, and the second input characteristic parameter of the at least one input characteristic parameter is the bit rate.
11. The method according to any one of claims 1 to 10, wherein Determining a processing parameter based at least on the at least two audio channels, wherein determining the processing parameter is controlled at least in part based on the at least one control parameter comprises: A processing matrix is generated based at least on the at least two audio channels, wherein the generation of the processing matrix is controlled at least in part based on the at least one control parameter.
12. The method according to claim 11, wherein Generating the processing matrix is controlled at least in part based on the control parameter, including regularizing the generation of the processing matrix based at least on the control parameter.
13. The method according to claim 12, wherein: Based at least on the at least two audio channels and the processing parameters, generating the audio signal includes processing the at least two audio channels using a regularized processing matrix to generate the audio signal.
14. The method according to any one of claims 12 or 13, wherein Generating the processing matrix to be controlled based at least in part on the at least one control parameter includes generating entries of a diagonal matrix based at least on at least two control parameter values.
15. The method according to any one of claims 1 to 14, wherein Obtaining at least one input characteristic parameter associated with the input audio signal comprises inferring the at least one input characteristic parameter from the input audio signal.
16. The method according to any one of claims 1 to 14, wherein Obtaining at least one input characteristic parameter associated with the input audio signal includes receiving the at least one input characteristic parameter as part of the input audio signal or as configuration information.
17. The method according to claim 16, wherein Receiving the at least one input characteristic parameter as part of the input audio signal further comprises: receiving encoded input characteristic parameters as part of the input audio signal; and The encoded input characteristic parameter is decoded to obtain the input characteristic parameter.
18. The method according to any one of claims 1 to 17, wherein Generating the audio signal based at least on the at least two audio channels and the processing parameters includes generating at least one of: Binaural audio signals; and Multi-channel audio signal.
19. An apparatus comprising means for performing the method according to any one of claims 1 to 18.
20. A computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform the method according to any one of claims 1 to 18.
21. An apparatus for generating an audio signal, the apparatus comprising means configured to: obtaining an input audio signal comprising at least two audio channels; obtaining at least one input characteristic parameter associated with the input audio signal; determining at least one control parameter based at least on the at least one input characteristic parameter; determining a processing parameter based at least on the at least two audio channels, wherein the determination of the processing parameter is controlled at least in part based on the at least one control parameter; as well as The audio signal is generated based at least on the at least two audio channels and the processing parameters.
22. An apparatus for generating an audio signal, the apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: obtaining an input audio signal comprising at least two audio channels; obtaining at least one input characteristic parameter associated with the input audio signal; determining at least one control parameter based at least on the at least one input characteristic parameter; determining a processing parameter based at least on the at least two audio channels, wherein the determination of the processing parameter is controlled at least in part based on the at least one control parameter; as well as The audio signal is generated based at least on the at least two audio channels and the processing parameters.
Citation Information
Patent Citations
Spatial audio parameter encoding and associated decoding
EP4091166A1
Analysis of spatial metadata from multi-microphones having asymmetric geometry in devices
US10873814B2
Determination of spatial audio parameter encoding and associated decoding
US20210295855A1
Selection of quantisation schemes for spatial audio parameter encoding
US20220036906A1
Determination of spatial audio parameter encoding and associated decoding
US20220343928A1