Parametric Spatial Audio Coding
By determining quantization bit allocations based on object priorities and using a spherical grid quantization configuration, the method optimizes encoding efficiency for immersive audio codecs, ensuring high-quality spatial audio reproduction across varying bitrates.
Patent Information
- Application Number
- JP2025531258
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-29
- Filing Date
- 2023-11-07
- Publication Date
- 2025-12-16
AI Technical Summary
Existing immersive audio codecs face challenges in efficiently encoding spatial audio data, particularly in managing bit allocation and quantization for directional metadata values across multiple audio objects, leading to suboptimal performance in varying bitrate conditions.
A method and apparatus for determining quantization bit allocations based on object priorities, using a spherical grid quantization configuration and varying bit allocation based on the importance of audio objects, with higher priority objects receiving more bits, to optimize encoding efficiency.
This approach enhances encoding efficiency by prioritizing critical audio objects, improving the quality of spatial audio reproduction across different bitrate conditions, ensuring high-quality spatial audio rendering even in low-bitrate scenarios.
Smart Images

Figure 2025540764000001_ABST
Abstract
Description
[Technical Field]
[0001] This application relates to an apparatus and method for spatial audio representation and encoding, but is not limited to audio representation for audio encoders. [Background technology]
[0002] Parametric spatial audio processing is a field of audio signal processing in which spatial aspects of sound are described using a set of parameters. For example, in parametric spatial audio capture from a microphone array, it is a typical and effective choice to estimate a set of parameters from the microphone array signal, such as the direction of sound within a frequency band and the ratio between the directional and non-directional portions of the captured sound within the frequency band. These parameters are known to well describe the perceived spatial characteristics of the sound captured at the microphone array location. These parameters can be used to synthesize spatial audio binaurally for headphones, for loudspeakers, or in other formats such as Ambisonics.
[0003] Therefore, direction within a frequency band and the direct-to-total energy ratio are particularly effective parameterizations for capturing spatial audio.
[0004] A parameter set consisting of frequency band direction parameters and frequency band energy ratio parameters (indicating sound directionality) can also be used as spatial metadata for an audio codec (which may also include surround coherence, spread coherence, number of directions, distance, etc.). For example, these parameters can be estimated from an audio signal captured by a microphone array, and a stereo or mono signal can be generated from the microphone array signal and conveyed along with the spatial metadata. The stereo signal can be encoded, for example, by an AAC encoder, and the mono signal can be encoded by an EVS encoder. A decoder can decode the audio signal into a PCM signal and process the sound in the frequency bands (using the spatial metadata) to obtain a spatial output, for example, a binaural output.
[0005] Immersive audio codecs are being implemented to support a variety of operating points, from low bitrate operation to transparency. An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, which is designed to be suitable for use over communication networks such as 3GPP's 4G / 5G networks, including for use in immersive services such as immersive voice and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general audio. Furthermore, it is expected to support channel-based and scene-based audio inputs, including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency to enable conversational services, as well as support high error resilience under various transmission conditions.
[0006] The aforementioned immersive audio codecs are particularly suitable for encoding spatial audio captured from microphone arrays (e.g., mobile phones, VR cameras, standalone microphone arrays), although such encoders can have other input types, e.g., loudspeaker signals, audio object signals, ambisonic signals. Summary of the Invention
[0007] According to a first aspect, there is provided an apparatus comprising: means for obtaining, for a plurality of audio objects in an audio environment, respective audio signals and directional metadata values; means for obtaining, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, at least one proportion parameter for identifying a distribution of the audio objects within the object portion of the overall audio environment, the frame comprising at least one time element and at least one frequency element, the at least one proportion parameter being configured to identify a distribution of the audio objects within the object portion of the overall audio environment; means for obtaining a further proportion parameter related to the distribution of the object portion of the overall audio environment; means for determining object priorities based on the proportion parameters of the plurality of audio objects and the further proportion parameter for the at least one time element and at least one frequency element of the frame; means for determining, based on the determined object priorities, a quantization bit allocation for directional metadata values associated with the plurality of audio objects; and means for quantizing the directional metadata values associated with the plurality of audio objects based on the quantization bit allocation.
[0008] The means may further be means for obtaining bit allocations for the plurality of objects, and the means for determining quantization bit allocations for directional metadata values associated with the plurality of audio objects based on the determined object priorities is further means for determining quantization bit allocations based on the bit allocations.
[0009] The means for determining quantization bit allocations for directional metadata values associated with the plurality of audio objects based on the determined object priorities may be means for determining a spherical grid quantization configuration based on the quantization bit allocations.
[0010] The means for quantizing directional metadata values associated with the plurality of audio objects based on a quantization bit allocation may be means for generating, for the directional metadata value, an index value representing a nearest node on the determined spherical grid quantization configuration.
[0011] The means for determining quantization bit allocations for directional metadata values associated with the plurality of audio objects based on the determined object priorities may be means for determining or allocating fewer bits to objects with lower priority.
[0012] The means for determining quantization bit allocations for directional metadata values associated with the plurality of audio objects based on the determined object priorities may be means for determining or allocating a first number of bits to an object with the highest priority, a second number of bits to a further object with the lowest priority, and a further number of bits between the first and second number to objects between the highest and lowest priority.
[0013] An additional number of bits between the first and second number for objects between the highest and lowest priority may be linearly scaled within the range between the first and second number.
[0014] An additional number of bits between the first and second numbers for objects between the highest and lowest priority may be scaled non-linearly within the range between the first and second numbers.
[0015] The means for determining object priority based on the ratio parameter and the further ratio parameter of a plurality of audio objects for at least one time element and at least one frequency element of a frame may be means for determining object priority based on an average of the ratio parameter and the further ratio parameter of a particular object over at least two time elements, or at least two frequency elements, or at least two time and frequency elements of a frame.
[0016] The means for determining object priority based on ratio parameters and further ratio parameters of a plurality of audio objects for at least one time element and at least one frequency element of a frame may be means for determining object priority based on a selection from ratio parameters of a particular audio object across at least two time elements, or at least two frequency elements, or at least two time and frequency elements.
[0017] The selection may be the maximum value of the ratio parameter of a particular audio object over at least two time components, or at least two frequency components, or at least two time and frequency components.
[0018] According to a second aspect, there is provided an apparatus comprising: means for obtaining, for a plurality of audio objects in an audio environment, respective encoded audio signals and encoded directional metadata values; means for obtaining, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, at least one encoded proportion parameter; means for obtaining, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, the frame comprising at least one time element and at least one frequency element, the at least one proportion parameter being configured to identify a distribution of the audio objects within the object portion of the overall audio environment; means for obtaining a further encoded proportion parameter related to the distribution of the object portion of the overall audio environment; means for determining, for the at least one time element and at least one frequency element of the frame, object priorities based on the encoded proportion parameters and the further proportion parameter of the plurality of audio objects; means for determining, based on the determined object priorities, quantization bit allocations of the encoded directional metadata values associated with the plurality of audio objects; and means for decoding the encoded directional metadata values associated with the plurality of audio objects based on the quantization bit allocations.
[0019] The means may further be means for obtaining bit allocations for the plurality of objects, and the means for determining quantization bit allocations for encoded directional metadata values associated with the plurality of audio objects based on the determined object priorities may further be means for determining quantization bit allocations based on the bit allocations.
[0020] The means for determining quantization bit allocations for encoded directional metadata values associated with the plurality of audio objects based on the determined object priorities may be means for determining a spherical grid quantization configuration based on the quantization bit allocations.
[0021] The means for decoding the encoded directional metadata values associated with the plurality of audio objects based on the quantization bit allocation may be means for determining directional values associated with the encoded directional metadata values representing nodes on the determined quantization structure of the spherical grid.
[0022] The means for determining quantization bit allocations for directional metadata values associated with the plurality of audio objects based on the determined object priorities may be means for determining or allocating fewer bits to objects with lower priority.
[0023] The means for determining quantization bit allocations for directional metadata values associated with the plurality of audio objects based on the determined object priorities may be means for determining or allocating a first number of bits to an object with the highest priority, a second number of bits to a further object with the lowest priority, and a further number of bits between the first and second number to objects between the highest and lowest priority.
[0024] An additional number of bits between the first and second number for objects between the highest and lowest priority may be linearly scaled within the range between the first and second number.
[0025] An additional number of bits between the first and second numbers for objects between the highest and lowest priority may be scaled non-linearly within the range between the first and second numbers.
[0026] The means for determining object priority based on the ratio parameters of a plurality of audio objects for a particular time element and a particular frequency element of a frame and the further ratio parameter may be means for determining object priority based on an average of the coded ratio parameters of the particular object and the further coded ratio parameter over at least two time elements, or at least two frequency elements, or at least two time and frequency elements of the frame.
[0027] The means for determining object priority based on ratio parameters of a plurality of audio objects for a particular time element and a particular frequency element of a frame and on further ratio parameters may be means for determining object priority based on a selection from encoded ratio parameters of a particular audio object across at least two time elements, or at least two frequency elements, or at least two time and frequency elements.
[0028] The selection may be the maximum value of the encoded ratio parameter of a particular audio object over at least two time elements, or over at least two frequency elements, or over at least two time and frequency elements.
[0029] According to a third aspect, there is provided a method comprising: obtaining, for a plurality of audio objects in an audio environment, respective audio signals and directional metadata values; obtaining, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, at least one proportion parameter for identifying a distribution of the audio objects within the object portion of the overall audio environment, the frame comprising at least one time element and at least one frequency element, the at least one proportion parameter being configured to identify a distribution of the audio objects within the object portion of the overall audio environment; obtaining a further proportion parameter related to the distribution of the object portion of the overall audio environment; determining, for the at least one time element and at least one frequency element of the frame, object priorities based on the proportion parameters of the plurality of audio objects and the further proportion parameter; determining, based on the determined object priorities, quantization bit allocations for directional metadata values associated with the plurality of audio objects; and quantizing the directional metadata values associated with the plurality of audio objects based on the quantization bit allocations.
[0030] The method may include obtaining bit allocations for a plurality of objects, and determining quantization bit allocations for directional metadata values associated with the plurality of audio objects based on the determined object priorities may further include determining quantization bit allocations based on the bit allocations.
[0031] Determining quantization bit allocations for directional metadata values associated with the plurality of audio objects based on the determined object priorities may include determining a spherical grid quantization configuration based on the quantization bit allocations.
[0032] Quantizing directional metadata values associated with the plurality of audio objects based on the quantization bit allocation may include generating, for the directional metadata value, an index value that represents a nearest node on the determined spherical grid quantization configuration.
[0033] Determining quantization bit allocations for directional metadata values associated with the plurality of audio objects based on the determined object priorities may include determining or allocating fewer bits to lower priority objects.
[0034] Determining quantization bit allocations for directional metadata values associated with the plurality of audio objects based on the determined object priorities may include determining or allocating a first number of bits to the highest priority object, a second number of bits to a further object of the lowest priority, and a further number of bits between the first and second number to an object between the highest and lowest priority.
[0035] An additional number of bits between the first and second number for objects between the highest and lowest priority may be linearly scaled within the range between the first and second number.
[0036] An additional number of bits between the first and second numbers for objects between the highest and lowest priority may be scaled non-linearly within the range between the first and second numbers.
[0037] Determining object priority based on the ratio parameter and the further ratio parameter of a plurality of audio objects for at least one time element and at least one frequency element of a frame may include determining object priority based on an average of the ratio parameter and the further ratio parameter of a particular object over at least two time elements, or at least two frequency elements, or at least two time and frequency elements.
[0038] Determining object priority based on ratio parameters and further ratio parameters of a plurality of audio objects for at least one time element and at least one frequency element of a frame may include determining object priority based on a selection from ratio parameters of a particular audio object across at least two time elements, or at least two frequency elements, or at least two time and frequency elements.
[0039] The selection may be the maximum value of the ratio parameter of a particular audio object over at least two time components, or at least two frequency components, or at least two time and frequency components.
[0040] According to a fourth aspect, there is provided a method comprising: obtaining, for a plurality of audio objects in an audio environment, respective coded audio signals and coded directional metadata values; obtaining, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, at least one coded proportion parameter; obtaining, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, the frame comprising at least one time element and at least one frequency element, the at least one proportion parameter being configured to identify a distribution of the audio objects within the object portion of the overall audio environment; obtaining a further coded proportion parameter related to the distribution of the object portion of the overall audio environment; determining, for the at least one time element and at least one frequency element of the frame, object priorities based on the coded proportion parameters of the plurality of audio objects; determining, based on the determined object priorities, quantization bit allocations for the coded directional metadata values associated with the plurality of audio objects; and decoding the coded directional metadata values associated with the plurality of audio objects based on the quantization bit allocations.
[0041] The method may include obtaining a bit allocation for a plurality of objects, and determining a quantization bit allocation for encoded directional metadata values associated with the plurality of audio objects based on the determined object priorities may include determining a quantization bit allocation based on the bit allocation.
[0042] Determining quantization bit allocations for encoded directional metadata values associated with the plurality of audio objects based on the determined object priorities may include determining a spherical grid quantization configuration based on the quantization bit allocations.
[0043] Decoding the encoded directional metadata values associated with the plurality of audio objects based on the quantization bit allocation may include determining directional values associated with the encoded directional metadata values that represent nodes on the determined spherical grid quantization configuration.
[0044] Determining quantization bit allocations for directional metadata values associated with the plurality of audio objects based on the determined object priorities may include determining or allocating fewer bits to lower priority objects.
[0045] Determining quantization bit allocations for directional metadata values associated with the plurality of audio objects based on the determined object priorities may include determining or allocating a first number of bits to the highest priority object, a second number of bits to a further object of the lowest priority, and a further number of bits between the first and second number to an object between the highest and lowest priority.
[0046] An additional number of bits between the first and second number for objects between the highest and lowest priority may be linearly scaled within the range between the first and second number.
[0047] An additional number of bits between the first and second numbers for objects between the highest and lowest priority may be scaled non-linearly within the range between the first and second numbers.
[0048] Determining object priority based on ratio parameters of a plurality of audio objects for a particular time element and a particular frequency element of a frame and further ratio parameters may include determining object priority based on coded ratio parameters of a particular object and further coded ratio parameters across at least two time elements, or at least two frequency elements, or at least two time and frequency elements.
[0049] Determining the object priority based on the ratio parameters of the plurality of audio objects for a particular time element and a particular frequency element of the frame and the further ratio parameters may include determining the object priority based on a selection from the encoded ratio parameters of the particular audio object across at least two time elements, or at least two frequency elements, or at least two time and frequency elements.
[0050] The selection may be the maximum value of the encoded ratio parameter of a particular audio object over at least two time elements, or over at least two frequency elements, or over at least two time and frequency elements.
[0051] According to a fifth aspect, there is provided an apparatus comprising at least one processor and at least one memory that stores instructions that, when executed by the at least one processor, cause a system to at least perform the following: obtain, for a plurality of audio objects in an audio environment, respective audio signals and directional metadata values; obtain, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, at least one ratio parameter configured to identify a distribution of the audio objects within the object portion of the overall audio environment; obtain a further ratio parameter related to the distribution of the object portion of the overall audio environment; determine, for the at least one time element and the at least one frequency element of the frame, object priorities based on the ratio parameters of the plurality of audio objects and the further ratio parameter; determine, based on the determined object priorities, quantization bit allocations for directional metadata values associated with the plurality of audio objects; and quantize the directional metadata values associated with the plurality of audio objects based on the quantization bit allocations.
[0052] The apparatus may further be adapted to obtain bit allocations for the plurality of objects, and the apparatus adapted to determine quantization bit allocations for directional metadata values associated with the plurality of audio objects based on the determined object priorities may be adapted to determine the quantization bit allocations based on the bit allocations.
[0053] An apparatus adapted to determine quantization bit allocations for directional metadata values associated with a plurality of audio objects based on the determined object priorities may be adapted to determine a spherical grid quantization configuration based on the quantization bit allocations.
[0054] An apparatus adapted to quantize directional metadata values associated with a plurality of audio objects based on a quantization bit allocation may be adapted to generate, for the directional metadata value, an index value representing the nearest node on the determined spherical grid quantization configuration.
[0055] Based on the determined object priorities, an apparatus adapted to determine quantization bit allocations for directional metadata values associated with a plurality of audio objects may be adapted to determine or allocate fewer bits to objects with lower priority.
[0056] An apparatus adapted to determine quantization bit allocations for directional metadata values associated with a plurality of audio objects based on determined object priorities may be adapted to determine or allocate a first number of bits to the highest priority object, a second number of bits to a further object of the lowest priority, and a further number of bits between the first and second number to an object between the highest and lowest priority.
[0057] An additional number of bits between the first and second number for objects between the highest and lowest priority may be linearly scaled within the range between the first and second number.
[0058] An additional number of bits between the first and second numbers for objects between the highest and lowest priority may be scaled non-linearly within the range between the first and second numbers.
[0059] An apparatus configured to determine object priority based on a ratio parameter and a further ratio parameter of a plurality of audio objects for at least one time element and at least one frequency element of a frame may be configured to determine object priority based on an average of the ratio parameter and the further ratio parameter of a particular object over at least two time elements, or at least two frequency elements, or at least two time and frequency elements.
[0060] An apparatus adapted to determine object priority based on ratio parameters and further ratio parameters of a plurality of audio objects for at least one time element and at least one frequency element of a frame may be adapted to determine object priority based on a selection from ratio parameters of a particular audio object across at least two time elements, or at least two frequency elements, or at least two time and frequency elements.
[0061] The selection may be the maximum value of the ratio parameter of a particular audio object over at least two time components, or at least two frequency components, or at least two time and frequency components.
[0062] According to a sixth aspect, there is provided an apparatus comprising at least one processor and at least one memory that stores instructions that, when executed by the at least one processor, cause a system to at least perform the following: obtain, for a plurality of audio objects in an audio environment, respective encoded audio signals and encoded directional metadata values; obtain, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, at least one encoded proportion parameter; obtain, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, the frame comprising at least one time element and at least one frequency element, the at least one proportion parameter being configured to identify a distribution of the audio objects within an object portion of the overall audio environment; obtain a further encoded proportion parameter related to the distribution of the object portion of the overall audio environment; determine, for the at least one time element and at least one frequency element of the frame, object priorities based on the encoded proportion parameters and the further proportion parameter of the plurality of audio objects; determine, based on the determined object priorities, quantization bit allocations for the encoded directional metadata values associated with the plurality of audio objects; and decode the encoded directional metadata values associated with the plurality of audio objects based on the quantization bit allocations.
[0063] The apparatus may be adapted to obtain bit allocations for a plurality of objects and to determine quantization bit allocations for encoded directional metadata values associated with a plurality of audio objects based on the determined object priorities, and the apparatus may be further adapted to determine quantization bit allocations based on the bit allocations.
[0064] An apparatus adapted to determine quantization bit allocations for encoded directional metadata values associated with a plurality of audio objects based on the determined object priorities may be adapted to determine a spherical grid quantization configuration based on the quantization bit allocations.
[0065] An apparatus adapted to decode encoded directional metadata values associated with a plurality of audio objects based on a quantization bit allocation may be adapted to determine directional values associated with the encoded directional metadata values that represent nodes on the determined spherical grid quantization configuration.
[0066] Based on the determined object priorities, an apparatus adapted to determine quantization bit allocations for directional metadata values associated with a plurality of audio objects may be adapted to determine or allocate fewer bits to objects with lower priority.
[0067] An apparatus adapted to determine quantization bit allocations for directional metadata values associated with a plurality of audio objects based on determined object priorities may be adapted to determine or allocate a first number of bits to the highest priority object, a second number of bits to a further object of the lowest priority, and a further number of bits between the first and second number to an object between the highest and lowest priority.
[0068] An additional number of bits between the first and second number for objects between the highest and lowest priority may be linearly scaled within the range between the first and second number.
[0069] An additional number of bits between the first and second numbers for objects between the highest and lowest priority may be scaled non-linearly within the range between the first and second numbers.
[0070] An apparatus configured to determine object priorities based on ratio parameters of a plurality of audio objects for a particular time element and a particular frequency element of a frame and a further ratio parameter may be configured to determine object priorities based on an average of the coded ratio parameters of the particular object and the further coded ratio parameter over at least two time elements, or at least two frequency elements, or at least two time and frequency elements.
[0071] An apparatus adapted to determine object priorities based on ratio parameters of a plurality of audio objects for a particular time element and a particular frequency element of a frame and further ratio parameters may be adapted to determine object priorities based on a selection from encoded ratio parameters of a particular audio object across at least two time elements, or at least two frequency elements, or at least two time and frequency elements.
[0072] The selection may be the maximum value of the encoded ratio parameter of a particular audio object over at least two time elements, or over at least two frequency elements, or over at least two time and frequency elements.
[0073] According to a seventh aspect, there is provided an apparatus for encoding audio object parameters, the apparatus comprising: means for obtaining, for a plurality of audio objects in an audio environment, each audio signal and a directional metadata value; means for obtaining, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, at least one ratio parameter for identifying a distribution of the audio objects within the object portion of the overall audio environment; means for obtaining a further ratio parameter related to the distribution of the object portion of the overall audio environment; means for determining, for the at least one time element and the at least one frequency element of the frame, object priorities based on the ratio parameters of the plurality of audio objects and the further ratio parameter; means for determining, based on the determined object priorities, a quantization bit allocation for directional metadata values associated with the plurality of audio objects; and means for quantizing the directional metadata values associated with the plurality of audio objects based on the quantization bit allocation.
[0074] According to an eighth aspect, there is provided an apparatus for encoding audio object parameters, the apparatus comprising: means for obtaining, for a plurality of audio objects in an audio environment, respective encoded audio signals and encoded directional metadata values; means for obtaining, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, at least one encoded proportion parameter; means for obtaining, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, the frame comprising at least one time element and at least one frequency element, the at least one proportion parameter being configured to identify a distribution of the audio object within the object portion of the overall audio environment; means for obtaining a further encoded proportion parameter related to the distribution of the object portion of the overall audio environment; means for determining object priorities based on the encoded proportion parameters and the further proportion parameter of the plurality of audio objects, for the at least one time element and the at least one frequency element of the frame; means for determining, based on the determined object priorities, quantization bit allocations of the encoded directional metadata values associated with the plurality of audio objects; and means for decoding the encoded directional metadata values associated with the plurality of audio objects based on the quantization bit allocations.
[0075] According to a ninth aspect, there is provided an apparatus comprising: an acquisition circuit configured to acquire, for a plurality of audio objects in an audio environment, respective audio signals and directional metadata values; the acquisition circuit configured to acquire, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, the frame comprising at least one time element and at least one frequency element, the at least one ratio parameter configured to identify a distribution of the audio objects within an object portion of an overall audio environment; the acquisition circuit configured to acquire a further ratio parameter related to the distribution of the object portion of the overall audio environment; a decision circuit configured to determine, for the at least one time element and the at least one frequency element of the frame, an object priority based on the ratio parameters of the plurality of audio objects and the further ratio parameter; a decision circuit configured to determine, based on the determined object priority, a quantization bit allocation of directional metadata values associated with the plurality of audio objects; and a quantization circuit configured to quantize the directional metadata values associated with the plurality of audio objects based on the quantization bit allocation.
[0076] According to a tenth aspect, there is provided an apparatus including: an acquisition circuit configured to acquire a plurality of audio objects in an audio environment, respective encoded audio signals, and encoded directional metadata values; the acquisition circuit configured to acquire at least one encoded proportion parameter for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, the frame comprising at least one time element and at least one frequency element, the at least one proportion parameter configured to identify a distribution of the audio objects within an object portion of an overall audio environment; the acquisition circuit configured to acquire a further encoded proportion parameter related to the distribution of the object portion of the overall audio environment; a decision circuit configured to determine object priorities based on the encoded proportion parameters and the further proportion parameter of the plurality of audio objects for the at least one time element and at least one frequency element of the frame; a decision circuit configured to determine quantization bit allocations for encoded directional metadata values associated with the plurality of audio objects based on the determined object priorities; and a decoding circuit configured to decode the encoded directional metadata values associated with the plurality of audio objects based on the quantization bit allocations.
[0077] According to an eleventh aspect, there is provided a computer program comprising instructions [or a computer-readable medium comprising instructions] causing an apparatus to at least perform the following: obtain, for a plurality of audio objects in an audio environment, each audio signal and a directional metadata value; obtain, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, at least one ratio parameter configured to identify a distribution of the audio objects within the object portion of the overall audio environment; obtain a further ratio parameter related to the distribution of the object portion of the overall audio environment; determine, for the at least one time element and at least one frequency element of the frame, object priorities based on the ratio parameters of the plurality of audio objects and the further ratio parameter; determine, based on the determined object priorities, quantization bit allocations for directional metadata values associated with the plurality of audio objects; and quantize the directional metadata values associated with the plurality of audio objects based on the quantization bit allocations.
[0078] According to a twelfth aspect, there is provided a computer program comprising instructions [or a computer-readable medium comprising instructions] to cause at least an apparatus to perform the following: obtain, for a plurality of audio objects in an audio environment, respective encoded audio signals and encoded directional metadata values; obtain, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, at least one encoded proportion parameter; obtain, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, the frame comprising at least one time element and at least one frequency element, the at least one proportion parameter being configured to identify a distribution of the audio object within the object portion of the overall audio environment; obtain a further encoded proportion parameter related to the distribution of the object portion of the overall audio environment; determine, for the at least one time element and at least one frequency element of the frame, object priorities based on the encoded proportion parameters and the further proportion parameter of the plurality of audio objects; determine, based on the determined object priorities, quantization bit allocations for the encoded directional metadata values associated with the plurality of audio objects; and decode the encoded directional metadata values associated with the plurality of audio objects based on the quantization bit allocations.
[0079] According to a thirteenth aspect, there is provided a non-transitory computer-readable medium comprising program instructions that cause an apparatus to at least perform the following: obtain, for a plurality of audio objects in an audio environment, respective audio signals and directional metadata values; obtain, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, at least one ratio parameter configured to identify a distribution of the audio objects within the object portion of the overall audio environment; obtain a further ratio parameter related to the distribution of the object portion of the overall audio environment; determine, for the at least one time element and the at least one frequency element of the frame, object priorities based on the ratio parameters of the plurality of audio objects and the further ratio parameter; determine, based on the determined object priorities, quantization bit allocations for directional metadata values associated with the plurality of audio objects; and quantize the directional metadata values associated with the plurality of audio objects based on the quantization bit allocations.
[0080] According to a fourteenth aspect, there is provided a non-transitory computer-readable medium comprising program instructions that cause an apparatus to at least perform the following: obtain, for a plurality of audio objects in an audio environment, respective encoded audio signals and encoded directional metadata values; obtain, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, at least one encoded proportion parameter; obtain, for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, the frame comprising at least one time element and at least one frequency element, the at least one proportion parameter being configured to identify a distribution of the audio object within the object portion of the overall audio environment; obtain a further encoded proportion parameter related to the distribution of the object portion of the overall audio environment; determine, for the at least one time element and the at least one frequency element of the frame, object priorities based on the encoded proportion parameters and the further proportion parameter of the plurality of audio objects; determine, based on the determined object priorities, quantization bit allocations for the encoded directional metadata values associated with the plurality of audio objects; and decode the encoded directional metadata values associated with the plurality of audio objects based on the quantization bit allocations.
[0081] An apparatus comprising means for performing the actions of the above method.
[0082] An apparatus configured to perform the actions of the above method.
[0083] A computer program comprising program instructions for causing a computer to carry out the above method.
[0084] A computer program product stored on a medium can cause an apparatus to perform the methods described herein.
[0085] The electronic device may include the apparatus described herein.
[0086] The chipset may include the devices described herein.
[0087] SUMMARY OF THE INVENTION Embodiments of the present application aim to address problems associated with the state of the art.
[0088] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings, in which: [Brief explanation of the drawings]
[0089] [Figure 1] 1 illustrates a schematic diagram of a system of apparatus suitable for implementing some embodiments. [Figure 2] 2 illustrates schematically an exemplary encoding mode selector in the system of devices shown in FIG. 1, according to some embodiments. [Figure 3] 3 illustrates a flow diagram of the operation of the exemplary encoding mode selector shown in FIG. 2, according to some embodiments. [Figure 4] 5 illustrates a flow diagram of the operation of the exemplary first, lowest, or only MASA bitrate encoding mode shown in FIG. 4, according to some embodiments. [Figure 5] 5 illustrates a flow diagram of the operation of the exemplary second, low, or object information coding mode shown in FIG. 4, according to some embodiments. [Figure 6] 5 illustrates a flow diagram of the exemplary third, high, or single object coding mode of operation shown in FIG. 4, according to some embodiments. [Figure 7] 5 illustrates a flow diagram of the operation of the exemplary fourth, top, or independent object and multiple input coding mode shown in FIG. 4, according to some embodiments. [Figure 8] 2 illustrates schematically the exemplary audio object metadata encoder shown in FIG. 1 for a fourth encoding mode, according to some embodiments; [Figure 9]9 illustrates a flow diagram of the operation of the encoding mode selector of the exemplary audio object metadata encoder shown in FIG. 8, according to some embodiments. [Figure 10] 1 illustrates an exemplary device suitable for implementing the apparatus shown in the above figures. DETAILED DESCRIPTION OF THE INVENTION
[0090] In the following, suitable devices and possible mechanisms for encoding a parametric spatial audio signal including a transport audio signal and spatial metadata are described in more detail. As mentioned above, immersive audio codecs (such as 3GPP's IVAS) are planned that support multiple operating points, from low-bitrate operation to transparency. They are expected to support channel-based audio input and scene-based audio input, including spatial information about the sound field and sound sources. In the following, an exemplary codec is configured to receive multiple input formats. In particular, the codec is configured to acquire or receive multiple audio signals (e.g., received from a microphone array, or as a multi-channel audio format input, Ambisonic format input) and one or more audio object signals (which may also be referred to as Independent Stream with Metadata (ISM) formats). Furthermore, in some situations, the codec is configured to process multiple input formats at once. This combined (input) format mode can, for example, enable simultaneous encoding of two different audio input formats. An example of two different audio input formats currently considered is the combination of the MASA format and the audio object format. Metadata-Assisted Spatial Audio (MASA) is an example of a parametric spatial audio format and representation that is suitable as an input format for IVAS.
[0091] It can be seen as an audio representation consisting of "N channels + spatial metadata". It is a scene-based audio format particularly suited for spatial audio capture on practical devices such as smartphones. The idea is to describe the acoustic scene in terms of the direction of sound sources and, for example, energy ratios, which vary with time and frequency. Acoustic energy that is not defined (described) by direction is described as diffuse (coming from all directions).
[0092] As mentioned above, spatial metadata associated with an audio signal may include multiple parameters for each time-frequency tile (multiple directions and a direct-to-total energy ratio, spread coherence, distance, etc. associated with each direction (or direction value)). The spatial metadata may also include other parameters or may be associated with other parameters that are considered non-directional (such as surround coherence, diffuse-to-total energy ratio, residual-to-total energy ratio, etc.), but which, combined with the directional parameters, can be used to define the characteristics of the audio scene. For example, a reasonable design choice that can produce high-quality output is one in which the spatial metadata includes one or more directions for each time-frequency subframe (and a direct-to-total ratio, spread coherence, distance value, etc. associated with each direction).
[0093] A concept described in further detail herein is the definition of a multi-rate coding model that provides for encoding of the combined format at various bit rates. This coding model allows for parametric coding of audio object input, including variable-rate coding of object direction parameters based on the determined priorities of the objects and the available bit rates. Accordingly, embodiments described in further detail herein include apparatus and method embodiments for defining the quantization resolution of object metadata as a function of both the MASA-to-total energy ratio and the ISM ratio. Based on these parameters, a priority value for each object is calculated, which is used to vary the quantization resolution of the object metadata.
[0094] As mentioned above, the parametric spatial metadata representation can use multiple simultaneous spatial directions. For MASA, the suggested maximum number of simultaneous directions is two. For each simultaneous direction, there may be a direction index and associated parameters such as direct-to-total ratio, spread coherence, and distance. In some embodiments, other parameters are defined, such as diffuse-to-total energy ratio, surround coherence, and residual-to-total energy ratio.
[0095] In this regard, Fig. 1 illustrates an exemplary apparatus 100 and system for implementing embodiments of the present application. The system is shown with an "analysis" part, which is the part from receiving the multi-channel signal to encoding the metadata and downmix signal.
[0096] The input to the "analysis" portion of the system is a multi-channel audio signal 102. In the examples below, microphone channel signal inputs are described, but in other embodiments, any suitable input (or synthesized multi-channel) format may be implemented. For example, in some embodiments, the spatial analyzer and spatial analysis may be implemented external to the encoder. For example, in some embodiments, spatial (MASA) metadata associated with the audio signal may be provided to the encoder as a separate bitstream. In some embodiments, the spatial (MASA) metadata may be provided as a set of spatial (directional) index values.
[0097] 1 also shows a number of audio objects 104 as further inputs to the analysis portion. As mentioned above, these multiple audio objects (or audio object streams) 104 may represent various sound sources in physical space. Each audio object may be characterized by an audio (object) signal and accompanying metadata including directional data (in the form of azimuth and elevation values), which indicates the position or direction of the audio object in physical space for each audio frame.
[0098] The multi-channel signal 102 is passed to an analyzer and encoder 101 , in particular to a transport signal generator 105 and a metadata generator 103 .
[0099] In some embodiments, the metadata generator 103 is also configured to receive the multi-channel signal and analyze the signal to generate metadata 104 related to the multi-channel signal and, therefore, related to the transport signal 106. The analysis processor 103 may be configured to generate, for each time-frequency analysis interval, metadata that may include direction parameters, energy ratio parameters, and coherence parameters (and, in some embodiments, diffuseness parameters). The direction, energy ratio, and coherence parameters may, in some embodiments, be considered to be MASA spatial audio parameters (or MASA metadata). In other words, spatial audio parameters include parameters that aim to characterize the sound field created / captured by the multi-channel signal (or two or more audio signals in general).
[0100] In some embodiments, the generated parameters may differ for each frequency band. Thus, for example, in band X, all of the parameters are generated and transmitted, while in band Y, only one of the parameters is generated and transmitted, and in band Z, no parameters are generated or transmitted. A practical example of this may be that in some frequency bands, such as the highest band, some of the parameters are not needed for perceptual reasons. The transport signal 106 and metadata 104 may be passed to a combined encoder core 109.
[0101] In some embodiments, the transport signal generator 105 is configured to receive the multi-channel signal, generate an appropriate transport signal including the determined number of channels, and output the transport signal 106 (MASA transport audio signal). For example, the transport signal generator 105 may be configured to generate a two-audio-channel downmix of the multi-channel signal. The determined number of channels may be any suitable number of channels. In some embodiments, the transport signal generator is configured to select or combine input audio signals into the determined number of channels, for example by beamforming techniques, and output these as a transport signal.
[0102] In some embodiments, the transport signal generator 105 is optional and the multi-channel signal is passed to the coupled encoder core 109 in the same manner as the transport signal, in this example without processing.
[0103] The audio objects 104 may be passed to the audio object analyzer 107 for processing. In some embodiments, the audio object analyzer 107 analyzes the object audio input stream 104 to generate an appropriate audio object transport signal and audio object metadata. For example, the audio object analyzer may be configured to generate the audio object transport signal by downmixing the audio signals of the audio objects together into stereo channels using amplitude panning based on the direction of the associated audio objects. Furthermore, the audio object analyzer may also be configured to generate audio object metadata associated with the audio object input stream 104. The audio object metadata may include direction values applicable to all subbands. Thus, if four objects are present, there are four directions. In the example described herein, the direction values also apply across all subframes of a frame, but in some embodiments, the time resolution of the direction values may be different, and the direction values apply to one or more subframes of a frame. Furthermore, an energy ratio (or ISM ratio) may be determined for each object. The energy ratio (ISM ratio) defines the contribution of the object within the object's portion of the overall audio environment. In the following examples, the energy ratio (or ISM ratio) is for each time-frequency tile of each object.
[0104] In some embodiments, the audio object analyzer 107 may be located elsewhere and the audio objects 104 input to the analyzer and encoder 101 are audio object transport signals and audio object metadata.
[0105] The analyzer and encoder 101 may comprise a combined encoder core 109 configured to receive a transport audio (e.g., downmix) signal 106 and an audio object transport signal 128 in order to generate appropriate encodings of these audio signals.
[0106] The analyzer and encoder 101 may also comprise an audio object metadata encoder 111, which is similarly configured to receive the audio object metadata 108 and to output the input information in coded or compressed form as coded audio object metadata 112.
[0107] In some embodiments, the combined encoder core may be configured to implement a stream separation metadata determinator and encoder, which may be configured to determine the relative contributions of the multi-channel signal 102 (the MASA audio signal) and the audio objects 104 to the overall audio scene. This measure of proportionality generated by the stream separation metadata determinator and encoder may be used to determine the proportion of quantization and encoding "effort" to be spent on the input multi-channel signal 102 and the audio objects 104. In other words, the stream separation metadata determinator and encoder may generate a metric that quantifies the proportion of encoding effort spent on the multi-channel audio signal 102 compared to the encoding effort spent on the audio objects 104. This metric may be used to drive the encoding of the audio object metadata 108 and the audio object metadata 104. Furthermore, the metric determined by the separation metadata determinator and encoder may also be used as an influencing factor in the process of encoding the transport audio signal 106 and the audio object transport audio signal 128 performed by the combined encoder core 109. The output metrics from the stream separation metadata determiner and encoder can further be represented as coded stream separation metadata and combined into a coded metadata stream from the combined encoder core 109.
[0108] In some embodiments, the analyzer and encoder 101 comprises a bitstream generator 113 configured to obtain the encoded metadata 116, the encoded transport audio signal 138, and the encoded audio object metadata 112 and to generate a bitstream 118 for potential transmission or storage.
[0109] In some embodiments, the analyzer and encoder 101 comprises an encoder controller 115, which in some embodiments may control the encoding implemented by the audio object metadata encoder 111 and the coupled encoder core 109. In some embodiments, the encoder controller 115 is configured to determine a bit rate for the bit stream 118 and control the encoding based on the bit rate. In some embodiments, the encoder controller 115 is further configured to control at least one of the audio object analyzer 107, the transport signal generator 105, and the metadata generator in generating the parameters.
[0110] The analyzer and encoder 101 may, in some embodiments, be a computer or mobile device (executing appropriate software stored in memory and at least one processor), or alternatively, a specific device utilizing, for example, an FPGA or ASIC. The encoding may be performed using any appropriate scheme. In some embodiments, the encoder 107 may further interleave, multiplex into a single data stream, or embed the encoded MASA metadata, audio object metadata, and stream separation metadata within the encoded (downmixed) transport audio signal before transmission or storage, as indicated by the dashed lines in FIG. 1. The multiplexing may be performed using any appropriate scheme.
[0111] 1 also shows an associated decoder and renderer 109 which is arranged to take the bitstream 118 containing the encoded metadata 116, the encoded transport audio signal 138 and the encoded audio object metadata 112 and to generate therefrom an appropriate spatial audio output signal. The decoding and processing of such audio signals is known in principle and will not be described in detail hereinafter, apart from the decoding of the encoded ISM ratio metadata.
[0112] With reference to FIG. 2, the encoder controller 115 is shown in further detail according to some embodiments.
[0113] In this example, the encoder controller 115 comprises a bitrate determiner / monitor 201 configured to determine and / or monitor the bitrate available for the encoded audio and metadata bandwidth. This may be determined based on an estimate of the transmission path bandwidth (and, for example, based on estimated signal strength), or based on a determination of bandwidth storage so as to keep the file below the required size for a determined period of time, or by any suitable method.
[0114] The bitrate determiner / monitor 201 may further be configured to control the coding mode selector 203. The encoder controller 115 may comprise the coding mode selector 203, which is configured to select a coding mode based on, for example, a determined bandwidth or bitrate, and then control an encoder, for example, the combined encoder core 109 and audio object metadata encoder 111.
[0115] With reference to Figure 3, there is shown a flow diagram of an exemplary operation of the encoder controller shown in Figure 2. In this example, there is an initial operation of receiving, obtaining or otherwise determining the encoded parameters and bit rate or bandwidth of the audio data, as shown in step 301 of Figure 3.
[0116] Once the available bandwidth or bit rate has been obtained, a check can then be made to determine whether the bit rate is below a first (or lowest, or object minimum) threshold, as shown in step 303 of FIG. 3.
[0117] If the available bandwidth or bit rate is below a first (or lowest, or object minimum) threshold, the encoder can be controlled to encode only the transport channel and MASA metadata (also shown as Mode A), as shown in step 304 of FIG.
[0118] If the available bandwidth or bit rate is above the first (or lowest, or object's smallest) threshold, a further check can be made to determine whether the bit rate is below a second (or lower, or one object's) threshold, as shown in step 305 of FIG. 3 .
[0119] If the available bandwidth or bit rate is below a second (or lower, or one object) threshold, the encoder may be controlled to encode the transport channel, MASA metadata, ISM metadata (all objects), MASA-to-overall ratio, and ISM ratio (also shown as Mode B), as shown in step 306 of FIG. 3 .
[0120] If the available bandwidth or bit rate is above the second (or low, or one object) threshold, a further check can be made to determine whether the bit rate is below a third, high, or full object threshold, as shown in step 307 of FIG. 3 .
[0121] If the available bandwidth or bit rate is below a third, high, or full object threshold, the encoder may be controlled to encode the transport channel, MASA metadata, ISM metadata (all objects), MASA-to-total ratio, ISM ratio, and audio data of one object, along with an identifier of the one object, as shown in step 308 of FIG. 3 (also shown as Mode C).
[0122] If the available bandwidth or bit rate is above a third, high, or maximum full object threshold, the encoder may be controlled to encode the transport channel, MASA metadata, ISM metadata (all objects), and audio data of all objects (also shown as Mode D), as shown in step 310 of FIG. 3 .
[0123] The encoding modes can be summarized, for example, by the following table: [Table 1]
[0124] It will be understood that the bit rates given herein are examples and that they may be other specific values.
[0125] 4-7, flow diagrams are shown illustrating the first (or lowest or combined) encoding mode shown in step 304 of FIG. 3, the second (or low or object metadata) encoding mode shown in step 306 of FIG. 3, the third (or high or one object) encoding mode shown in step 308 of FIG. 3, and the fourth (or highest or all objects) encoding mode shown in step 310 of FIG. 3, respectively.
[0126] For example, the Mode A encoding method of Figure 4 illustrates in more detail the first (or lowest, or combined) encoding mode shown in step 304 of Figure 3. Thus, for very low total bit rates (e.g., 32 kbps or less), all encoding is performed using the MASA representation.
[0127] Thus, as shown in step 401 of FIG. 4, there is an operation of receiving / obtaining, for example, object-based streams (independent streams with metadata) and multi-channel-based (MASA stream) transport audio signals and metadata.
[0128] Next, there is the act of generating an object-based MASA stream from the object stream (an independent stream with metadata), as shown in step 403 of Figure 4. In some embodiments, this object-based MASA stream can be created from the object stream using, for example, the methods presented in WO2019086757A1.
[0129] The object-based MASA stream and the multi-channel-based MASA stream are then combined, as shown in step 405 of Figure 4. In some embodiments, the original MASA stream and the MASA stream created from the objects can be combined using the methods set out in GB2574238. The decoder obtains the objects and the MASA audio content in MASA format.
[0130] The combined stream is then output, as shown in step 407 of Figure 4. In such an embodiment, the object audio content is present in the decoded audio scene (along with the MASA audio content), but it is not possible to edit the object, nor to separate the object from the scene at the decoder.
[0131] Figure 5 shows the encoding method of Mode B shown in step 306 of Figure 3, the second (or low or object metadata) encoding mode. Thus, for low bit rates (e.g., between 48 kbps and 80 kbps), more bits are available, so it is possible to parameterize the audio scene by transmitting, for each time-frequency tile, one common audio data downmix, MASA metadata, ISM metadata, and an additional set of parameters indicating how much of the signal from the entire audio scene corresponds to the MASA component (in other words, this can be represented or indicated by a MASA to total energy ratio) and a ratio indicating how the audio scene corresponding to the object is distributed among the ISMs (in other words, this can be represented or indicated by an ISM ratio).
[0132] Thus, as shown in step 501 of FIG. 5, there are method steps for receiving / obtaining, for example, object-based streams (independent streams with metadata) and multi-channel-based (MASA stream) transport audio signals and metadata.
[0133] Next, the combined MASA and object-based downmix (channel pair element) audio signal are generated, as shown in step 503 of Figure 5. In other words, the audio content of the MASA and the object is downmixed to two channels (channel pair element CPE).
[0134] As shown in step 505 of FIG. 5, the MASA to total ratio and ISM ratio may be determined.
[0135] The MASA-to-total ratio and ISM ratio can then be coded based on any suitable coding method. For example, the ISM ratio can be coded using lattice coding or other vector quantization methods, or the MASA-to-total ratio can be coded by a DCT transform followed by entropy coding (such as described in WO2022 / 200666). The coding of the MASA-to-total ratio and ISM ratio is shown in step 507 of FIG. 5.
[0136] Furthermore, the MASA metadata may then be encoded based on any suitable MASA metadata encoding method, as shown in step 509 of FIG.
[0137] The combined audio signal may then be encoded based on any suitable audio signal encoding method, as shown in step 511 of FIG.
[0138] The encoder may then output the encoded MASA metadata, the MASA-to-total ratio, the ISM ratio, and the combined transport audio signal, as shown in step 513 of FIG.
[0139] FIG. 6 illustrates the Mode C coding method shown in step 308 of FIG. 3, the third (or high, or single-object) coding mode. Thus, at medium or higher bit rates (e.g., bit rates above 96 kbps and below 160 kbps), the audio content of one object is separated and transmitted independently. Furthermore, the downmix formed from the MASA transport channel and the remainder of the object are transmitted in MASA format, along with additional parameters for the MASA-to-total energy ratio and the ISM ratio. Furthermore, ISM metadata is transmitted, and an identifier describes which object has been separated. For each frame, a decision is made as to which object to separate. The decision may be based, for example, on the relative level of the object relative to the other objects (e.g., isolating the loudest object). This is described in detail in WO 2022 / 214730.
[0140] Thus, as shown in step 601 of FIG. 6, there are method steps for receiving / obtaining, for example, object-based streams (independent streams with metadata) and multi-channel-based (MASA stream) transport audio signals and metadata.
[0141] Next, as shown in step 603 of FIG. 6, one audio object is selected, and an object identifier is generated based on the selected audio object. Furthermore, the audio signal associated with the selected audio object is encoded. Any appropriate audio signal encoder may be used to encode the audio signal of the selected object. For example, the same or similar audio signal encoder as used to encode the MASA audio signal(s) may be employed. Next, a combined MASA and the remaining (or unselected) object-based transport audio signal (or downmix) is generated, as shown in step 605 of FIG. 6. The object transport signal may be created in the same manner as presented in the previous mode, Mode B, except that the selected or separated object is not included in the mix. For example, the multi-channel or MASA audio signal and the (unselected) object transport signal may be summed to generate the combined transport audio signal.
[0142] The MASA to overall ratio and ISM ratio may be determined as shown in step 607 of FIG.
[0143] Next, as shown in step 609 of Figure 6, the object identifiers, MASA metadata, object metadata for all objects, the MASA-to-total ratio, and the ISM ratio may be encoded based on any suitable encoding method. For example, the encoding may employ any scalar or vector quantizer, with or without entropy encoding. The encoding of the MASA-to-total energy ratio may be performed in the manner described in WO2022 / 200666. The encoding of the ISM ratio is described in more detail below.
[0144] The combined audio signal may then be encoded based on any suitable MASA audio signal encoding method, as shown in step 611 of Figure 6. The encoding of the combined transport audio signal may employ any suitable transport audio signal encoding, for example, a MASA encoder.
[0145] In other words, the separated objects are determined, separated and coded as described in WO2022 / 214730, and for the remaining objects and MASA streams the processing works as described in WO2022 / 200666.
[0146] The encoder may then output the encoded object identifiers, MASA metadata, MASA-to-total ratio, ISM ratio, object metadata (for all objects), selected single object audio signals, and the combined transport audio signal, as shown in step 613 of FIG. 6.
[0147] Figure 7 illustrates the Mode D coding method, the fourth (or highest, or all objects) coding mode, as shown in step 310 of Figure 3. Thus, at high bit rates (e.g., bit rates above 160 kbps), the two input audio formats, MASA and ISM, are coded and transmitted independently.
[0148] Thus, as shown in step 701 of FIG. 7, there are method steps for receiving / obtaining, for example, object-based streams (independent streams with metadata) and multi-channel-based (MASA stream) transport audio signals and metadata.
[0149] Next, as shown in step 703 of FIG. 7, the multi-channel based (MASA stream) transport audio signal and metadata are encoded based on any suitable MASA encoding method.
[0150] The objects (independent streams with metadata) and associated metadata can be further encoded as shown in step 705 of Figure 7. The encoding can be performed employing any suitable mono encoder (as part of the main encoder), such as, for example, an EVS-based mono encoder.
[0151] Next, as shown in step 707 of FIG. 7, the encoder can output the independently encoded objects (independent streams with metadata) and associated metadata, as well as the independently encoded multi-channel-based (MASA stream) transport audio signal and metadata.
[0152] In the following, the generation and encoding of ISM ratio values as determined and encoded within encoding modes B and C will be described in further detail.
[0153] The following is described with respect to encoding of object-related directional metadata parameters (determined from ISM metadata), and in some embodiments, the encoding of object-related directional metadata parameters when the encoder is operating in Mode B and Mode C (second or third) encoding modes. However, it will be understood that in some embodiments the following may also apply to any encoding mode in which audio object metadata (and specifically directional metadata) is encoded. In embodiments in which the MASA-to-overall ratio and ISM ratio are not transmitted, further information regarding encoding details (bit allocation) is transmitted. Furthermore, in some embodiments the ratios could in principle be calculated from the encoded signal, both at the encoder and at the decoder.
[0154] 8, an audio object metadata encoder 111 according to some embodiments is shown in more detail. In the example below, the audio object metadata encoder 111 is configured to receive as input ISM metadata, in particular the MASA-to-total energy ratio m2t 812, the ISM ratio r 814, and the direction value 802. In other words, the ISM metadata contains direction information (elevation and azimuth) for each object in each frame. There is a standard resolution of 11 bits per elevation-azimuth pair. Furthermore, the MASA-to-total ratio and the ISM ratio can be determined or obtained from the ISM and also from the multi-channel audio signal (e.g., as described in WO2022 / 200666).
[0155] In some embodiments, the ISM ratio can be obtained, for example, as follows:
[0156] First, the object audio signal S obj (t,i) is the time-frequency domain S obj (b,n,i) where t is the time sample index, b is the frequency bin index, n is the time frame index, and i is the object index. The time-frequency domain signal can be obtained, for example, via a short-time Fourier transform (STFT) or a complex-modulated quadrature filter bank (QMF) (or their low-delay variants).
[0157] The energy of the object is then calculated in the following frequency bands:
number
number
[0158] In some embodiments, the temporal resolution of the ISM ratio is determined by the time-frequency domain audio signal S obj The temporal resolution of (b,n,i) may differ (i.e., the temporal resolution of the spatial metadata may differ from the temporal resolution of the time-frequency transform). In such cases, the calculation (of the energy ratio and / or ISM ratio) may involve summing over multiple time frames of the time-frequency domain audio signals and / or energy values.
[0159] ISM ratios are numbers between 0 and 1, and correspond to the proportion to which an object is active within the audio scene created by all objects. For each object, there is one ISM ratio per frequency subband and time subframe. As mentioned above, the ISM ratios are passed to the Audio Object Metadata Encoder 111.
[0160] In some embodiments, the audio object metadata encoder 111 is configured to encode the MASA-to-total ratio and the ISM ratio, the particular encoding of the ISM ratio and the MASA-to-total ratio not being described in further detail herein, for example WO2022 / 200666 describes a suitable MASA-to-total ratio encoding method, and GB applications 2217884.2 and 2217905.5 describe a suitable ISM ratio encoding method.
[0161] In some embodiments, there is an (audio) object priority determiner 803. The object priority determiner 803 is configured to generate priority values for objects within a time-frequency tile. In some embodiments, the object priority determiner 803 is configured to obtain a MASA-to-total ratio 812 and an ISM ratio 814. For each time-frequency (TF) tile (i.e., subband-subframe combination), there is one MASA-to-total energy ratio and N ISM ratios, where N is the number of objects.
[0162] Furthermore, in some embodiments, the priority value may be defined as the maximum value across all time-frequency tiles of the object contribution.
[0163] For example, the priority value may be generated based on the following:
number
[0164] In the above equation, there are N objects, B subbands, and M temporal subframes. mt(b,k) denotes the quantized MASA-to-total energy ratio for subband b and subframe k. r(i,b,k) is the quantized ISM ratio for object i in subband b and subframe k. In the following, quantized versions of the MASA-to-total and ISM ratios are used to make the same values available at the decoder.
[0165] In some embodiments, some other parameters (other than the MASA-to-total energy ratio) can be employed. For example, the object-to-total energy ratio may be something like o2t(b,k)=1−m2t(b,k), and thus may effectively convey the same information.
[0166] In such embodiments where alternative parameters are employed, the above equations are updated accordingly. Thus, in the object-to-total energy ratio example above, object-to-total o2t(b,k) may simply replace the (1-m2t(b,k)) portion.
[0167] Alternatively, in some embodiments, the (weighted) average of object contributions within a TF tile can also be considered to select the priority. Also, in some other embodiments, the max operator can be replaced by a second or third maximum value (or any similar value), so that the highest contribution in a single TF tile does not alone determine the priority.
[0168] The object priority value 804 may then be passed to a bit determiner 805 .
[0169] In some embodiments, the audio object metadata encoder 111 includes a bit determiner 805. The bit determiner 805 is configured to obtain or receive the object priorities 804 and to determine, based on these object priority values, the number of bits available for encoding the direction parameters. In other words, it sets the number of bits that defines the quantization grid used to encode the direction values 802.
[0170] In some embodiments, the bit determiner 805 is configured to determine or allocate fewer bits to objects with lower priority.
[0171] As an example, based on object priority, the number of bits allocated per object orientation metadata can be calculated as follows: bits(i)=11-[(1-p(i))*7],i=0:N-1
[0172] The [] operator represents rounding to the nearest integer arithmetic. The maximum number of bits per object direction is 11, and the minimum number of bits is 4. The maximum number of bits and the minimum number of bits may be other values based on implementation details, and thus the value 7 (the difference between the maximum and minimum number of bits per object direction) may vary in some other embodiments. While this example shows a linear scale, the above formula providing the number of bits can be replaced with any other increasing linear or non-linear function of object priority, ensuring that the number of bits is in or similar to the range [4,11].
[0173] The bit allocation 806 may then be passed to a spherical grid determiner and encoder 807 .
[0174] In some embodiments, the audio object metadata encoder 111 comprises a spherical grid determiner and encoder 807 configured to receive a bit allocation 806 and direction values 802. The spherical grid determiner and encoder 807 is then configured to generate output index values based on the closest points to the direction values 802 in the determined spherical grid defined by the object's bit allocation 806. The encoded direction index values 808 can be output or further processed.
[0175] Other embodiments may be configured to quantize the orientation of each object using a corresponding number of bits to output an elevation index and an azimuth index. The object metadata may be coded independently for each object, or the orientation metadata of the objects may be coded jointly, for example using some weighting to accommodate different bit allocations.
[0176] The highest resolution (11 bit) spherical grid used for directional quantization guarantees a quantization resolution of 5 degrees. The structure of the spherical grid is the same as that used for quantizing MASA directional metadata (and methods for defining grids and encoding indices for MASA directional metadata values are known, as described in PCT / EP2017 / 078948, GB1811071.8).
[0177] In some embodiments, if an object's priority is exactly zero, the number of bits for that object is set to zero and no directional metadata is sent for that object.
[0178] With reference to FIG. 9, a flow diagram summarizing the operation of the exemplary audio object metadata encoder 111 shown in FIG. 8 is shown.
[0179] The initial operation is to receive / acquire an independent stream with metadata and the determined (quantized) MASA to total energy ratio (m2t) and ISM ratio (r), as shown in step 901 of Figure 9.
[0180] The next operation is then to determine the object priority based on the (quantized) ISM ratio value (r) and the (quantized) MASA to total energy ratio (m2t), as shown in step 903 of Figure 9.
[0181] Once the object priority is determined, the bit allocation for the direction parameters is then determined based on the determined object priority, as shown in step 905 in FIG.
[0182] Next, from the bit allocation, a spherical grid for encoding the object's orientation parameters is determined, as shown in step 907 of FIG.
[0183] Next, as shown in step 909 of FIG. 9, we generate coded directional indices within the determined spherical grid.
[0184] The encoded direction index values can then be output for inclusion in the bitstream, as shown in step 911 of FIG.
[0185] For a decoder, decoding of the coded directional information can be performed by determining a similar prioritization decision and determining the associated quantization grid. Thus, for example, a pseudo-code representation of the coding and decoding can be as follows: Encoding Object Metadata 1. For each object a. Calculate the priority p(i) b. Calculate the number of bits (i) 2.End for 3. For each object a. Encode the spherical grid orientation metadata corresponding to the number of bits allocated to the object 4.End for Decrypting object metadata 1. For each object a. Calculate the priority p(i) b. Calculate the number of bits (i) 2.End for 3. For each object a. Decode the spherical grid orientation metadata corresponding to the number of bits allocated to the object 4.End for
[0186] In some embodiments, the ISM ratio and / or MASA to total energy ratio may be determined solely to determine the priority for coding modes other than B and C (e.g., coding mode D). In such embodiments, some additional information regarding bit allocation may also be transmitted.
[0187] 10 is an exemplary electronic device that may be used as any of the apparatus portions of the system as described above. The device may be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc. The device may be configured to implement, for example, the encoder / analyzer portion and / or decoder portion shown in FIG. 1 or any of the functional blocks described above.
[0188] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. Processor 1407 can be configured to execute various program code, such as the methods described herein.
[0189] In some embodiments, device 1400 comprises at least one memory 1411. In some embodiments, at least one processor 1407 is coupled to memory 1411. Memory 1411 may be any suitable storage means. In some embodiments, memory 1411 includes program code sections for storing program code executable on processor 1407. Additionally, in some embodiments, memory 1411 may further include stored data sections for storing data, e.g., data that has been processed or to be processed according to embodiments described herein. The implemented program code stored in the program code sections and the data stored in the stored data sections can be retrieved by processor 1407 when needed via the memory-processor coupling.
[0190] In some embodiments, device 1400 comprises a user interface 1405. User interface 1405, in some embodiments, may be coupled to a processor 1407. In some embodiments, processor 1407 may control the operation of user interface 1405 and receive input from user interface 1405. In some embodiments, user interface 1405 may allow a user to input commands to device 1400, for example, via a keypad. In some embodiments, user interface 1405 may allow a user to obtain information from device 1400. For example, user interface 1405 may include a display configured to display information from device 1400 to the user. User interface 1405, in some embodiments, may include a touch screen or touch interface that can both allow information to be input into device 1400 and also display information to the user of device 1400. In some embodiments, user interface 1405 may be a user interface for communication.
[0191] In some embodiments, device 1400 includes an input / output port 1409. In some embodiments, input / output port 1409 includes a transceiver. The transceiver in such embodiments may be coupled to processor 1407 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. The transceiver, or any suitable transceiver or transmitting and / or receiving means, may in some embodiments be configured to communicate with other electronic devices or apparatuses via a wire or wired coupling.
[0192] The transceiver may communicate with the further device via any suitable known communication protocol. For example, in some embodiments, the transceiver may use a suitable radio access architecture based on Long Term Evolution Advanced (LTE-Advanced, LTE-A) or New Radio (NR) (also referred to as 5G), Universal Mobile Telecommunications System (UMTS) Radio Access Network (UTRAN or E-UTRAN), Long Term Evolution (same as LTE, E-UTRA), 2G network (legacy network technology), Wireless Local Area Network (WLAN or Wi-Fi), Worldwide Interoperability for Microwave Access (WiMAX), Bluetooth®, Personal Communications Services (PCS), ZigBee®, Wideband Code Division Multiple Access (WCDMA), systems using Ultra Wideband (UWB) technology, sensor networks, Mobile Ad Hoc Networks (MANET), Cellular Internet of Things (IoT) RAN, Internet Protocol Multimedia Subsystem (IMS), any other suitable options, and / or any combination thereof.
[0193] The input / output port 1409 of the transceiver may be configured to receive a signal.
[0194] In some embodiments, device 1400 may be employed as at least part of a synthesis device. Input / output port 1409 may be coupled to headphones (which may be head-tracked or non-tracked headphones) or the like, and loudspeakers.
[0195] Generally, various embodiments of the present invention may be implemented in hardware or special purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the present invention may be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it will be appreciated that these blocks, apparatus, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing device, or some combination thereof.
[0196] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device, such as within a processor entity, or by hardware, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that any block of logic flow in the diagrams may represent program steps, or interconnected logic circuits, blocks and functions, or combinations of program steps and logic circuits, blocks and functions. Software may be stored on physical media, such as memory chips or blocks implemented within a processor, magnetic media, such as hard disks or floppy disks, and optical media, such as DVDs and their data variants, CDs.
[0197] The memory may be any type of memory suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be any type of memory suitable for the local technology environment and may include, by way of non-limiting examples, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processor architecture.
[0198] Embodiments of the present invention may be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert logic-level designs into semiconductor circuit designs ready to be etched onto semiconductor substrates.
[0199] Programs such as those offered by Synopsys, Inc. of Mountain View, Calif., and Cadence Design of San Jose, Calif., automatically route conductors and place components on semiconductor chips using well-established design rules and pre-stored libraries of design modules. Once the design of a semiconductor circuit is complete, the resulting design may be sent in a standardized electronic format (e.g., Opus, GDSII, etc.) to a semiconductor manufacturing facility or "fab" for fabrication.
[0200] As used in this application, the term "circuitry" may refer to one or more or all of the following: (a) Hardware-only circuit implementations (e.g., implementations using only analog and / or digital circuits); (b) Combinations of hardware circuitry and software, such as (where applicable): (i) a combination of analog and / or digital hardware circuitry(s) and software / firmware; (ii) any portion of software-based hardware processor(s) (including digital signal processor(s)), software, and memory(s) that cooperate to cause a device, such as a mobile phone or server, to perform various functions; and Hardware circuitry(s) and / or processor(s), such as microprocessor(s) or portions of microprocessor(s), that require software (e.g., firmware) to operate, but software may be absent if not necessary for operation.
[0201] This definition of circuit applies to all uses of the term in this application, including any claims. As a further example, as used in this application, the term circuit also encompasses implementations of simply a hardware circuit or processor(s), or portions of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuit also encompasses, for example, baseband or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, other computing devices, or other network devices, where applicable to certain claim elements.
[0202] The term "non-transitory" as used herein is not a limitation regarding the permanence of the data storage (eg, RAM vs. ROM), but rather a limitation of the medium itself (ie, tangible rather than signal).
[0203] As used herein, "at least one of:" and "at least one of" and similar phrases, when a list of two or more elements is joined by "and" or "or," mean at least any one of the elements, or at least any two or more of the elements, or at least all of the elements.
[0204] The foregoing description provides a complete and detailed description of exemplary embodiments of the present invention, by way of illustrative and non-limiting example. However, various modifications and adaptations may become apparent to those skilled in the art in light of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of the present invention, as defined by the appended claims.
Claims
1. 1. An apparatus comprising: means for obtaining respective audio signal and directional metadata values for a plurality of audio objects within an audio environment; means for obtaining at least one proportion parameter for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, the frame comprising at least one time element and at least one frequency element, the at least one proportion parameter being configured to identify a distribution of the audio object within an object portion of an overall audio environment; means for obtaining a further proportion parameter related to the distribution of the object parts in the overall audio environment; means for determining object priorities based on the ratio parameter and the further ratio parameter of the plurality of audio objects for at least one time element and at least one frequency element of the frame; means for determining a quantization bit allocation for the directional metadata values associated with the plurality of audio objects based on the determined object priorities; means for quantizing the directional metadata values associated with the plurality of audio objects based on the quantization bit allocation; The device comprising:
2. 2. The apparatus of claim 1, wherein the means for determining a quantization bit allocation for the plurality of objects based on the determined object priorities further comprises: means for determining a quantization bit allocation for the plurality of audio objects based on the determined object priorities.
3. 3. The apparatus of claim 1, wherein the means for determining quantization bit allocations for the directional metadata values associated with the plurality of audio objects based on the determined object priorities is means for determining a spherical grid quantization configuration based on the quantization bit allocations.
4. 4. The apparatus of claim 3, wherein the means for quantizing the directional metadata values associated with the plurality of audio objects based on the quantization bit allocation comprises means for generating, for the directional metadata value, an index value representing a nearest node on the determined spherical grid quantization configuration.
5. 5. The apparatus of claim 1, wherein the means for determining quantization bit allocations for the directional metadata values associated with the plurality of audio objects based on the determined object priorities is means for determining or allocating fewer bits to objects with lower priority.
6. 6. The apparatus of claim 1, wherein the means for determining a quantization bit allocation for the directional metadata values associated with the plurality of audio objects based on the determined object priorities is means for determining or allocating a first number of bits to an object with a highest priority, a second number of bits to a further object with a lowest priority, and a further number of bits between the first and second number to objects between the highest and lowest priorities.
7. 7. The apparatus of claim 6, wherein the additional number of bits between the first number and the second number of objects between the highest priority and the lowest priority are linearly scaled within the range between the first number and the second number.
8. 7. The apparatus of claim 6, wherein the additional number of bits between the first number and the second number of objects between the highest priority and the lowest priority are scaled nonlinearly within a region between the first number and the second number.
9. 9. The apparatus of claim 1, wherein the means for determining the object priority based on the ratio parameter and the further ratio parameter of the plurality of audio objects for at least one time element and at least one frequency element of the frame is means for determining the object priority based on an average of the ratio parameter and the further ratio parameter of a particular object over at least two time elements, or at least two frequency elements, or at least two time and frequency elements of the frame.
10. 9. The apparatus according to claim 1, wherein the means for determining the object priority based on the ratio parameters and the further ratio parameters of the plurality of audio objects for at least one time element and at least one frequency element of the frame is means for determining the object priority based on a selection from the ratio parameters of a particular audio object across at least two time elements, or at least two frequency elements, or at least two time and frequency elements.
11. 11. The apparatus of claim 10, wherein the selection is the maximum value of the ratio parameter of the particular audio object over the at least two time elements, or the at least two frequency elements, or the at least two time and frequency elements.
12. 1. An apparatus comprising: means for obtaining, for a plurality of audio objects in an audio environment, respective encoded audio signals and encoded directional metadata values; means for obtaining at least one encoded proportion parameter for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, the frame comprising at least one time element and at least one frequency element, the at least one proportion parameter being configured to identify a distribution of the audio object within an object portion of an overall audio environment; means for obtaining further encoded proportion parameters related to the distribution of the object parts in the overall audio environment; means for determining object priorities based on the coded ratio parameters of the plurality of audio objects and the further ratio parameter with respect to at least one time element and at least one frequency element of the frame and the further ratio parameter; means for determining a quantization bit allocation for the encoded directional metadata values associated with the plurality of audio objects based on the determined object priorities; means for decoding the encoded directional metadata values associated with the plurality of audio objects based on the quantization bit allocation; The device comprising:
13. 13. The apparatus of claim 12, further comprising: means for obtaining a bit allocation for the plurality of objects; and means for determining a quantization bit allocation for the encoded directional metadata values associated with the plurality of audio objects based on the determined object priorities, further comprising means for determining a quantization bit allocation based on the bit allocation.
14. 14. The apparatus of claim 12, wherein the means for determining quantization bit allocations for the encoded directional metadata values associated with the plurality of audio objects based on the determined object priorities is means for determining a spherical grid quantization configuration based on the quantization bit allocations.
15. 15. The apparatus of claim 14, wherein the means for decoding the encoded directional metadata values associated with the plurality of audio objects based on the quantization bit allocation is means for determining directional values associated with encoded directional metadata values that represent nodes on the determined spherical grid quantization configuration.
16. 16. The apparatus of claim 12, wherein the means for determining quantization bit allocations for the directional metadata values associated with the plurality of audio objects based on the determined object priorities is means for determining or allocating fewer bits to objects with lower priority.
17. 17. The apparatus of claim 12, wherein the means for determining a quantization bit allocation for the directional metadata values associated with the plurality of audio objects based on the determined object priorities is means for determining or allocating a first number of bits to an object with a highest priority, a second number of bits to a further object with a lowest priority, and a further number of bits between the first and second number to objects between the highest and lowest priority.
18. 18. The apparatus of claim 17, wherein the additional number of bits between the first number and the second number of objects between the highest priority object and the lowest priority object is linearly scaled within a range between the first number and the second number.
19. 18. The apparatus of claim 17, wherein the additional number of bits between the first number and the second number of objects between the highest priority and the lowest priority are scaled nonlinearly within a region between the first number and the second number.
20. 20. The apparatus of claim 12, wherein the means for determining the object priority based on the ratio parameters and the further ratio parameters of the plurality of audio objects for at least one time element and at least one frequency element of the frame is means for determining the object priority based on an average of the coded ratio parameters and the further coded ratio parameters of a particular object over at least two time elements, or at least two frequency elements, or at least two time and frequency elements of the frame.
21. 20. The apparatus of claim 12, wherein the means for determining the object priority based on the ratio parameters and the further ratio parameters of the plurality of audio objects for at least one time element and at least one frequency element of the frame is means for determining the object priority based on a selection from the coded ratio parameters of a particular audio object across at least two time elements, or at least two frequency elements, or at least two time and frequency elements.
22. 22. The apparatus of claim 21, wherein the selection is the maximum value of the encoded ratio parameter of the particular audio object over the at least two time elements, or the at least two frequency elements, or the at least two time and frequency elements.
23. 1. A method comprising: obtaining respective audio signal and directional metadata values for a plurality of audio objects within an audio environment; - obtaining at least one proportion parameter for each audio object of the plurality of audio objects and for each time element and each frequency element of a frame, the frame comprising at least one time element and at least one frequency element, the at least one proportion parameter being configured to identify a distribution of the audio object within an object portion of the overall audio environment; obtaining a further proportion parameter related to the distribution of the object parts in the overall audio environment; determining an object priority based on the ratio parameter and the further ratio parameter of the plurality of audio objects for at least one time element and at least one frequency element of the frame; determining a quantization bit allocation for the directional metadata values associated with the plurality of audio objects based on the determined object priorities; quantizing the directional metadata values associated with the plurality of audio objects based on the quantization bit allocation; The method comprising:
24. 1. A method comprising: obtaining respective encoded audio signals and encoded directional metadata values for a plurality of audio objects in an audio environment; - obtaining at least one encoded proportion parameter for each audio object of said plurality of audio objects and for each time element and each frequency element of a frame, said frame comprising at least one time element and at least one frequency element, said at least one proportion parameter being configured to identify a distribution of said audio object within an object portion of an overall audio environment; - obtaining further encoded proportion parameters related to the distribution of the object parts in the overall audio environment; determining an object priority based on the coded ratio parameters and the further ratio parameters of the plurality of audio objects for at least one time element and at least one frequency element of the frame; determining a quantization bit allocation for the encoded directional metadata values associated with the plurality of audio objects based on the determined object priorities; decoding the encoded directional metadata values associated with the plurality of audio objects based on the quantization bit allocation; The method comprising:
Citation Information
Patent Citations
Voice encoding device and voice decoding device, and program
JP2021124719A
Spatial audio parameter coding and associated decoding decisions
JP2022548038A
Encoding device and method, decoding device and method, and program
WO2014192602A1
Combining spatial audio streams
WO2022200666A1