Parametric Spatial Audio Encoding
The method and apparatus utilize quantization and differential encoding of ratio parameters to enhance the encoding and decoding of spatial audio data, addressing inefficiencies in existing codecs and improving performance in immersive audio systems.
Patent Information
- Application Number
- JP2025531256
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-29
- Filing Date
- 2023-11-07
- Publication Date
- 2026-01-06
AI Technical Summary
Existing immersive audio codecs struggle with efficient encoding and decoding of spatial audio data, particularly in environments with varying bitrates and transmission conditions, leading to suboptimal performance in terms of latency and error robustness.
A method and apparatus for encoding and decoding audio object parameters using ratio parameters, such as ISM and MASA ratios, through quantization and differential encoding techniques, including Golomb-Rice entropy coding, to efficiently represent spatial audio data.
Enhances the encoding and decoding of spatial audio data, improving latency and error robustness while maintaining high fidelity across different bitrates and transmission conditions.
Smart Images

Figure 2026500131000001_ABST
Abstract
Description
[Technical Field]
[0001] This application relates to an apparatus and method for spatial audio representation and encoding, but is not limited to audio representation in an audio encoder. [Background technology]
[0002] Parametric spatial audio processing is a field of audio signal processing in which the spatial aspects of sound are described using a set of parameters. For example, in parametric spatial audio capture from a microphone array, it is a common and effective choice to estimate a set of parameters from the microphone array signal, such as the direction of the sound within a frequency band and the ratio between the directional and omnidirectional portions of the captured sound within the frequency band. These parameters are known to well describe the perceived spatial characteristics of the sound captured at the location of the microphone array. Accordingly, these parameters can be utilized for spatial sound synthesis, for binaural headphones, for loudspeakers, or other formats such as Ambisonics.
[0003] Therefore, the direct-to-total energy ratio within a direction and frequency band is a particularly effective parameterization for capturing spatial audio.
[0004] A parameter set consisting of a direction parameter within a frequency band and an energy ratio parameter within a frequency band (indicating the directionality of the sound) can also be utilized as spatial metadata of an audio codec (which may also include other parameters such as surround coherence, spread coherence, number of directions, distance, etc.) For example, these parameters can be estimated from an audio signal captured by a microphone array, and e.g., a stereo or mono signal can be generated from a microphone array signal conveyed together with spatial metadata.
[0005] Immersive audio codecs are being implemented to support a number of operating points, ranging from low-bitrate operation to transparency. An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, which is designed for use over communication networks such as 3GPP 4G / 5G networks, including for immersive services such as immersive voice and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general-purpose audio. It is also expected to support channel-based and scene-based audio inputs, including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency to enable conversational services, as well as support high error robustness under various transmission conditions.
[0006] A stereo signal may be encoded with, for example, an AAC encoder, and a mono signal may be encoded with an EVS encoder. The decoder may decode the audio signal into a PCM signal and process the sounds in the frequency band (using spatial metadata) to obtain a spatial output, for example a binaural output.
[0007] The aforementioned immersive audio codecs are particularly suited to encoding spatial sound captured from microphone arrays (e.g., in mobile phones, VR cameras, standalone microphone arrays), although such encoders can have other input types, e.g., loudspeaker signals, audio object signals, ambisonic signals. Summary of the Invention
[0008] According to a first aspect, there is provided an apparatus for encoding audio object parameters, the apparatus comprising: means for obtaining, for time-frequency elements of a frame comprising more than one time element and more than one frequency element, a plurality of ratio parameters of audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for a specific time-frequency element, a distribution of specific objects within object portions of the overall audio environment; quantizing a selection of the ratio parameters, the selections being associated with audio objects in the specific frame time-frequency element; encoding a first set of selections of ratio parameters based on indexing of the selections; and encoding the remaining selections of ratio parameters for the frame based on differential encoding of the selections based on the first set of selections of ratio parameters or on selections of previously indexed time elements or frequency elements of the ratio parameters.
[0009] The selection may be a vector of ratio parameters and the means may be further for generating a vector of ratio parameters representing the ratio parameters.
[0010] The means for encoding the first set of ratio parameter selections based on indexing of the selections may be for generating integer values based on indexing from the selections, the generated integer values representing the ratio parameters of the audio objects.
[0011] The means for generating integer values based on indexing from the selection of ratio parameters may be for generating a singular value by appending an element from the selection of ratio parameters, and for generating an index from the singular value by executing an iteration loop from the zeroth iteration up to a single iteration number and sequentially associating index values with iteration numbers of the iteration loop having valid selections of ratio parameters, where the integer value is the highest index value.
[0012] The means for quantizing the selection of ratio parameters may be for quantizing ratio values within a specific selection using minimum nearest neighbor scalar quantization to obtain quantization index values, calculating a reconstructed value of the ratio parameter for the specific selection, calculating an error value based on a difference between the reconstructed ratio value and the specific selection of ratio parameter value, determining a sum of quantization index values, and selecting at least one quantization index value to increment such that the sum of quantization index values equals the sum of expected index values.
[0013] The means for selecting at least one quantization index value to increment so that a sum of the quantization index values equals a sum of the expected index values may be for one of selecting the at least one quantization index value to increment based on identifying a maximum decrease in the error value when the index value is incremented, or selecting the at least one quantization index value to increment based on identifying a minimum increase in the error value when the index value is incremented.
[0014] The means for quantizing the selection of the ratio parameter may be for determining that an element is zero for a specific selection of the ratio parameter, and for generating a further ratio parameter configured to identify a distribution of the object part across the audio environment, the further ratio parameter value identifying an absence of object part contribution.
[0015] The means for differentially encoding the selections based on the first set of ratio parameter selections or the selection of the remaining ratio parameters of the frame based on the selection of a previously indexed time element or frequency element of the ratio parameter may be for performing, for the set of ratio parameter selections for a specific time element of the frame, determining, for the first and second entropy code parameters, the number of bits required to entropy code the differences between the quantized frequency elements; determining, for the first and second entropy code parameters, the number of bits required to entropy code the differences between the quantized time elements; selecting, for the specific time element, the first entropy code parameter or the second entropy code parameter based on the fewer number of bits required to code the differences within the specific time element of the frame; and selecting, for the selected first entropy code parameter or the second entropy code parameter based on the fewer number of bits required to code the differences within the specific time element of the frame, one of the entropy codes of the differences between the frequency elements or the time elements.
[0016] The means for differential encoding of a selection based on a first set of ratio parameters or a selection of a previously indexed time element or frequency element of the ratio parameters may be for encoding a selected one of the entropy codes of the difference between the frequency element or the time element for the selected first entropy code parameter or the second entropy code parameter based entropy code.
[0017] The means for differentially encoding selections of ratio parameters based on the first set of selections of ratio parameters or encoding remaining selections of ratio parameters for the frame based on selections of ratio parameters for previously indexed time elements or frequency elements may be for performing, for the set of selections of ratio parameters for specific frequency elements of the frame, determining, for the first and second entropy code parameters, the number of bits required to entropy code quantized differences between frequency elements; determining, for the first and second entropy code parameters, the number of bits required to entropy code quantized differences between time elements; selecting, for the specific frequency elements, the first entropy code parameter or the second entropy code parameter based on the fewer number of bits required to code differences within the specific time element of the frame; and selecting, for the specific frequency elements, one of the entropy codes for differences between frequency elements or time elements based on the selected first entropy code parameter or the second entropy code parameter based on the fewer number of bits required to code differences within the specific time element of the frame.
[0018] The means for differential encoding of a selection based on a first set of ratio parameters or a selection of a previously indexed time element or frequency element of the ratio parameters may be for encoding a selected one of the entropy codes of the difference between the frequency element or the time element for the selected first entropy code parameter or the second entropy code parameter based entropy code.
[0019] The means for differentially encoding selections of ratio parameters based on the first set of ratio parameter selections or selections of remaining ratio parameters of the frame based on selections of previously indexed time elements or frequency elements of the ratio parameters may be for generating an indicator indicating the selected first entropy code parameter or the second entropy code parameter, and for the selected first entropy code parameter or the second entropy code parameter based entropy code, generating an indicator indicating a selected one of the differential entropy codes between the frequency elements or the time elements.
[0020] The entropy code may be a Golomb-Rice entropy code, the first entropy code parameter being a Golomb-Rice entropy code order of 0, and the second entropy code parameter being a Golomb-Rice entropy code order of 1.
[0021] The means for differentially encoding selections of ratio parameters based on the first set of selections of ratio parameters or encoding selections of remaining ratio parameters of the frame based on selections of previously indexed time elements or frequency elements of the ratio parameters may be for differential encoding selections of ratio parameters based on selections of previously indexed time elements of the ratio parameters where there are no selections of previously indexed frequency elements of the ratio parameters.
[0022] The ratio parameter configured to identify the distribution of specific objects within the object portion of the overall audio environment may be an ISM ratio.
[0023] A further ratio parameter configured to identify the distribution of object parts across the audio environment may be the MASA to total energy ratio.
[0024] According to a second aspect, there is provided an apparatus for decoding audio object parameters, the apparatus comprising: means for obtaining a bitstream comprising encoded ratio parameters for time-frequency elements of a frame comprising more than one time element and more than one frequency element, the ratio parameters being associated with audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for a specific time-frequency element, a distribution of specific objects within object portions of the overall audio environment; decoding a first set of selections of ratio parameters based on indexing of the selections; and decoding remaining selections of ratio parameters for the frame based on differential decoding of the first set of selections of ratio parameters or selections based on previously indexed time element or frequency element selections of ratio parameters.
[0025] The selection may be a vector of ratio parameters.
[0026] The means for decoding the first set of ratio parameter selections based on the selection indexing may be for obtaining integer values representing the encoded ratio parameters, converting the integer values to ratio parameter selections based on the vector indexing, and regenerating at least one further ratio parameter from the ratio parameter selections.
[0027] The means for converting integer values to ratio parameter selections based on vector indexing may be for generating a single value by appending an element from the ratio parameter selection, and generating an index from the single number by executing an iteration loop from the zeroth iteration up to a single iteration number, and sequentially associating index values with iteration numbers of the iteration loop having valid ratio parameter selections, wherein the integer value is the most significant index value.
[0028] The means for decoding the remaining selections of ratio parameters for the frame based on differential decoding of the first set of selections of ratio parameters or selections based on selections of previously indexed time or frequency elements of the ratio parameters may be for obtaining a differential indicator identifying frequency differential or time differential encoding, obtaining an entropy encoding indicator identifying entropy encoding parameters, and decoding the remaining selections of ratio parameters for the frame based on the differential indicator and the entropy encoding indicator.
[0029] The ratio parameters configured to identify a distribution of specific objects within an object portion of the overall audio environment may be ISM ratios. According to a third aspect, there is provided a method for encoding audio object parameters, the method comprising: obtaining, for time-frequency elements of a frame comprising more than one time element and more than one frequency element, a plurality of ratio parameters of audio objects in an audio environment, the ratio parameters being configured to identify, for the specific time-frequency elements, a distribution of specific objects within an object portion of the overall audio environment; quantizing selections of the ratio parameters, the selections being associated with audio objects within the specific frame time-frequency elements; encoding a first set of selections of ratio parameters based on indexing of the selections; and encoding a remaining selection of ratio parameters for the frame based on differential encoding of the selections based on the first set of selections of ratio parameters or on selections of previously indexed time elements or frequency elements of the ratio parameters.
[0030] The selection may be a vector of ratio parameters, and the method may further include generating a vector of ratio parameters representing the ratio parameters.
[0031] Encoding the first set of ratio parameter selections based on indexing of the selections may include generating integer values based on indexing from the selections, the generated integer values representing the ratio parameters of the audio objects.
[0032] Generating integer values based on indexing from the selection of the ratio parameter may include generating a singular value by appending an element from the selection of the ratio parameter, and generating an index from the singular value by executing an iteration loop from the zeroth iteration up to a number of iterations equal to or less than the single iteration, and sequentially associating index values with iteration numbers of the iteration loop having valid selections of the ratio parameter, where the integer value is the highest index value.
[0033] Quantizing the selection of ratio parameters may include quantizing ratio values within a specific selection using minimum nearest neighbor scalar quantization to obtain quantization index values, calculating a reconstructed value of the ratio parameter for the specific selection, calculating an error value based on a difference between the reconstructed ratio value and the specific selection of ratio parameter value, determining a sum of quantization index values, and selecting at least one quantization index value to increment so that the sum of quantization index values equals the sum of expected index values.
[0034] Selecting at least one quantization index value to increment so that the sum of the quantization index values equals the sum of the expected index values may include one of selecting the at least one quantization index value to increment based on identifying a maximum decrease in the error value when the index value is incremented, or selecting the at least one quantization index value to increment based on identifying a minimum increase in the error value when the index value is incremented.
[0035] Quantizing the selection of the ratio parameter may include determining that an element is zero for a specific selection of the ratio parameter, and generating a further ratio parameter configured to identify a distribution of the object part across the audio environment, wherein the further ratio parameter value identifies an absence of a contribution of the object part.
[0036] Differentially encoding the selections based on the first set of ratio parameter selections, or encoding the remaining selections of ratio parameters for the frame based on selections of previously indexed time elements or frequency elements of the ratio parameters, may include, for the set of ratio parameter selections for specific time elements of the frame, determining the number of bits required to entropy code the differences between the quantized frequency elements for the first and second entropy code parameters; determining the number of bits required to entropy code the differences between the quantized time elements for the first and second entropy code parameters; selecting the first entropy code parameter or the second entropy code parameter for the specific time element based on the fewer number of bits required to code the differences within the specific time element of the frame; and selecting one of the entropy codes of the differences between the frequency elements or the time elements for the selected first entropy code parameter or the second entropy code parameter based on the fewer number of bits required to code the differences within the specific time element of the frame.
[0037] Differential encoding of the selection based on the first set of ratio parameter selections, or the selection of the previously indexed time or frequency elements of the ratio parameters, may include encoding a selected one of the entropy codes of the differences between the frequency or time elements for the selected first entropy code parameter or the second entropy code parameter based entropy code.
[0038] Differentially encoding the ratio parameter selections based on the first set of ratio parameter selections, or encoding the remaining ratio parameter selections of the frame based on selections of ratio parameters for previously indexed time elements or frequency elements, may include, for the set of ratio parameter selections for specific frequency elements of the frame, determining, for first and second entropy code parameters, the number of bits required to entropy code the quantized differences between the frequency elements; determining, for the first and second entropy code parameters, the number of bits required to entropy code the quantized differences between the time elements; selecting, for the specific frequency elements, the first entropy code parameter or the second entropy code parameter based on the fewer number of bits required to code the differences within the specific time element of the frame; and selecting, for the specific frequency elements, one of the entropy codes of the differences between the frequency elements or the time elements for the selected first entropy code parameter or the second entropy code parameter based on the fewer number of bits required to code the differences within the specific time element of the frame.
[0039] Differential encoding of the selection based on the first set of ratio parameter selections, or the selection of the previously indexed time or frequency elements of the ratio parameters, may include encoding a selected one of the entropy codes of the differences between the frequency or time elements for the selected first entropy code parameter or the second entropy code parameter based entropy code.
[0040] Differentially encoding the ratio parameter selections based on the first set of ratio parameter selections, or encoding the remaining ratio parameter selections of the ratio parameters of the frame based on the selection of a previously indexed time element or frequency element of the ratio parameters, may include generating an indicator indicating the selected first entropy code parameter or the second entropy code parameter, and for the selected first entropy code parameter or the second entropy code parameter-based entropy code, generating an indicator indicating a selected one of the differential entropy codes between the frequency element or the time element.
[0041] The entropy code may be a Golomb-Rice entropy code, the first entropy code parameter being a Golomb-Rice entropy code order of 0, and the second entropy code parameter being a Golomb-Rice entropy code order of 1.
[0042] Differentially encoding the selections of ratio parameters based on the first set of selections of ratio parameters, or encoding selections of remaining ratio parameters of the ratio parameters of the frame based on selections of previously indexed time elements or frequency elements of the ratio parameters may include differentially encoding selections of ratio parameters based on selections of previously indexed time elements of the ratio parameters where there are no selections of previously indexed frequency elements of the ratio parameters.
[0043] The ratio parameter configured to identify the distribution of specific objects within the object portion of the overall audio environment may be an ISM ratio.
[0044] A further ratio parameter configured to identify the distribution of object parts across the audio environment may be the MASA to total energy ratio.
[0045] According to a fourth aspect, there is provided a method for decoding audio object parameters, the method comprising: obtaining a bitstream comprising encoded ratio parameters for time-frequency elements of a frame comprising more than one time element and more than one frequency element, the ratio parameters being associated with audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for the specific time-frequency elements, a distribution of specific objects within object portions of the overall audio environment; decoding a first set of selections of ratio parameters based on indexing of the selections; and decoding remaining selections of ratio parameters for the frame based on differential decoding of the first set of selections of ratio parameters or selections based on previously indexed time element or frequency element selections of ratio parameters.
[0046] The selection may be a vector of ratio parameters.
[0047] Decoding the first set of ratio parameter selections based on the selection indexing may include obtaining integer values representing the encoded ratio parameters, converting the integer values to ratio parameter selections based on the vector indexing, and regenerating at least one further ratio parameter from the ratio parameter selections.
[0048] Converting integer values to ratio parameter selections based on vector indexing may include generating singular values by appending elements from the ratio parameter selections, and generating indices from the singular values by executing an iteration loop from the zeroth iteration up to a single iteration number, and sequentially associating index values with iteration numbers of the iteration loop having valid ratio parameter selections, where the integer value is the most significant index value.
[0049] Decoding the remaining selections of ratio parameters for the frame based on differential decoding of the first set of selections of ratio parameters or selections based on selections of previously indexed time or frequency components of the ratio parameters may include obtaining a differential indicator that identifies frequency differential or time differential encoding, obtaining an entropy encoding indicator that identifies entropy encoding parameters, and decoding the remaining selections of ratio parameters for the frame based on the differential indicator and the entropy encoding indicator.
[0050] The ratio parameter configured to identify the distribution of specific objects within the object portion of the overall audio environment may be an ISM ratio.
[0051] According to a fifth aspect, there is provided an apparatus for encoding audio object parameters, the apparatus including at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system to: obtain, for time-frequency elements of a frame comprising at least one time element and more than one frequency element, a plurality of ratio parameters of audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for the specific time-frequency elements, a distribution of specific objects within object portions of the overall audio environment; quantizing a selection of the ratio parameters, the selections being associated with audio objects in the specific frame time-frequency elements; encoding a first set of selections of ratio parameters based on indexing of the selections; and encoding a remaining selection of ratio parameters for the frame based on differential encoding of the selections based on the first set of selections of ratio parameters or on a selection of a previously indexed time element or frequency element of the ratio parameters.
[0052] The selection may be a vector of ratio parameters, and the apparatus may be further adapted to generate a vector of ratio parameters representing the ratio parameters.
[0053] The apparatus adapted to perform encoding the first set of selections of ratio parameters based on indexing of the selections may be adapted to perform generating integer values based on indexing from the selections, the generated integer values representing the ratio parameters of the audio objects.
[0054] An apparatus adapted to perform generating an integer value based on indexing from a selection of a ratio parameter may be adapted to generate a single value by appending an element from the selection of a ratio parameter, and to generate an index from the single number by executing an iteration loop from the zeroth iteration up to a single iteration number, and sequentially associating index values with iteration numbers of the iteration loop having a valid selection of a ratio parameter, wherein the integer value is the most significant index value.
[0055] An apparatus adapted to perform quantizing a selection of ratio parameters may be adapted to quantize ratio values within a specific selection using minimum nearest neighbor scalar quantization to obtain quantization index values, calculate a reconstructed value of the ratio parameter for the specific selection, calculate an error value based on a difference between the reconstructed ratio value and the specific selection of ratio parameter value, determine a sum of quantization index values, and select at least one quantization index value to increment such that the sum of quantization index values equals the sum of expected index values.
[0056] An apparatus adapted to select at least one quantization index value to increment such that a sum of the quantization index values equals a sum of expected index values may be adapted to perform one of: selecting the at least one quantization index value to increment based on identifying a maximum decrease in an error value when the index value is incremented; or selecting the at least one quantization index value to increment based on identifying a minimum increase in the error value when the index value is incremented.
[0057] An apparatus adapted to perform quantizing a selection of ratio parameters may be adapted to perform: determining that an element is zero for a specific selection of ratio parameters; generating a further ratio parameter configured to identify a distribution of object parts across the audio environment, the further ratio parameter value identifying an absence of object part contribution.
[0058] An apparatus configured to perform differential encoding of selections based on a first set of ratio parameter selections, or encoding remaining selections of ratio parameters of a frame based on selections of previously indexed time elements or frequency elements of the ratio parameters, may be configured to perform, for a set of ratio parameter selections for a specific time element of the frame, determining, for the first and second entropy code parameters, the number of bits required to entropy code differences between quantized frequency elements; determining, for the first and second entropy code parameters, the number of bits required to entropy code differences between quantized time elements; selecting, for the specific time element, the first entropy code parameter or the second entropy code parameter based on the fewer number of bits required to code differences within the specific time element of the frame; and selecting one of the entropy codes of differences between frequency elements or time elements for the selected first entropy code parameter or the second entropy code parameter based on the fewer number of bits required to code differences within the specific time element of the frame.
[0059] An apparatus adapted to perform differential encoding of a selection based on a first set of ratio parameter selections or selection of previously indexed time elements or frequency elements of the ratio parameters may be adapted to perform encoding a selected one of the entropy codes of the difference between the frequency elements or the time elements for the selected first entropy code parameter or the second entropy code parameter based entropy code.
[0060] An apparatus configured to perform differential encoding of ratio parameter selections based on a first set of ratio parameter selections, or encoding remaining ratio parameter selections of a frame based on selections of ratio parameters of previously indexed time elements or frequency elements, may be configured to perform, for a set of ratio parameter selections for specific frequency elements of the frame, determining, for first and second entropy code parameters, the number of bits required to entropy code quantized differences between frequency elements; determining, for the first and second entropy code parameters, the number of bits required to entropy code quantized differences between time elements; selecting, for the specific frequency elements, the first entropy code parameter or the second entropy code parameter based on the fewer number of bits required to code differences within the specific time element of the frame; and selecting, for the specific frequency elements, one of the entropy codes of differences between frequency elements or time elements for the selected first entropy code parameter or the second entropy code parameter based on the fewer number of bits required to code differences within the specific time element of the frame.
[0061] An apparatus adapted to perform differential encoding of a selection based on a first set of ratio parameter selections or selection of previously indexed time elements or frequency elements of the ratio parameters may be adapted to perform encoding a selected one of the entropy codes of the difference between the frequency elements or the time elements for the selected first entropy code parameter or the second entropy code parameter based entropy code.
[0062] An apparatus adapted to perform differential encoding of ratio parameter selections based on a first set of ratio parameter selections, or encoding selections of remaining ratio parameters of the frame based on selections of previously indexed time elements or frequency elements of the ratio parameters, may be adapted to generate an indicator indicating the selected first entropy code parameter or the second entropy code parameter, and for the selected first entropy code parameter or second entropy code parameter-based entropy code, generate an indicator indicating a selected one of the differential entropy codes between the frequency elements or the time elements.
[0063] The entropy code may be a Golomb-Rice entropy code, the first entropy code parameter being a Golomb-Rice entropy code order of 0, and the second entropy code parameter being a Golomb-Rice entropy code order of 1.
[0064] An apparatus adapted to perform differential encoding of ratio parameter selections based on the first set of ratio parameter selections or encoding remaining ratio parameter selections of the ratio parameters of the frame based on selections of previously indexed time elements or frequency elements of the ratio parameters may be adapted to perform differential encoding of ratio parameter selections based on selections of previously indexed time elements of the ratio parameters for which there are no selections of previously indexed frequency elements of the ratio parameters.
[0065] The ratio parameter configured to identify the distribution of specific objects within the object portion of the overall audio environment may be an ISM ratio.
[0066] A further ratio parameter configured to identify the distribution of object parts across the audio environment may be the MASA to total energy ratio.
[0067] According to a sixth aspect, there is provided an apparatus for decoding audio object parameters, the apparatus comprising: at least one processor; and at least one memory storing instructions which, when executed by the at least one processor, cause the system to: obtain a bitstream comprising encoded ratio parameters for at least a time-frequency element of a frame comprising more than one time element and more than one frequency element, the ratio parameters being associated with audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for the specific time-frequency element, a distribution of specific objects within object portions of the overall audio environment; decode a first set of selections of ratio parameters based on indexing of the selections; and decode remaining selections of ratio parameters for the frame based on differential decoding of the first set of selections of ratio parameters or selections based on previously indexed time element or frequency element selections of ratio parameters.
[0068] The selection may be a vector of ratio parameters.
[0069] An apparatus adapted to decode the first set of ratio parameter selections based on the selection indexing may be adapted to obtain integer values representing the encoded ratio parameters, convert the integer values to ratio parameter selections based on the vector indexing, and regenerate at least one further ratio parameter from the ratio parameter selections.
[0070] An apparatus adapted to perform converting integer values to ratio parameter selections based on vector indexing may be adapted to generate a single value by appending an element from the ratio parameter selection, and to generate an index from the single number by executing an iteration loop from the zeroth iteration up to a single iteration number, sequentially associating index values with iteration numbers of the iteration loop having valid ratio parameter selections, wherein the integer value is the most significant index value.
[0071] An apparatus adapted to perform decoding a selection of remaining ratio parameters for the frame based on differential decoding of the first set of selections of ratio parameters, or selections based on selections of previously indexed time or frequency components of the ratio parameters, may be adapted to obtain a differential indicator identifying frequency differential or time differential encoding, obtain an entropy encoding indicator identifying entropy encoding parameters, and decode the selection of remaining ratio parameters for the frame based on the differential indicator and the entropy encoding indicator.
[0072] The ratio parameter configured to identify the distribution of specific objects within the object portion of the overall audio environment may be an ISM ratio.
[0073] According to a seventh aspect, there is provided an apparatus for encoding audio object parameters, the apparatus comprising: acquiring circuitry configured to acquire, for time-frequency elements of a frame comprising more than one time element and more than one frequency element, a plurality of ratio parameters of audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for a specific time-frequency element, a distribution of specific objects within object portions of the overall audio environment; quantizing circuitry configured to quantize selections of the ratio parameters, the selections being associated with audio objects in the specific frame time-frequency element; encoding circuitry configured to encode a first set of selections of the ratio parameters based on indexing of the selections; and encoding circuitry configured to encode the remaining selections of ratio parameters for the frame based on differential encoding of the selections based on the first set of selections of the ratio parameters or on selections of previously indexed time elements or frequency elements of the ratio parameters.
[0074] According to an eighth aspect, there is provided an apparatus for decoding audio object parameters, the apparatus comprising: acquiring circuitry configured to acquire a bitstream comprising encoded ratio parameters for time-frequency elements of a frame comprising more than one time element and more than one frequency element, the ratio parameters being associated with audio objects in an audio environment, the audio environment comprising more than one audio object, and the ratio parameters being configured to identify, for the specific time-frequency elements, a distribution of specific objects within object portions of the overall audio environment; decoding circuitry configured to decode a first set of ratio parameter selections based on indexing of the selections; and decoding circuitry configured to decode remaining selections of ratio parameters for the frame based on differential decoding of the first set of ratio parameter selections or selections based on previously indexed time element or frequency element selections of the ratio parameters.
[0075] According to a ninth aspect, there is provided a computer program comprising instructions (or a computer readable medium comprising program instructions) for causing an apparatus for encoding audio object parameters to perform at least the following: obtaining, for time-frequency elements of a frame comprising more than one time element and more than one frequency element, a plurality of ratio parameters of audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for the specific time-frequency element, a distribution of specific objects within object portions of the overall audio environment; quantizing a selection of ratio parameters, the selections being associated with audio objects in the specific frame time-frequency element; encoding a first set of selections of ratio parameters based on indexing of the selections; and encoding the remaining selections of ratio parameters for the frame based on differential encoding of the selections based on the first set of selections of ratio parameters or on a selection of a previously indexed time element or frequency element of the ratio parameters.
[0076] According to a tenth aspect, there is provided a computer program comprising instructions (or a computer readable medium comprising program instructions) for causing an apparatus for decoding audio object parameters to perform at least the following: obtaining a bitstream comprising encoded ratio parameters for time-frequency elements of a frame comprising more than one time element and more than one frequency element, the ratio parameters being associated with audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured for identifying, for the specific time-frequency elements, a distribution of specific objects within object portions of the overall audio environment; decoding a first set of selections of ratio parameters based on indexing of the selections; and decoding remaining selections of ratio parameters for the frame based on differential decoding of the first set of selections of ratio parameters or selections based on previously indexed time element or frequency element selections of ratio parameters.
[0077] According to an eleventh aspect, there is provided a non-transitory computer-readable medium comprising program instructions for causing an apparatus for encoding audio object parameters to perform at least the following: obtaining, for time-frequency elements of a frame comprising more than one time element and more than one frequency element, a plurality of ratio parameters of audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for the specific time-frequency element, a distribution of specific objects within object portions of the overall audio environment; quantizing a selection of the ratio parameters, the selections being associated with audio objects in the specific frame time-frequency element; encoding a first set of selections of ratio parameters based on indexing of the selections; and encoding the remaining selections of ratio parameters for the frame based on differential encoding of the selections based on the first set of selections of ratio parameters or on selections of previously indexed time elements or frequency elements of the ratio parameters.
[0078] According to a twelfth aspect, there is provided a non-transitory computer-readable medium comprising program instructions for causing an apparatus for decoding audio object parameters to perform at least the following: obtaining a bitstream comprising encoded ratio parameters for time-frequency elements of a frame comprising more than one time element and more than one frequency element, the ratio parameters being associated with audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for the specific time-frequency elements, a distribution of specific objects within object portions of the overall audio environment; decoding a first set of selections of ratio parameters based on indexing of the selections; and decoding remaining selections of ratio parameters for the frame based on differential decoding of the first set of selections of ratio parameters or selections based on previously indexed time element or frequency element selections of ratio parameters.
[0079] According to a thirteenth aspect, there is provided an apparatus for encoding audio object parameters, the apparatus comprising: means for obtaining, for a time-frequency element of a frame comprising more than one time element and more than one frequency element, a plurality of ratio parameters of audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for a specific time-frequency element, a distribution of a specific object within an object portion of the overall audio environment; means for quantizing a selection of the ratio parameters, the selection being associated with an audio object in the specific frame time-frequency element; means for encoding a first set of selections of the ratio parameters based on indexing of the selections; and means for encoding the remaining selections of ratio parameters of the frame based on differential encoding of the selections based on the first set of selections of the ratio parameters or on selections of previously indexed time elements or frequency elements of the ratio parameters.
[0080] According to a fourteenth aspect, there is provided an apparatus for decoding audio object parameters, the apparatus comprising: means for obtaining a bitstream comprising encoded ratio parameters for time-frequency elements of a frame comprising more than one time element and more than one frequency element, the ratio parameters being associated with audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for a specific time-frequency element, a distribution of specific objects within object portions of the overall audio environment; means for decoding a first set of ratio parameter selections based on indexing of the selections; and means for decoding remaining selections of ratio parameters for the frame based on differential decoding of the first set of ratio parameter selections or selections based on previously indexed time element or frequency element selections of the ratio parameters.
[0081] According to a fifteenth aspect, there is provided a computer-readable medium comprising program instructions for causing an apparatus for encoding audio object parameters to perform at least the following: obtaining, for time-frequency elements of a frame comprising more than one time element and more than one frequency element, a plurality of ratio parameters of audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for the specific time-frequency element, a distribution of specific objects within object portions of the overall audio environment; quantizing a selection of ratio parameters, the selections being associated with audio objects in the specific frame time-frequency element; encoding a first set of selections of ratio parameters based on indexing of the selections; and encoding the remaining selections of ratio parameters for the frame based on differential encoding of the selections based on the first set of selections of ratio parameters or on selections of previously indexed time elements or frequency elements of the ratio parameters.
[0082] According to a sixteenth aspect, there is provided a computer-readable medium comprising program instructions for causing an apparatus for decoding audio object parameters to perform at least the following: obtaining a bitstream comprising encoded ratio parameters for time-frequency elements of a frame comprising more than one time element and more than one frequency element, the ratio parameters being associated with audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured for identifying, for the specific time-frequency elements, a distribution of specific objects within object portions of the overall audio environment; decoding a first set of selections of ratio parameters based on indexing of the selections; and decoding remaining selections of ratio parameters for the frame based on differential decoding of the first set of selections of ratio parameters or selections based on previously indexed time element or frequency element selections of ratio parameters.
[0083] According to a seventeenth aspect, there is provided an apparatus for encoding an audio signal, the apparatus comprising: means for obtaining a plurality of audio object audio signals; obtaining a spatial audio signal; determining an available bit rate for encoding the plurality of audio object audio signals and the spatial audio signal; selecting an encoding mode based on the available bit rate; and encoding the plurality of audio object audio signals and the spatial audio signal based on the encoding mode.
[0084] The encoding modes may include a first encoding mode in which metadata of at least one transport audio signal and associated spatial audio signals is encoded.
[0085] The first encoding mode may be selected when the available bitrate is below a first bitrate threshold.
[0086] Further, the means may be for generating at least one transport audio signal by combining the plurality of audio object audio signals and at least one audio signal from the spatial audio signal.
[0087] The encoding modes may include a second encoding mode in which at least one transport audio signal, associated spatial audio metadata, associated audio object metadata, a ratio parameter configured to identify a distribution of specific audio objects within audio object portions of the overall audio environment, and a further ratio parameter configured to identify a distribution of audio object portions of the overall audio environment are encoded.
[0088] The second encoding mode may be selected when the available bitrate is below a second bitrate threshold, the second bitrate threshold being greater than the first bitrate threshold.
[0089] Further, the means may be for generating at least one transport audio signal by combining the plurality of audio object audio signals and at least one audio signal from the spatial audio signal.
[0090] The encoding modes may include a third encoding mode in which at least one transport audio signal, a selected single audio object audio signal, associated spatial audio metadata, associated audio object metadata, an object identifier for identifying the selected single audio object audio signal from the plurality of audio object audio signals, a ratio parameter configured to identify a distribution of specific audio objects within object portions of the overall audio environment, and a further ratio parameter configured to identify a distribution of object portions of the overall audio environment are encoded.
[0091] The third encoding mode may be selected when the available bitrate is below a third bitrate threshold, the third bitrate threshold being greater than the second bitrate threshold.
[0092] Further, the means may be for selecting one of the plurality of audio objects, generating a selected single object audio signal based on an audio object audio signal from the selected one of the plurality of audio objects, and generating at least one transport audio signal by combining the remaining portion of the plurality of audio object audio signals and the spatial audio signal.
[0093] Further, the means may be for analyzing the audio object audio signals and the spatial audio signal to determine a ratio parameter configured to identify a distribution of specific audio objects within the object portion of the overall audio environment.
[0094] Further means may be for analysing the audio object audio signal and the spatial audio signal to determine a further proportion parameter configured to identify a distribution of object parts across the audio environment.
[0095] The encoding modes may include a fourth encoding mode in which a plurality of audio object audio signals, a transport audio signal based on the spatial audio signal, associated spatial audio metadata, and associated object metadata are encoded separately.
[0096] The fourth encoding mode may be selected when the available bitrate is above a third bitrate threshold.
[0097] The spatial audio signal may include one of a multi-channel audio signal, a MASA audio signal, a single-channel audio signal, a stereo audio signal, and a parametric spatial audio signal.
[0098] The spatial audio signal may include associated spatial audio metadata, which may include at least one of a direction parameter, an energy ratio parameter, a surround coherence parameter, a spread coherence parameter, and several direction and distance parameters.
[0099] According to an eighteenth aspect, there is provided a method for encoding an audio signal, the method comprising: obtaining a plurality of audio object audio signals; obtaining a spatial audio signal; determining an available bitrate for encoding the plurality of audio object audio signals and the spatial audio signal; selecting an encoding mode based on the available bitrate; and encoding the plurality of audio object audio signals and the spatial audio signal based on the encoding mode.
[0100] The encoding modes may include a first encoding mode in which metadata of at least one transport audio signal and associated spatial audio signals is encoded.
[0101] The first encoding mode may be selected when the available bitrate is below a first bitrate threshold.
[0102] The method may further include generating at least one transport audio signal by combining the plurality of audio object audio signals and at least one audio signal from the spatial audio signal.
[0103] The encoding modes may include a second encoding mode in which at least one transport audio signal, associated spatial audio metadata, associated audio object metadata, a ratio parameter configured to identify a distribution of specific audio objects within audio object portions of the overall audio environment, and a further ratio parameter configured to identify a distribution of audio object portions of the overall audio environment are encoded.
[0104] The second encoding mode may be selected when the available bitrate is below a second bitrate threshold, the second bitrate threshold being greater than the first bitrate threshold.
[0105] The method may further include generating at least one transport audio signal by combining the plurality of audio object audio signals and at least one audio signal from the spatial audio signal.
[0106] The encoding modes may include a third encoding mode in which at least one transport audio signal, a selected single audio object audio signal, associated spatial audio metadata, associated audio object metadata, an object identifier for identifying the selected single audio object audio signal from the plurality of audio object audio signals, a ratio parameter configured to identify a distribution of specific audio objects within object portions of the overall audio environment, and a further ratio parameter configured to identify a distribution of object portions of the overall audio environment are encoded.
[0107] The third encoding mode may be selected when the available bitrate is below a third bitrate threshold, the third bitrate threshold being greater than the second bitrate threshold.
[0108] The method may further include selecting one of the plurality of audio objects, generating a selected single object audio signal based on an audio object audio signal from the selected one of the plurality of audio objects, and generating at least one transport audio signal by combining a remainder of the plurality of audio object audio signals and the spatial audio signal.
[0109] The method may further include analyzing the audio object audio signals and the spatial audio signal to determine a ratio parameter configured to identify a distribution of specific audio objects within the object portion of the overall audio environment.
[0110] Further, the method may include analysing the audio object audio signal and the spatial audio signal to determine a further ratio parameter configured to identify a distribution of object parts across the audio environment.
[0111] The encoding modes may include a fourth encoding mode in which a plurality of audio object audio signals, a transport audio signal based on the spatial audio signal, associated spatial audio metadata, and associated object metadata are encoded separately.
[0112] The fourth encoding mode may be selected when the available bitrate is above a third bitrate threshold.
[0113] The spatial audio signal may include one of a multi-channel audio signal, a MASA audio signal, a single-channel audio signal, a stereo audio signal, and a parametric spatial audio signal.
[0114] The spatial audio signal may include associated spatial audio metadata, which may include at least one of a direction parameter, an energy ratio parameter, a surround coherence parameter, a spread coherence parameter, and several direction and distance parameters.
[0115] According to a nineteenth aspect, there is provided an apparatus for encoding audio signals, the apparatus including at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system to at least obtain a plurality of audio object audio signals; obtain a spatial audio signal; determine an available bitrate for encoding the plurality of audio object audio signals and the spatial audio signal; select an encoding mode based on the available bitrate; and encode the plurality of audio object audio signals and the spatial audio signal based on the encoding mode.
[0116] The encoding modes may include a first encoding mode in which metadata of at least one transport audio signal and associated spatial audio signals is encoded.
[0117] The first encoding mode may be selected when the available bitrate is below a first bitrate threshold.
[0118] The apparatus may further be adapted to generate at least one transport audio signal by combining the plurality of audio object audio signals and at least one audio signal from the spatial audio signal.
[0119] The encoding modes may include a second encoding mode in which at least one transport audio signal, associated spatial audio metadata, associated audio object metadata, a ratio parameter configured to identify a distribution of specific audio objects within audio object portions of the overall audio environment, and a further ratio parameter configured to identify a distribution of audio object portions of the overall audio environment are encoded.
[0120] The second encoding mode may be selected when the available bitrate is below a second bitrate threshold, the second bitrate threshold being greater than the first bitrate threshold.
[0121] The apparatus may further be adapted to generate at least one transport audio signal by combining the plurality of audio object audio signals and at least one audio signal from the spatial audio signal.
[0122] The encoding modes may include a third encoding mode in which at least one transport audio signal, a selected single audio object audio signal, associated spatial audio metadata, associated audio object metadata, an object identifier for identifying the selected single audio object audio signal from the plurality of audio object audio signals, a ratio parameter configured to identify a distribution of specific audio objects within object portions of the overall audio environment, and a further ratio parameter configured to identify a distribution of object portions of the overall audio environment are encoded.
[0123] The third encoding mode may be selected when the available bitrate is below a third bitrate threshold, the third bitrate threshold being greater than the second bitrate threshold.
[0124] The apparatus may further be configured to select one of the plurality of audio objects, generate a selected single object audio signal based on an audio object audio signal from the selected one of the plurality of audio objects, and generate at least one transport audio signal by combining the remaining portions of the plurality of audio object audio signals and the spatial audio signal.
[0125] The device may further be configured to perform an analysis of the audio object audio signals and the spatial audio signal to determine a ratio parameter configured to identify a distribution of specific audio objects within the object portion of the overall audio environment.
[0126] Furthermore, the device may be adapted to perform an analysis of the audio object audio signal and the spatial audio signal to determine a further proportion parameter configured to identify a distribution of object parts across the audio environment.
[0127] The encoding modes may include a fourth encoding mode in which a plurality of audio object audio signals, a transport audio signal based on the spatial audio signal, associated spatial audio metadata, and associated object metadata are encoded separately.
[0128] The fourth encoding mode may be selected when the available bitrate is above a third bitrate threshold.
[0129] The spatial audio signal may include one of a multi-channel audio signal, a MASA audio signal, a single-channel audio signal, a stereo audio signal, and a parametric spatial audio signal.
[0130] The spatial audio signal may include associated spatial audio metadata, which may include at least one of a direction parameter, an energy ratio parameter, a surround coherence parameter, a spread coherence parameter, several directions, and a distance parameter. According to a twentieth aspect, there is provided an apparatus for encoding an audio signal, the apparatus comprising: means for obtaining a plurality of audio object audio signals; means for obtaining a spatial audio signal; means for determining an available bitrate for encoding the plurality of audio object audio signals and the spatial audio signal; means for selecting an encoding mode based on the available bitrate; and means for encoding the plurality of audio object audio signals and the spatial audio signal based on the encoding mode.
[0131] According to a twenty-first aspect, there is provided an apparatus for encoding an audio signal, the apparatus comprising: acquiring a circuit configured to acquire a plurality of audio object audio signals; acquiring a circuit configured to acquire a spatial audio signal; determining a circuit configured to determine an available bit rate for encoding the plurality of audio object audio signals and the spatial audio signal; selecting a circuit configured to select an encoding mode based on the available bit rate; and encoding a circuit configured to encode the plurality of audio object audio signals and the spatial audio signal based on the encoding mode.
[0132] According to a twenty-second aspect, there is provided a computer program (or a computer readable medium containing instructions) comprising instructions for causing an apparatus for encoding an audio signal to perform at least the following: obtain a plurality of audio object audio signals; obtain a spatial audio signal; determine an available bitrate for encoding the plurality of audio object audio signals and the spatial audio signal; select an encoding mode based on the available bitrate; and encode the plurality of audio object audio signals and the spatial audio signal based on the encoding mode.
[0133] According to a twenty-third aspect, there is provided a non-transitory computer-readable medium comprising program instructions for causing an apparatus for encoding an audio signal to perform at least the following: acquiring a plurality of audio object audio signals; acquiring a spatial audio signal; determining an available bitrate for encoding the plurality of audio object audio signals and the spatial audio signal; selecting an encoding mode based on the available bitrate; and encoding the plurality of audio object audio signals and the spatial audio signal based on the encoding mode.
[0134] An apparatus comprising means for performing the actions of the method as described above.
[0135] An apparatus configured to perform the actions of the method as described above.
[0136] A computer program comprising program instructions for causing a computer to carry out the method as described above.
[0137] A computer program product stored on the medium may cause an apparatus to perform the methods described herein.
[0138] The electronic device may include the apparatus described herein.
[0139] The chipset may include the devices described herein.
[0140] SUMMARY OF THE INVENTION Embodiments of the present application aim to address problems associated with the current state of the art.
[0141] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings, in which: [Brief explanation of the drawings]
[0142] [Figure 1] 1 illustrates a schematic diagram of a system of apparatus suitable for implementing some embodiments. [Figure 2] 2 illustrates a schematic diagram of an exemplary encoding mode selector shown in the system of the device illustrated in FIG. 1, according to some embodiments. [Figure 3] 3 illustrates a flow diagram of the operation of the exemplary encoding mode selector shown in FIG. 2, according to some embodiments. [Figure 4] 5 illustrates a flow diagram of the exemplary first, lowest, or single MASA bitrate encoding mode of operation shown in FIG. 4, according to some embodiments. [Figure 5] 5 illustrates a flow diagram of the exemplary second, lower level, or object information encoding mode of operation shown in FIG. 4, according to some embodiments. [Figure 6] 5 illustrates a flow diagram of the exemplary third, higher level, or single object encoding mode of operation shown in FIG. 4, according to some embodiments. [Figure 7] 5 illustrates a flow diagram of the operation of the exemplary fourth, top-level, or independent object and multiple input encoding mode shown in FIG. 4, according to some embodiments. [Figure 8]2A illustrates a schematic diagram of the exemplary audio object analyzer and audio object metadata encoder shown in FIG. 1 for a fourth encoding mode, according to some embodiments. [Figure 9] 9 illustrates a flow diagram of the operation of the exemplary audio object analyzer and audio object metadata encoder encoding mode selector shown in FIG. 8, according to some embodiments. [Figure 10] 9 illustrates a flow diagram of the operation of the exemplary ISM rate quantization optimizer shown in FIG. 8, according to some embodiments. [Figure 11] 9 illustrates a schematic diagram of an exemplary ISM vector index generator shown in FIG. 8, in accordance with some embodiments. [Figure 12] 12 illustrates a flow diagram of the operation of the exemplary ISM vector index generator shown in FIG. 11, according to some embodiments. [Figure 13] 1 illustrates an exemplary device suitable for implementing the apparatus shown in the previous figures. DETAILED DESCRIPTION OF THE INVENTION
[0143] In the following, suitable devices and possible mechanisms for encoding a parametric spatial audio signal including a transport audio signal and spatial metadata are described in more detail. As mentioned above, immersive audio codecs (such as 3GPP IVAS) are planned that support a large number of operating points ranging from low-bitrate operation to transparency. They are expected to support channel-based audio and scene-based audio inputs that include spatial information about the sound field and sound sources. In the following, an exemplary codec is configured to receive multiple input formats. In particular, the codec is configured to acquire or receive multi-audio signals (e.g., received from a microphone array, or as multi-channel audio format input, or Ambisonics format input) and audio object signals (which can also be referred to as independent streams with metadata—ISM format). Furthermore, in some situations, the codec is configured to process more than one input format at a time. This combined (input) format mode can, for example, enable simultaneous encoding of two different audio input formats. An example of two different audio input formats currently considered is the MASA format combined with an audio object format. Metadata-Assisted Spatial Audio (MASA) is an example of a parametric spatial audio format and representation that is suitable as an input format for IVAS.
[0144] It can be seen as an audio representation consisting of "N channels + spatial metadata". It is a scene-based audio format particularly suited for spatial audio capture on practical devices such as smartphones. The idea is to describe the sound scene in terms of the direction and, for example, energy ratio of time- and frequency-varying sound sources. Sound energy not specified (described) by direction is described as diffuse (coming from all directions).
[0145] As mentioned above, spatial metadata associated with an audio signal may include multiple parameters per time-frequency tile (e.g., multiple directions, and a direct-to-total energy ratio, spread coherence, distance, etc. associated with each direction (or directional value)). The spatial metadata may also include other parameters, or may be associated with other parameters that are considered omnidirectional (e.g., surround coherence, diffuse-to-total energy ratio, residual-to-total energy ratio, etc.), but which, when combined with the directional parameters, can be used to define characteristics of the audio scene. For example, a reasonable design choice that can produce quality output is one in which the spatial metadata includes determining one or more directions per time-frequency subframe (and a direct-to-total ratio, spread coherence, distance value, etc. associated with each direction).
[0146] A concept described in further detail herein is the definition of a multi-rate coding model that provides combined format encoding at various bit rates, allowing parametric encoding of audio object inputs, including encoding ISM energy proportion parameters configured to define the proportion of an audio scene created by each object in the audio scene created by all objects.
[0147] In the example below, for each time-frequency tile, a set of O such ISM energy ratio parameter values is shown, where O is the number of objects in the scene. Because there can be a significant number of such values in a frame (e.g., 20×O), efficient encoding of the values provided by embodiments herein can result in significant savings in bandwidth and bitrate.
[0148] As mentioned above, the parametric spatial metadata representation can use multiple parallel spatial directions. For MASA, the maximum number of parallel directions proposed is two. For each parallel direction, there may be associated parameters such as a direction index, a direct-to-sum ratio, a spread coherence, and a distance. In some embodiments, other parameters such as a diffuse-to-sum energy ratio, a surround coherence, and a residue-to-sum energy ratio are defined.
[0149] In this regard, Fig. 1 shows an exemplary apparatus 100 and system for implementing embodiments of the present application. The system is shown with an "analysis" part, which is the part from receiving the multi-channel signal to encoding the metadata and the downmix signal.
[0150] The input to the system "analysis" portion is a multi-channel audio signal 102. In the following example, microphone channel signal input is described, but in other embodiments, any suitable input (or synthesized multi-channel) format may be implemented. For example, in some embodiments, the spatial analyzer and spatial analysis may be implemented external to the encoder. For example, in some embodiments, spatial (MASA) metadata associated with the audio signal may be provided to the encoder as a separate bitstream. In some embodiments, the spatial (MASA) metadata may be provided as a set of spatial (directional) index values.
[0151] 1 also shows a number of audio objects 104 as further inputs to the analysis portion. As mentioned above, these multiple audio objects (or audio object streams) 104 may represent various sound sources in a physical space. Each audio object may be characterized by an audio (object) signal and accompanying metadata including directional data (in the form of azimuth and altitude angle values) that indicate the location or direction of the audio object in the physical space for each audio frame.
[0152] The multi-channel signal 102 is passed to an analyzer and encoder 101 , specifically to a transport signal generator 105 and a metadata generator 103 .
[0153] In some embodiments, the metadata generator 103 is also configured to receive the multi-channel signal and analyze the signal in order to generate metadata 104 associated with the multi-channel signal and, therefore, associated with the transport signal 106. The analysis processor 103 may be configured to generate, for each time-frequency analysis interval, metadata that may include a direction parameter, an energy ratio parameter, and a coherence parameter (and, in some embodiments, a diffuseness parameter). The direction, energy ratio, and coherence parameters may, in some embodiments, be considered to be MASA spatial audio parameters (or MASA metadata). In other words, spatial audio parameters include parameters that aim to characterize the sound field created / captured by the multi-channel signal (or two or more audio signals in general).
[0154] In some embodiments, the generated parameters may differ for each frequency band. Thus, for example, in band X, all of the parameters are generated and transmitted, while in band Y, only one of the parameters is generated and transmitted, and further, in band Z, no parameters are generated or transmitted. A practical example of this may be that in some frequency bands, such as the highest band, some of the parameters are not needed for perceptual reasons. The transport signal 106 and metadata 104 may be passed to a combined encoder core 109.
[0155] In some embodiments, transport signal generator 105 is configured to receive the multi-channel signal, generate an appropriate transport signal including the determined number of channels, and output transport signal 106 (MASA transport audio signal). For example, transport signal generator 105 may be configured to generate a two-audio-channel downmix of the multi-channel signal. The determined number of channels may be any appropriate number of channels. In some embodiments, transport signal generator is configured to select or combine input audio signals for the determined number of channels in other ways, such as by beamforming techniques, and output them as a transport signal.
[0156] In some embodiments, the transport signal generator 105 is optional and the multi-channel signal is passed unprocessed to the coupled encoder core 109 in the same manner as the transport signal, in this example.
[0157] The audio objects 104 may be passed to an audio object analyzer 107 for processing. In some embodiments, the audio object analyzer 107 analyzes the object audio input stream 104 to generate an appropriate audio object transport signal and audio object metadata. For example, the audio object analyzer may be configured to generate an audio object transport signal by combining and downmixing the audio signals of the audio objects to stereo channels using amplitude panning based on the associated audio object direction. Furthermore, the audio object analyzer may also be configured to generate audio object metadata associated with the audio object input stream 104. The audio object metadata may include direction values applicable to all subbands. Thus, if four objects are present, there are four directions. In the example described herein, the direction values also apply across all of the subframes of the frame, but in some embodiments, the temporal resolution of the direction values can be different, and the direction values apply to one or more subframes of the frame. Furthermore, an energy ratio (or ISM ratio) may be determined for each object. The energy ratio (ISM ratio) defines the contribution of an object within the object portion of the overall audio environment. In the examples below, the energy ratio (or ISM ratio) is per time-frequency tile per object.
[0158] In some embodiments, the audio object analyzer 107 may be located elsewhere and the audio objects 104 input to the analyzer and encoder 101 are audio object transport signals and audio object metadata.
[0159] The analyzer and encoder 101 may include a combined encoder core 109 configured to receive a transport audio (e.g., downmix) signal 106 and an audio object transport signal 128 and generate appropriate encodings of these audio signals.
[0160] The analyzer and encoder 101 may also include an audio object metadata encoder 111, which is similarly configured to receive the audio object metadata 108 and output an encoded or compressed form of the input information as encoded audio object metadata 112.
[0161] In some embodiments, the combined encoder core 109 may be configured to implement a stream separation metadata determiner and encoder that may be configured to determine the relative contributions of the multi-channel signal 102 (sometimes known as a MASA audio signal) and the audio objects 104 to an overall audio scene. While the following examples describe the combination of a multi-channel audio signal and audio objects, in some embodiments, the multi-channel audio signal may be generalized as a spatial audio signal. This measure of proportionality generated by the stream separation metadata determiner and encoder may be used to determine the proportion of quantization and encoding “effort” spent on the input multi-channel signal 102 and audio objects 104. In other words, the stream separation metadata determiner and encoder may generate a metric that quantifies the proportion of encoding effort spent on the multi-channel audio signal 102 compared to the encoding effort spent on the audio objects 104. This metric may be used to drive the encoding of the audio object metadata 108 and the metadata 104. Furthermore, the metrics determined by the separation metadata determiner and encoder may also be used as influencing factors in the process of encoding the transport audio signal 106 and the audio object transport audio signal 128 performed by the combined encoder core 109. The output metrics from the stream separation metadata determiner and encoder may further be represented as encoded stream separation metadata and may be combined into the encoded metadata stream from the combined encoder core 109.
[0162] In some embodiments, the analyzer and encoder 101 includes a bitstream generator 113 configured to take the encoded metadata 116, the encoded transport audio signal 138, and the encoded audio object metadata 112 and generate a bitstream 118 for potential transmission or storage.
[0163] In some embodiments, the analyzer and encoder 101 includes an encoder controller 115. The encoder controller 115, in some embodiments, can control the encoding implemented by the audio object metadata encoder 111 and the coupled encoder core 109. In some embodiments, the encoder controller 115 is configured to determine a bit rate for the bit stream 118 and control the encoding based on the bit rate. In some embodiments, the encoder controller 115 is further configured to control at least one of the audio object analyzer 107, the transport signal generator 105, and the metadata generator in generating parameters.
[0164] The analyzer and encoder 101, in some embodiments, can be a computer or mobile device (executing appropriate software stored in memory and at least one processor), or alternatively, a specialized device utilizing, for example, an FPGA or ASIC. The encoding can be implemented using any suitable scheme. In some embodiments, the encoder 107 can further interleave, multiplex into a single data stream, or embed the encoded MASA metadata, audio object metadata, and stream separation metadata within the encoded (downmixed) transport audio signal before transmission or storage, as indicated by the dashed lines in FIG. 1. The multiplexing can be implemented using any suitable scheme.
[0165] 1, an associated decoder and renderer 109 is shown, configured to take the bitstream 118 containing the encoded metadata 116, the encoded transport audio signal 138, and the encoded audio object metadata 112, and to generate therefrom an appropriate spatial audio output signal. The decoding and processing of such audio signals is known in principle and will not be described in detail hereinafter, other than the decoding of the encoded ISM ratio metadata.
[0166] With reference to FIG. 2, the encoder controller 115 according to some embodiments is shown in further detail.
[0167] In this example, encoder controller 115 includes a bitrate determiner / monitor 201 configured to determine and / or monitor the available bitrate for the encoded audio and metadata bandwidth, which may be determined based on transmission path bandwidth estimates (and, for example, based on estimated signal strength), or based on bandwidth storage determinations, or by any suitable method, to keep the file below the required size for a determined period of time.
[0168] Additionally, the bitrate determiner / monitor 201 can be configured to control the encoding mode selector 203. The encoder controller 115 can include the encoding mode selector 203, which can be configured to select an encoding mode, for example based on a determined bandwidth or bitrate, and then control an encoder, for example the combined encoder core 109 and audio object metadata encoder 111.
[0169] With reference to Figure 3, there is shown a flow diagram of an exemplary operation of the encoder controller shown in Figure 2. In this example, there is an initial operation of receiving, obtaining, or otherwise determining the encoded parameters and the bit rate or bandwidth of the audio data, as shown in Figure 3 at step 301.
[0170] Once the available bandwidth or bit rate has been obtained, a check can then be made to determine whether the bit rate is below the first (or lowest or object minimum) threshold limit, as shown by step 303 in FIG. 3 .
[0171] If the available bandwidth or bit rate is below the first (or lowest or object minimum) threshold limit, the encoder can be controlled to encode only the transport channel and MASA metadata (also shown as Mode A), as shown by step 304 in FIG. 3.
[0172] If the available bandwidth or bit rate is above the first (or lowest or object minimum) threshold limit, a further check can be made to determine whether the bit rate is below a second (or lower or one object) threshold limit, as shown by step 305 in FIG. 3 .
[0173] If the available bandwidth or bit rate falls below a second (or lower or one object) threshold limit, the encoder can be controlled to encode the transport channel, MASA metadata, ISM metadata (all objects), MASA to total ratio, ISM ratio (also shown as Mode B), as shown by step 306 in FIG. 3 .
[0174] If the available bandwidth or bit rate is above the second (or lower or one object) threshold limit, a further check can be made to determine whether the bit rate is below a third, higher or full object threshold limit, as shown by step 307 in FIG. 3 .
[0175] If the available bandwidth or bit rate is below a third, higher, or full object threshold limit, the encoder can be controlled to encode the transport channel, MASA metadata, ISM metadata (all objects), MASA to total ratio, ISM ratio, and one object audio data with one object identifier (also shown as Mode C), as shown by step 308 in FIG. 3 .
[0176] If the available bandwidth or bit rate is above a third, higher, or full object threshold limit, the encoder can be controlled to encode the transport channel, MASA metadata, ISM metadata (all objects), and all object audio data (also shown as Mode D), as shown by step 310 in FIG. 3 .
[0177] 4-7, flow diagrams are shown illustrating each of the first (or lowest or combined) encoding mode indicated by step 304 of Figure 3, the second (or lower or object metadata) encoding mode indicated by step 306 of Figure 3, the third (or higher or one object) encoding mode indicated by step 308 of Figure 3, and the fourth (or top or all objects) encoding mode indicated by step 310 of Figure 3. The encoding modes can be summarized, for example, by the following table: [Table 1]
[0178] It will be understood that the bit rates given herein are examples and that they can be other specific values.
[0179] For example, Figure 4 shows in more detail the Mode A encoding method, the first (or lowest or combined) encoding mode indicated in step 304 of Figure 3. Thus, for very low total bit rates (e.g., 32 kbps or less), the entire encoding is implemented using the MASA representation.
[0180] Thus, for example, there are operations for receiving / obtaining object-based streams (independent streams with metadata) and multi-channel-based (MASA stream) transport audio signals and metadata, as indicated by step 401 in FIG. 4 .
[0181] Next, there is the operation of generating an object-based MASA stream from the object stream (an independent stream with metadata), as indicated by step 403 in Figure 4. In some embodiments, this object-based MASA stream can be created from the object stream using, for example, the method presented in WO2019086757A1.
[0182] The object-based MASA stream and the multi-channel-based MASA stream are then combined, as shown by step 405 in Figure 4. In some embodiments, the original MASA stream and the MASA stream created from the objects can be combined using the methods set out in GB2574238. The decoder obtains the objects and the MASA audio content in MASA format.
[0183] The combined stream is then output as indicated by step 407 in Figure 4. In such embodiments, the object audio content is present in the decoded audio scene (along with the MASA audio content), but the object cannot be edited or separated from the scene at the decoder.
[0184] Figure 5 illustrates the Mode B encoding method, the second (or lower or object metadata) encoding mode as indicated by step 306 in Figure 3. Thus, for low bit rates (e.g., between 48 kbps and 80 kbps), due to the large number of bits available, there is the possibility to parameterize the audio scene by transmitting one common audio data downmix, MASA metadata, ISM metadata, and, per time-frequency tile, an additional set of parameters indicating the amount of signal corresponding to MASA components outside the total audio scene (in other words, this can be represented or indicated by a MASA to total energy ratio) and a ratio indicating how the audio scene corresponding to the object is distributed among the ISMs (in other words, this can be represented or indicated by an ISM ratio).
[0185] Thus, there are method steps for receiving / acquiring object-based streams (independent streams with metadata) and multi-channel-based (MASA stream) transport audio signals and metadata, for example as shown by step 501 in FIG. 5.
[0186] Next, a combined MASA and object-based downmix (channel-to-element) audio signal is generated, as shown by step 503 in Figure 5. In other words, the audio content of the MASA and objects is downmixed to two channels (channel-to-element CPE).
[0187] The MASA to total ratio and ISM ratio may be determined as shown by step 505 in FIG.
[0188] The MASA-to-sum ratio and the ISM ratio can then be encoded based on any suitable encoding method. For example, the ISM ratio can be encoded using a lattice encoding method, or by entropy coding the MASA-to-sum ratio after a DCT transform (e.g., as described in WO2022 / 200666). The encoding of the MASA-to-sum ratio and the ISM ratio is illustrated by step 507 in FIG. 5.
[0189] The further MASA metadata may then be encoded based on any suitable MASA metadata encoding method, as indicated by step 509 of FIG.
[0190] The combined audio signal may then be encoded based on any suitable audio signal encoding method, as indicated by step 511 in FIG.
[0191] The encoder may then output the encoded MASA metadata, the MASA-to-total ratio, the ISM ratio, and the combined transport audio signal, as indicated by step 513 in FIG.
[0192] FIG. 6 illustrates the Mode C encoding method, the third (or higher or single object) encoding mode indicated by step 308 in FIG. 3. Thus, at medium or higher bit rates (e.g., bit rates above 96 kbps and below 160 kbps), the audio content of one object is separated and transmitted independently. Furthermore, the shaped downmix from the MASA transport channel and the remainder of the object are transmitted in MASA format, with additional parameters of the MASA-to-total energy ratio and the ISM ratio. Furthermore, ISM metadata is transmitted, along with an identifier describing which object has been separated. For each frame, it is determined which object needs to be separated. This determination may be based, for example, on the relative level of the object to the other objects (e.g., isolating the loudest object). This is described in detail in WO 2022 / 214730.
[0193] Thus, there are method steps for receiving / acquiring object-based streams (independent streams with metadata) and multi-channel-based (MASA stream) transport audio signals and metadata, for example as shown by step 601 in FIG. 6.
[0194] Next, as shown by step 603 in Figure 6, one audio object is selected, and an object identifier is generated based on the selected audio object. Furthermore, an audio signal associated with the selected audio object is encoded. Any suitable audio signal encoder may be used to encode the audio signal of the selected object. For example, an audio signal encoder the same as or similar to that used to encode the MASA audio signal(s) may be employed.
[0195] A combined MASA and remaining (or unselected) object-based transport audio signal (or downmix) is then generated, as indicated by step 605 in Figure 6. The object transport signal can be created in the same way as presented in the previous mode, Mode B, with the difference that the selected or separated objects are not included in the mix. For example, the multichannel or MASA audio signal and the (unselected) object transport signal can be summed together to generate the combined transport audio signal.
[0196] The MASA to total ratio and ISM ratio may be determined as shown by step 607 in FIG.
[0197] The object identifier, MASA metadata, MASA-to-total ratio, and ISM ratio may then be encoded based on any suitable lattice encoding or entropy encoding method, as shown by step 609 of Figure 6. The encoding of the MASA-to-total energy ratio may be implemented in the manner described in WO2022 / 200666. The encoding of the ISM ratio is described in further detail below.
[0198] The combined audio signal may then be encoded based on any suitable MASA audio signal encoding method, as indicated by step 611 of Figure 6. The encoding of the combined transport audio signal may employ any suitable transport audio signal encoding, for example, an audio signal(s) encoder of an IVAS encoder.
[0199] In other words, the separated objects are determined, separated and encoded as described in WO2022 / 214730, and for the remaining objects and MASA streams the processing works as described in WO2022 / 200666.
[0200] The encoder may then output the encoded object identifiers, MASA metadata, MASA-to-total ratio, ISM ratio, object metadata (for all objects), selected single object audio signals, and the combined transport audio signal, as indicated by step 613 in FIG. 6.
[0201] Figure 7 illustrates the Mode D encoding method, the fourth (or top-level or all-object) encoding mode indicated by step 310 in Figure 3. Thus, at higher bit rates (e.g., bit rates of 160 kbps and above), the two input audio formats, MASA and ISM, are encoded independently and transmitted in the same bitstream (in other words, using a single instance of the IVAS codec).
[0202] Thus, there are method steps for receiving / acquiring object-based streams (independent streams with metadata) and multi-channel-based (MASA stream) transport audio signals and metadata, for example as shown by step 701 in FIG. 7.
[0203] Next, as indicated by step 703 of FIG. 7, the multi-channel-based (MASA stream) transport audio signal and metadata are encoded based on any suitable MASA encoding method.
[0204] The object (independent stream with metadata) and associated metadata can be further encoded, as indicated by step 705 in Figure 7. Any suitable mono encoder can be employed to implement the encoding, such as an EVS-based mono encoder block.
[0205] The encoder can then output the independently encoded objects (independent streams with metadata) and associated metadata, as well as the independently encoded multi-channel based (MASA stream) transport audio signal and metadata, as indicated by step 707 in FIG. 7.
[0206] In the following, the generation and encoding of ISM ratio values as determined and encoded within encoding modes B and C will be explained in further detail.
[0207] 8, the audio object analyzer 107 and audio object metadata encoder 111 are shown in further detail, according to some embodiments. In some embodiments, the MASA-to-sum ratio and direction (i.e., azimuth and altitude angles for each object) are forwarded and encoded by the audio object metadata encoder 111, although the specific encoding of direction and MASA-to-sum ratio will not be described in further detail herein. For example, WO2022 / 200666 describes a suitable MASA-to-sum ratio encoding method, and PCT / EP2017 / 078948 and US11475904 describe suitable direction value encoding methods.
[0208] In some embodiments, the audio object analyzer 107 comprises an ISM ratio generator 801. The ISM ratio generator 801 is configured to generate an independent stream with metadata (ISM) ratio associated with the audio object signal (independent stream with metadata) 104.
[0209] In some embodiments, the ISM ratio can be obtained as follows:
[0210] First, the object audio signal s obj (t,i) is the time-frequency domain S obj(b,n,i), where t is the time sample index, b is the frequency bin index, n is the time frame index, and i is the object index. The time-frequency domain signal can be obtained, for example, via a short-time Fourier transform (STFT) or a complex-modulated quadrature filter bank (QMF) (or their low-delay variants).
[0211] The energy of the object is then calculated in the frequency band as follows:
number
number
[0212] In some embodiments, the temporal resolution of the ISM ratio is determined by the time-frequency domain audio signal S obj The temporal resolution of (b,n,i) may differ (i.e., the temporal resolution of the spatial metadata may differ from the temporal resolution of the time-frequency transform). In such cases, the calculation (of the energy and / or ISM ratios) may involve summation over multiple time frames of the time-frequency domain audio signal and / or energy values.
[0213] ISM ratios are numbers between 0 and 1, and correspond to the proportion of times an object is active in the audio scene created by all objects. For each object, there is one ISM ratio per frequency subband and time subframe. In the following examples, assume that a time frame contains N subframes. In these examples, if there are N=4 subframes and the frame length is 20 milliseconds, the subframe length will be 5 milliseconds (i.e., there are four subframes in one frame). In other embodiments, the frame and subframe lengths may be different. Furthermore, the frame size generated, for example, by a time-frequency transform, may be different. In these embodiments, the ISM ratio may be calculated by summing values over multiple frames (or slots, as they may be called) of the time-frequency transform.
[0214] As mentioned above, the ISM ratio is passed to the audio object metadata encoder 111 .
[0215] As noted above, in some embodiments, the audio object metadata encoder 111 is configured to encode the ISM ratio.
[0216] In some embodiments, the audio object metadata encoder 111 includes an ISM ratio vector generator 803 configured to receive ISM ratio values and generate a vector representation of the ISM ratios for the subbands and subframes. In other words, the vector describes the ISM values of all objects in a given time-frequency tile. The vector of ISM ratio values 804 can then be passed to a vector (ISM ratio) quantizer 805. A vector may also be known as an array of ISM ratio values.
[0217] In some embodiments, the audio object metadata encoder 111 includes a vector (ISM ratio) quantizer 805 configured to receive a vector of ISM ratios 804 and quantize them. In some embodiments, for each subband and temporal subframe, the ratios can be scalar quantized with nb = 3 bits. Quantization of each ratio thus returns a positive integer value between 000 and 111 in binary form (or between 0 and 7 in decimal or base 10 form). In other embodiments, quantization can be performed using any suitable number of bits. Thus, the following example shows a uniform scalar quantizer based on 3 bits per value. It could also be a non-uniform scalar quantizer. The distribution of the indexes does not affect the indexing; however, this can be taken into account by observing that, in principle, some vector indices are more likely than others. In some embodiments, a quantizer based on more than 3 bits can be employed.
[0218] By definition, for each subband and subframe, the sum over all objects is 1. For each subband and temporal subframe, the values are scalar quantized with nb=3 bits. Since the ISM ratios sum to at most 1, there is a corresponding relationship between the quantization indexes, and they will sum to at most 2^nb-1 (=7). This allows for a reduction in the number of transmitted indexes; they can be transmitted with one less object per subband. However, due to the nonlinearity of the quantization operation, reconstruction at the decoder may not be optimal with respect to the condition of summing to a constant within the index domain. Thus, in some embodiments, the quantization operation may further include a quantization optimization operation, characterized by constrained vector quantization.
[0219] Thus, in some embodiments, the quantization of the per-subband and per-subframe indices may be implemented based on the following operations: 1. For statement o=0:O-1 a. Quantize the ISM ratio rISM(o) to the nearest minimum and obtain the index idx(o) (i.e., select the one with the lower value from two adjacent possible quantization values). 2. End 3. Equation
Number
[0220] Note that in scalar quantization, since the quantization operation is always forced to take the nearest lower neighbor, index modification is only performed by increasing those values.
[0221] This quantization process ensures that the sum of the indices of the entire object is equal to K. '
[0222] The vector of indices of the quantized ISM ratio values 806 can then be passed to a quantization vector encoder 807. In some embodiments, the audio object metadata encoder 111 can include a quantization vector encoder 807. The quantization vector encoder 807 can be configured to take the vector of indices of the quantized ISM ratio values 806 and generate therefrom appropriate encoded quantized ISM ratio values 808, which can be passed to a bitstream generator 113, for example, to be included in the bitstream 118.
[0223] With reference to FIG. 9, a flow diagram summarizing the operation of the exemplary audio object analyzer 107 and the exemplary audio object metadata encoder 111 shown in FIG. 8 is shown.
[0224] The initial operation is one of receiving / obtaining an independent stream with metadata, as indicated by step 901 in FIG.
[0225] Next, as shown by step 903 in FIG. 9, the following operations are performed to generate ISM ratio values from the independent streams with metadata.
[0226] The next operation from the ISM ratios is to generate a vector from the ISM ratio values, as shown by step 905 in FIG.
[0227] Once the vectors of ISM ratio values are determined, they can be quantized to generate a quantized vector of ISM ratio values, as indicated by step 907 in FIG.
[0228] From the quantized vector, an index value is then generated that represents the encoded quantized vector, as indicated by step 909 in FIG.
[0229] The encoded ISM vector index values can then be output for inclusion in the bitstream, as indicated by step 911 in FIG.
[0230] 10, the quantization of the vector of ISM ratio values to generate a quantized vector of ISM ratio values, as indicated by step 907 in FIG. 9, is further described.
[0231] Thus, initially, a vector of ISM ratio values is received or otherwise obtained, as indicated by step 1001 in FIG.
[0232] The vector of ISM ratio values is then quantized by a quantization function such that for an object o, there is a respective quantization element of the vector that has the least closest value rISM(o), and then obtain the index idx(o) associated with the least closest value as indicated by step 1003 in FIG. 10.
[0233] Further, as indicated by step 1005 in FIG. 10, there is a manipulation of the vector to regenerate or reconstruct the ISM ratio values from the index values.
[0234] Further, based on the reconstructed ISM ratio values and the original ISM ratio values, Euclidean distortion (error) values are generated, as indicated by step 1007 in FIG.
[0235] Additionally, a sum of quantization indexes, which may be referred to as SI, is determined as indicated by step 1009 in FIG.
[0236] Next, while the sum of quantization indices SI is less than the expected index sum value K, an optimization operation or step, indicated by step 1011 in FIG. 10, is implemented. This optimization may include selecting a quantization index and incrementing the index by one quantization unit. The quantization index, in some embodiments, is selected based on a reduction in the error value (or a minimum increase in the error value). The distortion value and the sum of quantization indices SI are then updated. As will be explained, these selection, increment, and update operations are performed until SI is at the expected value of K.
[0237] Once optimized, the quantized vector of ISM ratio values is then output as indicated by step 1013 in FIG.
[0238] With reference to FIG. 11, the quantized vector (ISM ratio value) encoder 807 is shown in more detail.
[0239] The quantization vector encoder 807, in some embodiments, includes a first subframe vector component encoder 1101. The first subframe vector component encoder 1101 is configured to obtain vector quantization ISM ratio index values, which can be defined as a B×N O-dimensional integer vector of ISM ratio indices. The first subframe vector component encoder 1101 can then encode the vector of O integer values as enumeration indices for each subband of the first subframe. Encoding ISM ratio index vectors using the enumeration index encoding method is described in further detail in co-pending GB application 2217884.2.
[0240] The quantized vector encoder 807, in some embodiments, includes a subframe difference and positive index generator 1103 configured to determine a difference index for a subsequent subframe vector with respect to a previous subframe, and further to change or convert the difference index to a positive index.
[0241] Further, in some embodiments, the quantization vector encoder 807 includes a positive index (sub-frame) entropy encoder 1105 configured to apply the parameters 0 and 1 to an entropy encoding (e.g., Golomb-Rice encoding) to determine or estimate the corresponding number of bits required.
[0242] In some embodiments, the quantization vector encoder 807 includes a subband difference and positive index generator 1113 configured to determine a difference index for a subsequent subband with respect to a previous subband and to change or convert the difference index to a positive index.
[0243] In some embodiments, the quantization vector encoder 807 includes a positive index (sub-band) entropy encoder 1115 configured to apply parameters 0 and 1 to entropy encoding (e.g., Golomb-Rice encoding) to determine or estimate the corresponding number of bits required.
[0244] Further, in some embodiments, the quantization vector encoder 807 includes an entropy parameter selector (across all subbands in the current subframe) 1107, which is configured to select optimal entropy encoding (GR) parameters for use in the mode across all subband data in the current subframe.
[0245] In some embodiments, the quantized vector encoder 807 includes a code mode selector 1109 configured to select a differential code mode for the current subframe (either subband or subframe differential encoding) as the mode that provides the shortest code length for the subframe.
[0246] The encoded quantized ISM ratio values 808 may be output from the quantized vector encoder 807 .
[0247] With reference to FIG. 12, a flow diagram of the operation of the exemplary quantized vector encoder 807 shown in FIG. 11 is shown, in accordance with some embodiments.
[0248] Thus, as indicated by step 1201 in FIG. 12, a vector of indices of quantized ISM ratio values 806 is shown being received or otherwise obtained.
[0249] Next, as indicated by step 1203 in FIG. 12, the first subframe vector components are encoded (for each subband) using the encoding of the enumeration index.
[0250] Next, as indicated by step 1205 of FIG. 12, a subframe loop (for subsequent subframes) may be initialized.
[0251] Further sub-band loops may then also be initiated, as indicated by step 1207 in FIG.
[0252] Next, as shown by step 1221 in FIG. 12, for each object, the method includes calculating a differential index (relative to the previous subframe), converting the differential index to a positive index, encoding the positive index with an entropy (GR) code with parameters 0 and 1, and estimating the corresponding number of bits.
[0253] Next, as shown by step 1223 in FIG. 12, for each object, the method includes calculating differential indices (relative to the previous subband), converting the differential indices to positive indices, encoding the positive indices with an entropy (GR) code with parameters 0 and 1, and estimating the corresponding number of bits.
[0254] Once this subband loop is complete, then for each differential code mode (subband-wise differential or subframe-wise differential), the "best" GR parameters are selected to use the mode across all subband data in the current subframe, as indicated by step 1209 in FIG. 12.
[0255] Then, when the subframe loop is finished, select the differential code mode (differential to previous subframe or previous subband) for the current subframe as giving the shortest code length for the subframe, as indicated by step 1211 in FIG.
[0256] Finally, as indicated by step 1213 in FIG. 12, output the selected differential code mode entropy (GR) parameter.
[0257] This operation can be expressed as follows: 1. The first subframe ISM rate quantization data is encoded with the enumeration index a. For each subband of the first subframe i. As the enumeration index, encode a vector of O integer values that sum up to the value 2^nb-1 (=7). b.End of for statement 2. For each subframe 1 to N-1 For statements for each subband from a.0 to B-1 i. Calculate the difference index for each object relative to the previous subframe ii. Convert the differential index to a regular index iii. Positive indices are encoded with GR code with parameter 0, and the corresponding number of bits is estimated. iv. Positive indices are encoded with a GR code with parameter 1, and the corresponding number of bits is estimated. v. Calculate the difference index for each object relative to the previous sub-band (if the previous sub-band does not exist, use the data from the previous sub-frame) vi. Convert differential indexes to positive indexes vii. Positive indices are encoded with GR code with parameter 0, and the corresponding number of bits is estimated. viii. Positive indices are encoded with GR code with parameter 1, and the corresponding number of bits is estimated. b.End of for statement c. For statements for each differential code mode (subband-related / subframe-related) i. Select the optimal GR parameters for using the mode across all subband data in the current subframe d.End of for statement e. Select the differential code mode (differential to previous subframe or previous subband) for the current subframe, given the shortest code length for that subframe. 3. End of for statement
[0258] The GR parameter values tested are 0 and 1, but other or more GR parameter values may also be considered.
[0259] Thus, the bitstream for one frame contains the following ISM rate related data: - Vector index of the first subframe of all subbands - For each subframe, except the first subframe 1 bit indicating differential code mode (for previous subframe or previous subband) ○ 1 bit indicating GR order (0 or 1) GR encoded differential index
[0260] In some embodiments for the first subband, there is no subband data to look back to, so a difference is taken against the previous subframe data. The GR parameters and difference code flags are determined for each subframe, and they are valid for all subbands corresponding to that subframe.
[0261] In some embodiments, the ISM ratio is a variable such as: float ism_ratios[num_subframes][num_bands][num_objects];
[0262] Next, quantization can employ "for loops" to quantize the values: for i=1:num_subframes for j=1:num_bands quantized_ism_ratios(i,j,:)=quantize_ratios(ism_ratios(i,j,:)); end end The quantized values become variables: int quantized_ism_ratios[num_subframes][num_bands][num_objects];
[0263] In such embodiments, there is no explicit generation of vectors, but the data is passed to the quantization operation in the appropriate format.
[0264] Alternatively, in some embodiments, the selection may also be implemented to be valid for the entire subframe and determined per subband, or determined separately per subframe and subband.
[0265] In some embodiments, when processing data, there is a special case where the total ISM ratio is 0, which does not obey the constant sum constraint shown above. This case corresponds to an instance where the object has no audio signal or no audio signal at all. The information that the object has no audio signal can be inferred from the MASA to total energy ratio. If the MASA to total energy ratio of a TF tile (identified by subband and subframe) is 1, there is no need to transmit the ISM ratio of that TF tile.
[0266] Furthermore, since the MASA to total energy ratio can be forced to 1 when no audio signal is present, it is possible to infer degeneracy from the value of the MASA to total energy ratio, which is the case when all values of the ISM ratios are zero.
[0267] In some embodiments, if there is no energy in any object, the ISM ratio can be set to 1 / num_objects, which forces the sum of the ISM ratios to 1 and no special processing is required since the information is present in the corresponding encoded MASA-to-sum ratio.
[0268] The decoder can be configured to decode the ISM values using the reverse process to that described above. In that way, the decoder can be configured to obtain the ISM ratio values from the encoded vector values based on the following operations: 1. sf=1:num_subframes for statement 1.1. Decode / read ISM ratio index vectors for all subbands 1.2. Save the current subframe data to the previous subframe data 1.3.Reconstructing the ISM Ratio from the ISM Ratio Index 2. End of for statement
[0269] Decode the ISM ratio index for subframe sf: 1. If statement in the first subframe 1.1.b=1:num_subbands for statement 1.1.1. Read the index of the ISM ratio index vector for subband b 1.1.2. Decode the index (according to NC327207) into a vector of indices 1.2.End of for statement 2. else statement 2.1.Read the differential mode bits 2.2.Reading Golomb Rice Order 2.3.b=1:num_subbands for statement 2.3.1. for statement with i=1:num_objects-1 2.3.1.1. Read the GR code and decode it to a positive index 2.3.1.2. Convert the positive index to an integer corresponding to the differential index 2.3.2.End of for statement 2.4.Ending the for statement 2.5.Diff mode if statement for previous subframe 2.5.1.b=1:num_subbands for statement 2.5.1.1.for statement with i=1:num_objects-1 2.5.1.1.1. Calculate the ISM rate index for subband b as the sum of the ISM rate index of the previous subframe plus the decoded difference 2.5.1.2.End of for statement 2.5.1.3. Calculate the index corresponding to the last object so that the sum of the indices of all objects is a constant K. 2.5.2.Ending the for statement 2.6.else Statement 2.6.1. Calculate the ISM ratio index of num_objects-1 for the first subband as the sum of the ISM ratio index of the first subband of the previous subframe plus the decoded difference. 2.6.2.b=2:num_subbands for statement 2.6.2.1.for statement with i=1:num_objects-1 2.6.2.1.1. Calculate the ISM ratio index for subband b as the sum of the ISM ratio index of the previous subband plus the decoded difference 2.6.2.2.Ending the for statement 2.6.2.3. Calculate the index corresponding to the last object so that the sum of the indices of all objects is a constant K. 2.6.3.End of for statement 2.7.End of if statement
[0270] 13 is an example of an electronic device that may be used as any of the apparatus portions of the system as described above. The device may be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc. The device may be configured to implement, for example, an encoder / analyzer portion and / or a decoder portion as shown in FIG. 1, or any of the functional blocks as described above.
[0271] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. Processor 1407 can be configured to execute various program code, such as the methods described herein.
[0272] In some embodiments, device 1400 includes at least one memory 1411. In some embodiments, at least one processor 1407 is coupled to memory 1411. Memory 1411 can be any suitable storage means. In some embodiments, memory 1411 includes program code sections for storing program code implementable on processor 1407. Additionally, in some embodiments, memory 1411 can further include a storage data section for storing data, e.g., data that has been processed or to be processed according to embodiments as described herein. The implemented program code stored in the program code sections and the data stored in the storage data section can be retrieved by processor 1407 whenever needed via the memory-processor coupling.
[0273] In some embodiments, device 1400 includes a user interface 1405. User interface 1405, in some embodiments, can be coupled to a processor 1407. In some embodiments, processor 1407 can control the operation of user interface 1405 and receive input from user interface 1405. In some embodiments, user interface 1405 allows a user to input commands into device 1400, for example, via a keypad. In some embodiments, user interface 1405 allows a user to obtain information from device 1400. For example, user interface 1405 may include a display configured to display information from device 1400 to the user. User interface 1405, in some embodiments, can include a touch screen or touch interface that can both allow information to be input into device 1400 and also display information to the user of device 1400. In some embodiments, user interface 1405 may be a user interface for communication.
[0274] In some embodiments, device 1400 includes input / output port 1409. In some embodiments, input / output port 1409 includes a transceiver. In such embodiments, the transceiver is coupled to processor 1407 and can be configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means can, in some embodiments, be configured to communicate with other electronic devices or apparatuses via a wire or wired coupling.
[0275] The transceiver may communicate with the further device via any suitable known communication protocol. For example, in some embodiments, the transceiver may use a suitable radio access architecture based on Long Term Evolution Advanced (LTE-Advanced, LTE-A) or New Radio (NR) (which may also be referred to as 5G), a Universal Mobile Telecommunications System (UMTS) Radio Access Network (UTRAN or E-UTRAN), Long Term Evolution (LTE, same as E-UTRA), a 2G network (legacy network technology), a wireless local area network (WLAN or Wi-Fi), a system using Worldwide Interoperability for Microwave Access (WiMAX), Bluetooth®, Personal Communications Services (PCS), ZigBee®, Wideband Code Division Multiple Access (WCDMA), Ultra-Wideband (UWB) technology, a sensor network, a mobile ad hoc network (MANET), a Cellular Internet of Things (IoT) RAN, and an Internet Protocol Multimedia Subsystem (IMS), or any other suitable option and / or any combination thereof.
[0276] The transceiver input / output port 1409 may be configured to receive a signal.
[0277] In some embodiments, device 1400 may be employed as at least a portion of a synthesis device. Input / output port 1409 may be coupled to headphones (which may be headphones with or without head tracking) or the like and loudspeakers.
[0278] Generally, various embodiments of the present invention may be implemented in hardware or special-purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Although various aspects of the present invention may be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it will be appreciated that these blocks, apparatus, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controller or other computing device, or some combination thereof.
[0279] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that any block of logic flow in the diagrams may represent program steps or interconnected logic circuits, blocks, and functions, or combinations of program steps and logic circuits, blocks, and functions. Software may be stored on physical media, such as memory chips or blocks implemented within a processor, magnetic media, such as hard disks or floppy disks, and optical media, such as DVDs and their data variants, CDs.
[0280] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technology environment and may include, by way of non-limiting example, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), gate-level circuitry, and a processor based on a multi-core processor architecture.
[0281] Embodiments of the present invention may be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert logic-level designs into semiconductor circuit designs ready to be etched into semiconductor substrates.
[0282] Programs, such as those offered by Synopsys, Inc. of Mountain View, California, and Cadence Design of San Jose, California, automatically route conductors and position components on semiconductor chips using well-established design rules as well as a library of pre-stored design modules. Once the design of a semiconductor circuit is complete, the resulting design can be transmitted in a standardized electronic format (e.g., Opus, GDSII, etc.) to a semiconductor manufacturing facility or "fab" for fabrication.
[0283] As used in this application, the term "circuitry" may refer to one or more or all of the following: (a) Hardware-only circuit implementations (e.g., implementations using only analog and / or digital circuits); (b) For example (where applicable), a combination of the following hardware circuitry and software: (i) a combination of analog and / or digital hardware circuitry(s) and software / firmware; (ii) any portion of software-based hardware processor(s) (including digital signal processor(s)), software, and memory(s) that cooperate to cause a device, such as a mobile phone or server, to perform various functions; and Hardware circuitry(s) and / or processor(s), such as microprocessor(s) or portions of microprocessor(s), that require software (e.g., firmware) for operation, but software may not be present if not necessary for operation.
[0284] This definition of circuit applies to all uses of the term in this application, including any claims. As a further example, as used in this application, the term circuit also encompasses implementations of simply a hardware circuit or processor(s), or portions of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuit also encompasses, for example, baseband or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, other computing devices, or other network devices, where applicable to certain claim elements.
[0285] As used herein, the term "non-transitory" is not a limitation regarding the permanence of the data storage (eg, RAM vs. ROM), but rather a limitation of the medium itself (ie, tangible rather than signal).
[0286] As used herein, "at least one of: " and "at least one of " and similar expressions mean at least any one of the elements, or at least any two or more of the elements, or at least all of the elements, where a list of two or more elements is joined by "and" or "or."
[0287] The foregoing description has provided a full and informative description of exemplary embodiments of the present invention, by way of illustrative and non-limiting example. However, various modifications and adaptations may become apparent to those skilled in the art in light of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of the present invention, as defined by the appended claims.
Claims
1. 1. An apparatus for encoding audio object parameters, comprising: obtaining, for time-frequency elements of a frame comprising more than one time element and more than one frequency element, a plurality of proportion parameters of audio objects in an audio environment, the audio environment comprising more than one audio object, the proportion parameters being configured to identify, for a specific time-frequency element, a distribution of specific objects within the object portion of the overall audio environment; quantizing the ratio parameter selection, the selection being associated with an audio object within a specific frame time-frequency element; encoding a first set of ratio parameter selections based on an indexing of the selections; encoding the remaining selections of the ratio parameters for the frame based on differential encoding of the selections based on the first set of selections of the ratio parameters or based on selections of previously indexed time or frequency elements of the ratio parameters; The apparatus comprising:
2. 2. The apparatus of claim 1, wherein the selection is a vector of the ratio parameters, and the means is further for generating the vector of the ratio parameters representing the ratio parameters.
3. 3. The apparatus of claim 1, wherein the means for encoding the first set of ratio parameter selections based on indexing of the selections is for generating integer values based on indexing from the selections, the generated integer values representing the ratio parameters of the audio objects.
4. The means for generating the integer value based on the indexing from the selection of the ratio parameter comprises: generating a single value by appending elements from said ratio parameter selection; generating the index from the number by executing an iteration loop from the zeroth iteration up to and including the number of iterations of the number, and sequentially associating index values with iteration numbers of the iteration loop having valid selections of ratio parameters, the integer value being a most significant index value; The device according to claim 3, for
5. The means for quantizing the selection of the ratio parameter comprises: quantizing the ratio values within the specific selection to obtain quantization index values using minimum nearest neighbor scalar quantization; calculating a reconstructed value of said ratio parameter for said specific selection; calculating an error value based on the difference between the reconstructed ratio value and the specific selection of ratio parameter values; determining a sum of quantization index values; selecting at least one quantization index value to increment so that a sum of said quantization index values equals a sum of expected indexes; The device according to any one of claims 1 to 4,
6. The means for selecting at least one quantization index value to increment so that the sum of the quantization index values equals the sum of expected indexes comprises: selecting the at least one quantization index value to increment based on identifying a maximum decrease in the error value as the index value is incremented; or selecting the at least one quantization index value to increment based on identifying a minimum increase in the error value when the index value is incremented; 6. The apparatus of claim 5, wherein the apparatus is for one of:
7. The means for quantizing the selection of the ratio parameter comprises: determining that said element is zero for a specific choice of ratio parameter; generating a further proportion parameter configured to identify a distribution of the object parts across the audio environment, the further proportion parameter value identifying an absence of an object part contribution; The device according to any one of claims 1 to 3,
8. The means for encoding the remaining selections of the ratio parameters for the frame based on a first set of selections of the ratio parameters or on selections of previously indexed time or frequency elements of the ratio parameters comprises: For a set of ratio parameter selections for specific time elements of said frame, determining a number of bits required to entropy code the differences between the quantized frequency components for a first entropy code parameter and a second entropy code parameter; determining a number of bits required to entropy code the difference between the quantized time elements for the first entropy code parameter and the second entropy code parameter; selecting, for the specific time element, the first entropy code parameter or the second entropy code parameter based on a fewer number of bits required to code the difference within the specific time element of the frame; selecting one of the entropy codes for the difference between frequency elements or time elements for the selected first entropy code parameter or the second entropy code parameter based entropy code based on a fewer number of bits required to code the difference within the specific time element of the frame; The apparatus according to any one of claims 1 to 7, for carrying out the above.
9. 9. The apparatus of claim 8, wherein the means for differential encoding of the selection based on the first set of ratio parameter selections or selection of previously indexed time or frequency elements of ratio parameters is for encoding the selected one of the entropy codes of the difference between frequency or time elements for the selected first entropy code parameter or second entropy code parameter based entropy code.
10. The means for differentially encoding the ratio parameter selections based on the first set of ratio parameter selections or based on selections of previously indexed time or frequency elements of ratio parameters comprises: For a set of selected ratio parameters for specific frequency components of the frame, determining a number of bits required to entropy code the quantized difference between the frequency components for the first entropy code parameter and the second entropy code parameter; determining a number of bits required to entropy code the quantized difference between time elements for the first entropy code parameter and the second entropy code parameter; selecting, for the specific frequency element, the first entropy code parameter or the second entropy code parameter based on a fewer number of bits required to code the difference within the specific time element of the frame; selecting one of the entropy codes of the difference between frequency elements or time elements for the selected first entropy code parameter or the second entropy code parameter based entropy code based on a smaller number of bits required to code the difference within the specific time element of the frame for the specific frequency element; The apparatus according to any one of claims 1 to 7, for carrying out the above.
11. 11. The apparatus of claim 10, wherein the means for differential encoding of the selection based on the first set of ratio parameter selections or selection of previously indexed time or frequency elements of ratio parameters is for encoding the selected one of the entropy codes of the difference between frequency or time elements for the selected first entropy code parameter or the second entropy code parameter based entropy code.
12. The means for differentially encoding selections of the ratio parameters based on the first set of ratio parameter selections or selections of previously indexed time or frequency elements of the ratio parameters for the frame, comprises: generating an indicator indicative of the selected first entropy code parameter or the second entropy code parameter; generating an indicator of the selected one of the entropy codes of the difference between frequency elements or time elements for the selected first entropy code parameter or the second entropy code parameter based entropy code; The device according to any one of claims 8 to 11, for
13. 13. The apparatus according to claim 8, wherein the entropy code is a Golomb-Rice entropy code, the first entropy code parameter is a Golomb-Rice entropy code order of 0, and the second entropy code parameter is a Golomb-Rice entropy code order of 1.
14. 14. The apparatus of claim 8, wherein the means for differentially encoding selections of the ratio parameters based on the first set of selections of the ratio parameters or encoding selections of the remaining ratio parameters of the ratio parameters of the frame based on selections of previously indexed time elements or frequency elements of the ratio parameters is for differentially encoding selections of the ratio parameters based on selections of previously indexed time elements of the ratio parameters for which there are no selections of previously indexed frequency elements of the ratio parameters.
15. The device according to any of the preceding claims, wherein the ratio parameter configured to identify a distribution of specific objects within the object portion of the overall audio environment is an ISM ratio.
16. 16. The apparatus of claim 15, when dependent on claim 7, wherein the further ratio parameter configured to identify a distribution of the object parts across the audio environment is a MASA to total energy ratio.
17. 1. An apparatus for decoding audio object parameters, comprising: obtaining a bitstream comprising encoded ratio parameters for time-frequency elements of a frame comprising more than one time element and more than one frequency element, the ratio parameters being associated with audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for a specific time-frequency element, a distribution of specific objects within the object portion of the overall audio environment; decoding a first set of ratio parameter selections based on the indexing of the selections; decoding the remaining selections of the ratio parameters for the frame based on differential decoding of the first set of selections of the ratio parameters or the selections based on selections of previously indexed time or frequency elements of the ratio parameters; The apparatus comprising:
18. 18. The apparatus of claim 17, wherein the selection is a vector of the ratio parameters.
19. The means for decoding the first set of ratio parameter selections based on the indexing of the selections comprises: obtaining an integer value representing the encoded ratio parameter; converting the integer value to a selection of a ratio parameter based on the indexing of the vector; regenerating at least one further ratio parameter from said selection of said ratio parameters; 19. The device according to claim 17 or 18, for
20. The means for converting the integer value into a selection of the ratio parameter based on the indexing of the vector comprises: generating a single value by appending elements from said ratio parameter selection; generating the index from the number by executing an iteration loop from the zeroth iteration up to and including the number of iterations of the number, and sequentially associating index values with iteration numbers of the iteration loop having valid selections of ratio parameters, the integer value being a most significant index value; 20. The device of claim 19 for:
21. The means for decoding the remaining selections of the ratio parameters for the frame based on differential decoding of the first set of selections of the ratio parameters or the selections based on selections of previously indexed time or frequency elements of the ratio parameters comprises: obtaining a differential indicator identifying frequency differential or time differential encoding; obtaining an entropy encoding designator that identifies entropy encoding parameters; decoding the remaining selection of the ratio parameter for the frame based on the difference indicator and the entropy encoding indicator; The device according to any one of claims 17 to 20, for
22. The device according to any of claims 17 to 21, wherein the ratio parameter configured to identify a distribution of specific objects within the object portion of the overall audio environment is an ISM ratio.
23. 1. A method for encoding audio object parameters, comprising: obtaining, for time-frequency elements of a frame comprising more than one time element and more than one frequency element, a plurality of proportion parameters of audio objects in an audio environment, the audio environment comprising more than one audio object, the proportion parameters being configured to identify, for a specific time-frequency element, a distribution of specific objects within the object portion of the overall audio environment; quantizing the ratio parameter selection, the selection being associated with an audio object within a specific frame time-frequency element; encoding a first set of ratio parameter selections based on an indexing of the selections; encoding the remaining selections of the ratio parameters for the frame based on differential encoding of the selections based on the first set of selections of the ratio parameters or based on selections of previously indexed time or frequency elements of the ratio parameters; The method comprising:
24. 1. A method for decoding audio object parameters, comprising: obtaining a bitstream comprising encoded ratio parameters for time-frequency elements of a frame comprising more than one time element and more than one frequency element, the ratio parameters being associated with audio objects in an audio environment, the audio environment comprising more than one audio object, the ratio parameters being configured to identify, for a specific time-frequency element, a distribution of specific objects within the object portion of the overall audio environment; decoding a first set of ratio parameter selections based on the indexing of the selections; decoding the remaining selections of the ratio parameters for the frame based on differential decoding of the first set of selections of the ratio parameters or the selections based on selections of previously indexed time or frequency elements of the ratio parameters; The method comprising:
Citation Information
Patent Citations
Coder, method therefor and storage medium
JP2001044848A
lossless multi-channel audio codec
JP2007531012A
Adaptive Grouping of Parameters to Improve Coding Efficiency
JP2008536182A
Method and apparatus for encoding and decoding object-based audio signals.
JP2010506231A
Audio signal processing method and apparatus
JP2010529500A