Parametric Spatial Audio Coding
The method addresses the challenge of efficiently encoding and decoding spatial audio parameters by using quantized ratio parameters and lattice indexing, enhancing the quality and efficiency of spatial audio representation and transmission in immersive audio applications.
Patent Information
- Application Number
- JP2025531242
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-29
- Filing Date
- 2023-11-07
- Publication Date
- 2025-12-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing immersive audio codecs face challenges in efficiently encoding and decoding spatial audio parameters, particularly the distribution of audio objects within an audio environment, which affects the quality and efficiency of spatial audio representation and transmission.
A method and apparatus for encoding and decoding audio object parameters using quantized ratio parameters, generating vectors from these parameters, and indexing them to represent the distribution of audio objects within an audio environment, utilizing techniques such as scalar quantization and lattice indexing to reduce the required memory and improve compression efficiency.
Enhances the encoding and decoding of spatial audio parameters, improving the quality and efficiency of spatial audio representation and transmission by reducing memory requirements and optimizing compression, suitable for immersive audio applications like virtual reality and mobile devices.
Smart Images

Figure 2025540763000001_ABST
Abstract
Description
[Technical Field]
[0001] This application relates to an apparatus and method for spatial audio representation and encoding, but is not limited to audio representation for audio encoders. [Background technology]
[0002] Parametric spatial audio processing is a field of audio signal processing in which spatial aspects of sound are described using a set of parameters. For example, in parametric spatial audio capture from a microphone array, it is a typical and effective choice to estimate a set of parameters from the microphone array signal, such as the direction of sound within a frequency band or the ratio of directional to omnidirectional portions of the captured sound within a frequency band. These parameters are known to well describe the perceived spatial characteristics of the sound captured at the microphone array location. These parameters can be appropriately utilized in synthesizing spatial audio for binaural headphones, loudspeakers, or other formats such as Ambisonics.
[0003] Therefore, the directional and direct-to-total energy ratio within a frequency band is a particularly effective parameterization for spatial audio capture.
[0004] A parameter set consisting of a direction parameter within a frequency band and an energy ratio parameter within a frequency band (which indicates the directionality of the sound) can also be used as spatial metadata for an audio codec (which may also include other parameters such as surround coherence, spread coherence, number of directions, distance, etc.) For example, these parameters can be estimated from an audio signal captured with a microphone array, and a stereo or mono signal can be generated from the microphone array signal, which will then be conveyed using the spatial metadata.
[0005] Immersive audio codecs are being implemented to support a wide range of operating points, from low bitrate operation to transparency. An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, which is designed for use over communication networks such as 3GPP 4G / 5G networks, including for immersive services such as immersive voice and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general audio. Furthermore, it is expected to support channel-based and scene-based audio inputs, including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency to enable conversational services, as well as support high error robustness under various transmission conditions.
[0006] A stereo signal may be encoded with, for example, an AAC encoder, and a mono signal may be encoded with an EVS encoder. The decoder decodes the audio signal into a PCM signal and processes the sound in the frequency band (using spatial metadata) to obtain a spatial output, for example a binaural output.
[0007] The aforementioned immersive audio codecs are particularly suitable for encoding spatial audio captured from microphone arrays (e.g., microphone arrays in mobile phones, VR cameras, standalone microphone arrays), however, such encoders may also have other input types, e.g., loudspeaker signals, audio object signals, ambisonic signals. Summary of the Invention
[0008] According to a first aspect, there is provided an apparatus for encoding audio object parameters, the apparatus comprising: means for obtaining ratio parameters associated with each audio object in an audio environment, the audio environment including at least two audio objects, the ratio parameters being configured to identify a distribution of each object within an object portion of the overall audio environment; quantizing the ratio parameters for the audio objects using a first number of bits; generating a vector from a selection of the quantized ratio parameters; and generating integer values based on indexing from the vector, the generated integer values representing the ratio parameters of the at least two audio objects.
[0009] The means for generating an integer value based on indexing from a vector, wherein the generated integer value represents a ratio parameter of at least two audio objects, may be means for: generating a single number by appending elements from the vector; and generating an index from the single number by executing an iteration loop including zero iterations through a single iteration and sequentially associating index values with iteration numbers of the iteration loop having a valid vector, wherein the integer value is the highest index value reached at the end of the iteration loop.
[0010] The means for generating a single number by appending elements from the vector may further be means for converting elements from the vector to a base representation based on a first number of bits.
[0011] The means for converting elements from the vector to a base representation based on the first number of bits may be means for converting elements to one of a decimal representation when the first number of bits is three, a hexadecimal representation when the first number of bits is four, or a base 32 representation when the first number of bits is five.
[0012] The means for generating vectors from selections of quantization rate parameters may be means for generating vectors from selections of all but one of the quantization rate parameters.
[0013] The means for generating vectors from all but one selection of quantization ratio parameters may be means for generating full vectors from quantization ratio parameters of the audio objects, and for generating vectors from all but one selection of quantization ratio parameters of the audio objects.
[0014] The means for quantizing the ratio parameter for the audio object using a first number of bits may be means for scalar quantizing the ratio parameter for the audio object using the first number of bits.
[0015] The first number of bits may be three, and the integer value may be a decimal integer value.
[0016] A valid vector may be one of the following: the sum of the vector element values may be seven or less, or none of the vector elements have a value greater than seven and the sum of the vector element values may be seven or less.
[0017] According to a second aspect, there is provided an apparatus for decoding ratio parameters of an audio object, the apparatus comprising means for obtaining integer values representing the ratio parameters of the audio object, converting the integer values into vectors representing selections of quantization ratio parameters based on vector indexing, regenerating at least one further quantization ratio parameter from the vector selection of quantization ratio parameters, and dequantizing the quantization ratio parameters to obtain ratio parameters of the audio object, the ratio parameters being configured to identify a distribution of a particular object within an object portion of an overall audio environment.
[0018] The means for converting the integer values into a vector representing the selection of the quantization ratio parameter based on the vector indexing may be means for generating a single number from the integer values by executing an iteration loop including zero iterations through a single iteration and sequentially associating index values with the iteration number of the iteration loop having a valid vector, where the integer value is the highest index value, and dividing the single number into vector component values to generate the vector.
[0019] The means for regenerating the at least one further quantization ratio parameter from the vector selection of quantization ratio parameters may be means for generating the at least one further quantization ratio parameter based on the value of a sum element of the vector subtracted from the expected sum value.
[0020] The means for dequantizing may be means for scalar dequantizing the ratio parameter for the audio object using a first number of bits, the ratio parameter being configured to identify a distribution of the particular object within the object portion of the overall audio environment.
[0021] The first number of bits may be three, the expected total value may be seven, and the integer value may be a decimal integer value.
[0022] According to a third aspect, there is provided a method for an apparatus for encoding audio object parameters, the method comprising: obtaining ratio parameters associated with each audio object in an audio environment, the audio environment including at least two audio objects, the ratio parameters being configured to identify a distribution of each object within an object portion of the overall audio environment; quantizing the ratio parameters for the audio objects using a first number of bits; generating a vector from a selection of the quantized ratio parameters; and generating integer values based on indexing from the vector, the generated integer values representing the ratio parameters of the at least two audio objects.
[0023] generating an integer value based on indexing from a vector, wherein the generated integer value represents a ratio parameter of at least two audio objects; generating may include generating a single number by appending elements from the vector; and generating an index from the single number by executing an iteration loop including zero iterations to a single iteration and sequentially associating index values with iteration numbers of the iteration loop having a valid vector, wherein the integer value is the highest index value reached at the end of the iteration loop.
[0024] Generating a single number by appending elements from the vector may further include converting the elements from the vector to a base representation based on the first number of bits.
[0025] Converting elements from the vector to a base representation based on the first number of bits may include converting the elements to one of a decimal representation when the first number of bits is three, a hexadecimal representation when the first number of bits is four, or a base 32 representation when the first number of bits is five.
[0026] Generating vectors from selections of quantization rate parameters may include generating vectors from selections of all but one of the quantization rate parameters.
[0027] Generating vectors from all but one selection of the quantization ratio parameters may include generating a full vector from the quantization ratio parameters of the audio object, and generating vectors from all but one selection of the quantization ratio parameters of the audio object.
[0028] Quantizing the ratio parameter for the audio object using a first number of bits may include scalar quantizing the ratio parameter for the audio object using the first number of bits.
[0029] The first number of bits may be three, and the integer value may be a decimal integer value.
[0030] A valid vector may be one of the following: the sum of the vector element values may be seven or less, or none of the vector elements have a value greater than seven and the sum of the vector element values may be seven or less.
[0031] According to a fourth aspect, there is provided a method for an apparatus for decoding ratio parameters of audio objects, the method comprising: obtaining integer values representing ratio parameters of the audio objects; converting the integer values into vectors representing selections of quantization ratio parameters based on vector indexing; regenerating at least one further quantization ratio parameter from the vector selection of quantization ratio parameters; and dequantizing the quantization ratio parameters to obtain ratio parameters of the audio objects, the ratio parameters being configured to identify a distribution of a particular object within an object portion of an overall audio environment.
[0032] Converting the integer values into a vector representing a selection of a quantization ratio parameter based on vector indexing may include generating a single number from the integer values by executing an iteration loop including zero iterations through a single iteration and sequentially associating index values with the iteration number of the iteration loop having a valid vector, where the integer value is the highest index value, and dividing the single number into vector component values to generate the vector.
[0033] Regenerating at least one further quantization ratio parameter from the vector selection of quantization ratio parameters may include generating the at least one further quantization ratio parameter based on the value of a sum element of the vector subtracted from the expected sum value.
[0034] Dequantizing the quantized ratio parameter to obtain a ratio parameter of the audio object, the ratio parameter being configured to identify a distribution of the particular object within the object portion of the overall audio environment, the dequantizing may include scalar dequantizing the ratio parameter for the audio object using a first number of bits.
[0035] The first number of bits may be three, the expected total value may be seven, and the integer value may be a decimal integer value.
[0036] According to a fifth aspect, there is provided an apparatus for encoding audio object parameters, the apparatus including at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the system to at least: obtain ratio parameters associated with each audio object in an audio environment, the audio environment including at least two audio objects, the ratio parameters being configured to identify a distribution of each object within an object portion of the overall audio environment; quantizing the ratio parameters for the audio objects using a first number of bits; generating a vector from a selection of the quantized ratio parameters; and generating integer values based on indexing from the vector, the generated integer values representing the ratio parameters of the at least two audio objects.
[0037] The device for generating integer values based on indexing from a vector, wherein the generated integer values represent ratio parameters of at least two audio objects, may further be caused to generate a single number by appending elements from the vector, and to generate an index from the single number by executing an iteration loop including zero iterations through a single iteration and sequentially associating index values with iteration numbers of the iteration loop having a valid vector, wherein the integer value is the highest index value reached at the end of the iteration loop.
[0038] The apparatus that generates a single number by appending elements from a vector may further be caused to convert the elements from the vector to a base representation based on the first number of bits.
[0039] The apparatus for converting elements from a vector into a base representation based on a first number of bits may be configured to convert the elements into one of a decimal representation when the first number of bits is three, a hexadecimal representation when the first number of bits is four, or a base 32 representation when the first number of bits is five.
[0040] The apparatus that is caused to generate vectors from selections of quantization rate parameters may be further caused to generate vectors from selections of all but one of the quantization rate parameters.
[0041] The apparatus that generates vectors from all but one selection of quantization ratio parameters may be further configured to generate full vectors from quantization ratio parameters of audio objects, and to generate vectors from all but one selection of quantization ratio parameters of audio objects.
[0042] The apparatus that causes the quantization of the ratio parameter for the audio object using the first number of bits may also cause the apparatus that causes the scalar quantization of the ratio parameter for the audio object using the first number of bits.
[0043] The first number of bits may be three, and the integer value may be a decimal integer value.
[0044] A valid vector may be one of the following: the sum of the vector element values may be seven or less, or none of the vector elements have a value greater than seven and the sum of the vector element values may be seven or less.
[0045] According to a sixth aspect, there is provided an apparatus for decoding ratio parameters of audio objects, the apparatus including at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system to at least obtain integer values representing ratio parameters of the audio objects; convert the integer values into vectors representing selections of quantization ratio parameters based on vector indexing; regenerate at least one further quantization ratio parameter from the vector selection of quantization ratio parameters; and dequantize the quantization ratio parameters to obtain ratio parameters of the audio objects, the ratio parameters being configured to identify a distribution of a particular object within an object portion of an overall audio environment.
[0046] The apparatus for converting integer values into vectors representing selection of quantization ratio parameters based on vector indexing may be configured to generate a single number from the integer values by executing an iteration loop including zero iterations through a single iteration and sequentially associating index values with iteration numbers of the iteration loop having a valid vector, where the integer value is the highest index value, and to separate the single number into vector component values to generate the vector.
[0047] The apparatus that is caused to regenerate at least one further quantization ratio parameter from a vector selection of quantization ratio parameters may further be caused to generate the at least one further quantization ratio parameter based on the value of a sum element of the vector subtracted from the expected sum value.
[0048] The apparatus that causes the dequantization to dequantize the quantization ratio parameter to obtain a ratio parameter of the audio object, the ratio parameter being configured to identify a distribution of the particular object within the object portion of the overall audio environment, may further be caused to scalar dequantize the ratio parameter for the audio object using a first number of bits.
[0049] The first number of bits may be three, the expected total value may be seven, and the integer value may be a decimal integer value.
[0050] According to a seventh aspect, there is provided an apparatus for encoding audio object parameters, the apparatus comprising: means for obtaining a ratio parameter associated with each audio object in an audio environment, the audio environment including at least two audio objects, the ratio parameter being configured to identify a distribution of each object within an object portion of the overall audio environment; means for quantizing the ratio parameters for the audio objects using a first number of bits; means for generating a vector from a selection of the quantized ratio parameters; and means for generating integer values based on indexing from the vector, the generated integer values representing the ratio parameters of the at least two audio objects.
[0051] According to an eighth aspect, there is provided an apparatus for decoding ratio parameters of an audio object, the apparatus comprising: means for obtaining integer values representing the ratio parameters of the audio object; means for converting the integer values into vectors representing selections of quantization ratio parameters based on vector indexing; means for regenerating at least one further quantization ratio parameter from the vector selections of quantization ratio parameters; and means for dequantizing the quantization ratio parameters to obtain ratio parameters of the audio object, the ratio parameters being configured to identify a distribution of a particular object within an object portion of an overall audio environment.
[0052] According to a ninth aspect, there is provided an apparatus for encoding audio object parameters, the apparatus comprising: an acquisition circuit configured to acquire ratio parameters associated with each audio object in an audio environment, the audio environment including at least two audio objects, the ratio parameters configured to identify a distribution of each object within an object portion of the overall audio environment; a quantization circuit configured to quantize the ratio parameters for the audio objects using a first number of bits; a vector generation circuit configured to generate a vector from a selection of the quantization ratio parameters; and an integer value generation circuit configured to generate integer values based on indexing from the vector, the generated integer values representing the ratio parameters of the at least two audio objects.
[0053] According to a tenth aspect, there is provided an apparatus for decoding ratio parameters of audio objects, the apparatus comprising: an obtaining circuit configured to obtain integer values representing the ratio parameters of the audio objects; a converting circuit configured to convert the integer values into vectors representing selections of quantization ratio parameters based on vector indexing; a regenerating circuit configured to regenerate at least one further quantization ratio parameter from the vector selection of quantization ratio parameters; and a dequantizing circuit configured to dequantize the quantization ratio parameters to obtain ratio parameters of the audio objects, the ratio parameters being configured to identify a distribution of a particular object within an object portion of an overall audio environment.
[0054] According to an eleventh aspect, there is provided a computer program (or a computer-readable medium containing instructions) comprising instructions for providing an apparatus for encoding audio object parameters, causing the apparatus to at least: obtain ratio parameters associated with each audio object in an audio environment, the audio environment comprising at least two audio objects, the ratio parameters being configured to identify a distribution of each object within an object portion of the overall audio environment; quantize the ratio parameters for the audio objects using a first number of bits; generate a vector from a selection of the quantized ratio parameters; and generate integer values based on indexing from the vector, the generated integer values representing the ratio parameters of the at least two audio objects.
[0055] According to a twelfth aspect, there is provided a computer program comprising instructions (or a computer readable medium comprising instructions) for providing an apparatus for decoding ratio parameters of audio objects, causing the apparatus to at least: obtain integer values representing ratio parameters of the audio objects; convert the integer values into vectors representing selections of quantization ratio parameters based on vector indexing; regenerate at least one further quantization ratio parameter from the vector selection of quantization ratio parameters; and dequantize the quantization ratio parameters to obtain ratio parameters of the audio objects, the ratio parameters being configured to identify a distribution of a particular object within an object portion of an overall audio environment.
[0056] According to a thirteenth aspect, there is provided a non-transitory computer-readable medium comprising program instructions for providing an apparatus for encoding audio object parameters, causing the apparatus to at least: obtain ratio parameters associated with each audio object in an audio environment, the audio environment including at least two audio objects, the ratio parameters being configured to identify a distribution of each object within an object portion of the overall audio environment; quantizing the ratio parameters for the audio objects using a first number of bits; generating a vector from a selection of the quantized ratio parameters; and generating integer values based on indexing from the vector, the generated integer values representing the ratio parameters of the at least two audio objects.
[0057] According to a fourteenth aspect, there is provided a non-transitory computer-readable medium comprising program instructions for providing an apparatus for decoding ratio parameters of audio objects, causing the apparatus to at least: obtain integer values representing ratio parameters of the audio objects; convert the integer values into vectors representing selections of quantization ratio parameters based on vector indexing; regenerate at least one further quantization ratio parameter from the vector selection of quantization ratio parameters; and dequantize the quantization ratio parameters to obtain ratio parameters of the audio objects, the ratio parameters being configured to identify a distribution of a particular object within an object portion of an overall audio environment.
[0058] An apparatus comprising means for performing the operations of the above method.
[0059] An apparatus configured to perform the operations of the above method.
[0060] A computer program comprising program instructions for causing a computer to carry out the above method.
[0061] A computer program product stored on the medium can cause the apparatus to perform the methods described herein.
[0062] The electronic device may include the apparatus described herein.
[0063] The chipset may include the devices described herein.
[0064] SUMMARY OF THE INVENTION Embodiments of the present application aim to address problems with the state of the art.
[0065] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings in which: [Brief explanation of the drawings]
[0066] [Figure 1] 1 illustrates a schematic diagram of a system of apparatus suitable for implementing some embodiments. [Figure 2] 2 illustrates a schematic diagram of an exemplary metadata extractor and metadata compressor and packer shown in the system of devices illustrated in FIG. 1, according to some embodiments. [Figure 3] 3 illustrates a flow diagram of the operation of the exemplary metadata extractor and metadata compressor and packer shown in FIG. 2, according to some embodiments. [Figure 4] 3 illustrates a schematic diagram of an exemplary ISM vector index generator shown in FIG. 2, in accordance with some embodiments. [Figure 5] 5 illustrates a flow diagram of the operation of the exemplary ISM vector index generator shown in FIG. 4, according to some embodiments. [Figure 6] 2 illustrates a schematic diagram of an exemplary metadata decoder shown in the system of devices illustrated in FIG. 1, according to some embodiments. [Figure 7] 7 illustrates a flow diagram of the operation of the exemplary metadata decoder shown in FIG. 6, according to some embodiments. [Figure 8] 7 illustrates a schematic flow diagram of the operation of the exemplary ISM vector index to vector generator shown in FIG. 6, in accordance with some embodiments. [Figure 9] 1 illustrates an exemplary device suitable for implementing the apparatus shown in the previous figures. DETAILED DESCRIPTION OF THE INVENTION
[0067] The following describes in more detail an apparatus and possible mechanisms suitable for encoding a parametric spatial audio signal including a transport audio signal and spatial metadata. In the following, a 3GPP IVAS codec is configured to receive a mixed input format mode. The mixed input format mode enables simultaneous encoding of two different audio input formats. An example of two different audio input formats currently under consideration is the combination of the MASA format and audio objects. Audio object data may also be known as an independent stream with metadata (ISM), and are referred to interchangeably herein. In mixed format encoding, a parameterized ISM ratio is used to describe the distribution of ISM-related audio content with respect to an object. Specifically, the ISM ratio identifies the distribution of a particular object within the object portion of the entire audio scene. Furthermore, there may be a parameter called the MASA-to-total energy ratio, which identifies the portion of the MASA stream within the entire audio scene (including the object and MASA). Thus, (1-MASA-to-total energy ratio) identifies the portion of all objects within the entire audio scene.
[0068] These parameters are transmitted to the decoder.
[0069] The next concept described in detail herein is the efficient encoding of these ISM ratios. These ISM ratios may be indexed within a pyramidal truncation of the Zn lattice, or may be encoded by a suitable entropy encoder (such as a Golomb-Rice coder) or a context arithmetic encoder. Encoding using a pyramidal truncation of the Zn lattice is more efficient in terms of compression efficiency, but requires memory to store the Zn lattice index offsets and vector descriptions. Arithmetic encoding methods are generally less efficient because there is usually not enough data in an audio frame to determine the distribution of index values.
[0070] The embodiments described herein seek to provide an indexing method for lattice Zn vectors that does not require storing index offsets or information related to layer values. Embodiments employing such a method are efficient for low-dimensional lattices and can be used to encode ISM ratio index vectors.
[0071] Metadata-Assisted Spatial Audio (MASA) is an example of a parametric spatial audio format and representation that is suitable as an input format for IVAS.
[0072] The format can be seen as an audio representation consisting of "N channels + spatial metadata". It is a scene-based audio format particularly suited for spatial audio capture on practical devices such as smartphones. Its aim is to describe the acoustic scene in terms of the direction of sound sources and, for example, energy proportions, which vary with time and frequency. Acoustic energy that is not defined (described) by direction is described as diffuse (coming from all directions).
[0073] As mentioned above, spatial metadata associated with an audio signal may include multiple parameters per time-frequency tile (multiple parameters, such as multiple directions, and multiple parameters related to each direction (or directivity value), direct-to-total energy ratio, spread coherence, distance, etc.). Although the spatial metadata may also include other parameters, or the spatial metadata may be associated with other parameters that are considered omnidirectional (surround coherence, diffuse-to-total energy ratio, residual-to-total energy ratio, etc.), the spatial metadata, when combined with the directional parameters, can be used to define characteristics of the audio scene. For example, a reasonable design choice that can produce high-quality output is one in which the spatial metadata includes one or more directions per time-frequency subframe (and associated with a direct-to-total ratio, spread coherence, distance value, etc. for each direction).
[0074] As mentioned above, the parametric spatial metadata representation can use multiple simultaneous spatial directions. For MASA, the maximum number of simultaneous directions proposed is two. For each simultaneous direction, there may be associated parameters such as direction index, direct-to-total ratio, spread coherence, and distance. In some embodiments, other parameters are defined, such as diffuse-to-total energy ratio, surround coherence, and residual-to-total energy ratio.
[0075] To have sufficient frequency and time resolution (e.g., five frequency bands and 20 millisecond time resolution), often only a few bits per value (e.g., direction parameter) can be used. In practice, this means that the quantization steps are relatively large. Thus, for example, in a particular time-frequency tile, the quantization points are 0 degrees, ±45 degrees, ±90 degrees, ±135 degrees, and 180 degrees in orientation.
[0076] Audio object input formats can contain independent streams with metadata (ISM). In some cases, metadata may not be available (in which case, for example, some default values may be assumed). Within the encoding of the composite format, a parameter called ISM ratio is defined, which identifies the distribution of a particular object within the object part of the entire audio scene. The concept discussed herein is the efficient encoding and decoding of these ISM ratio parameters.
[0077] In this regard, Fig. 1 shows an exemplary apparatus 100 and system for implementing embodiments of the present application. In this regard, Fig. 1 shows an exemplary apparatus and system for implementing embodiments of the present application. A system is shown with an "analysis" part. The "analysis" part is the part from receiving the multi-channel signal to encoding the metadata and the downmix signal.
[0078] The input to the "analysis" portion of the system is a multi-channel audio signal 102. In the examples below, microphone channel signal inputs are described, but in other embodiments, any suitable input (or synthesized multi-channel) format may be implemented. For example, in some embodiments, the spatial analyzer and spatial analysis may be implemented external to the encoder. For example, in some embodiments, spatial (MASA) metadata associated with the audio signal may be provided to the encoder as a separate bitstream. In some embodiments, the spatial (MASA) metadata may be provided as a set of spatial (directional) index values.
[0079] 1 also shows a number of audio objects 104 as further inputs to the analyzer. As noted above, these multiple audio objects (or audio object streams) 104 may represent various sound sources within a physical space. Each audio object may be characterized by an audio (object) signal and associated metadata including directional data (in the form of azimuth and altitude values) that indicate the location or direction of the audio object within the physical space on an audio frame-by-audio frame basis.
[0080] The multi-channel signal 102 is passed to an analyzer and encoder 101 , and in particular to a transport signal generator 105 and a metadata generator 103 .
[0081] In some embodiments, the metadata generator 103 is also configured to receive the multi-channel signal and analyze it to generate metadata 104 associated with the multi-channel signal and, in turn, with the transport signal 106. The analysis processor 103 may be configured to generate, for each time-frequency analysis interval, metadata that may include a direction parameter, an energy ratio parameter, and a coherence parameter (and, in some embodiments, a diffuseness parameter). The direction, energy ratio, and coherence parameters may, in some embodiments, be considered to be MASA spatial audio parameters (or MASA metadata). In other words, spatial audio parameters include parameters that aim to characterize the sound field created / captured by the multi-channel signal (or two or more audio signals in general).
[0082] In some embodiments, the generated parameters may differ for each frequency band. Thus, for example, in band X, all of the parameters may be generated and transmitted, while in band Y, only one of the parameters may be generated and transmitted, and in band Z, no parameters may be generated or transmitted. A practical example of this may be that in some frequency bands, such as the highest band, some of the parameters may not be needed for perceptual reasons. The transport signal 106 and metadata 104 may be passed to a composite encoder core 109.
[0083] In some embodiments, the transport signal generator 105 is configured to receive a multi-channel signal, generate an appropriate transport signal including a determined number of channels, and output the transport signal 106 (MASA transport audio signal). For example, the transport signal generator 105 may be configured to generate a two-audio-channel downmix of the multi-channel signal. The determined number of channels may be any suitable number of channels. The transport signal generator of some embodiments is configured to otherwise select or combine the input audio signal into the determined number of channels, for example by beamforming techniques, and output these as a transport signal.
[0084] In some embodiments, the transport signal generator 105 is optional and the multi-channel signal is passed to the composite encoder core 109 without processing, just like the transport signal in this example.
[0085] The audio objects 104 may be passed to an audio object analyzer 107 for processing. In some embodiments, the audio object analyzer 107 analyzes the object audio input stream 104 to generate an appropriate audio object transport signal and audio object metadata. For example, the audio object analyzer may be configured to generate an audio object transport signal by downmixing the audio signals of the audio objects to stereo channels using amplitude panning based on the direction of the associated audio object. Furthermore, the audio object analyzer may also be configured to generate audio object metadata associated with the audio object input stream 104. The audio object metadata may include direction values applicable to all subbands. Thus, if there are four objects, there are four directions. In the example described herein, the direction values also apply across all subframes of a frame, but in some embodiments, the time resolution of the direction values may be different, and the direction values apply to one or more subframes of a frame. Furthermore, the audio object metadata may include energy ratios (or ISM ratios). In the following example, the energy ratios (or ISM ratios) are for each time-frequency tile of each object.
[0086] In some embodiments, the audio object analyzer 107 may be located elsewhere and the audio objects 104 input to the analyzer and encoder 101 are audio object transport signals and audio object metadata.
[0087] The analyzer and encoder 101 may include an audio encoder core 109 configured to receive a transport audio (e.g., downmix) signal 106 and an audio object transport signal 128 for appropriate encoding of these audio signals. The audio encoder core 109 is further configured to receive MASA metadata 104, which is the output of the metadata generator, and to output an encoded or compressed form of the information as encoded (MASA) metadata 116.
[0088] The analyzer and encoder 101 may also include an audio object metadata encoder 111, which is similarly configured to receive the audio object metadata 108 and output an encoded or compressed form of the input information as encoded audio object metadata 112.
[0089] In some embodiments, the composite encoder core 109 can be configured to implement a stream separation metadata determiner and encoder, which can be configured to determine the relative contributions of the multi-channel signal 102 (the MASA audio signal) and the audio objects 104 to the overall audio scene. This proportional measure provided by the stream separation metadata determiner and encoder can be used to determine the proportion of quantization and encoding "effort" to be spent on the input multi-channel signal 102 and the audio objects 104. In other words, the stream separation metadata determiner and encoder can provide a metric that quantifies the proportion of the encoding effort spent on the multi-channel audio signal 102 compared to the encoding effort spent on the audio objects 104. This metric can be used to facilitate the encoding of the audio object metadata 108 and the audio object metadata 104. Furthermore, the metric determined by the separation metadata determiner and encoder can also be used as an influencing factor in the process of encoding the transport audio signal 106 and the audio object transport audio signal 128 performed by the composite encoder core 109. Additionally, output metrics from the stream separation metadata determiner and encoder may be represented as encoded stream separation metadata and combined into an encoded metadata stream from the combined encoder core 109 .
[0090] The analyzer and encoder 101, in some embodiments, may be a computer or mobile device (executing appropriate software stored in memory and at least one processor), or alternatively, a specialized device utilizing, for example, an FPGA or ASIC. The encoding may be implemented using any suitable scheme. In some embodiments, the encoder 107 may further interleave, multiplex, or embed the encoded MASA metadata, audio object metadata, and stream separation metadata into a single data stream, or within the encoded (downmixed) transport audio signal, prior to transmission or storage, as indicated by the dashed lines in FIG. 1. The multiplexing may be implemented using any suitable scheme.
[0091] 1, there is shown an associated decoder and renderer 119 that is configured to take the bitstream 118 containing the encoded metadata 116, the encoded transport audio signal 138, and the encoded audio object metadata 112, and to generate therefrom an appropriate spatial audio output signal. The decoding and processing of such audio signals is known in principle and will not be described in detail hereinafter, other than the decoding of the encoded ISM ratio metadata.
[0092] With reference to FIG. 2, the audio object analyzer 107 and audio object metadata encoder 111 are shown in further detail, according to some embodiments.
[0093] In some embodiments, the audio object analyzer 107 includes an ISM ratio generator 201. The ISM ratio generator 201 is configured to generate an independent stream with metadata (ISM) ratio associated with the audio object signal 104.
[0094] In some embodiments, the ISM ratio can be obtained as follows:
[0095] First, the object audio signal s obj (t,i) is the time-frequency domain S obj (b,n,i) where t is the time sample index, b is the frequency bin index, n is the time frame index, and i is the object index. The time-frequency domain signal can be obtained, for example, via a short-time Fourier transform (STFT) or a complex-modulated quadrature filter bank (QMF) (or their low-delay variants).
[0096] The energy of the object is then calculated in the following frequency bands:
number
number
[0097] In some embodiments, the temporal resolution of the ISM ratio is determined by the time-frequency domain audio signal S obj The temporal resolution of (b,n,i) may differ (i.e., the temporal resolution of the spatial metadata may differ from the temporal resolution of the time-frequency transform). In such cases, the calculation (of energy and / or ISM ratio) may involve summation over multiple time frames of the time-frequency domain audio signal and / or energy values.
[0098] The ISM ratios are numbers between 0 and 1, and correspond to the ratio at which an object is active within the audio scene created by all objects. For each object, there is one ISM ratio per frequency subband and time subframe. As mentioned above, the ISM ratios are passed to the Audio Object Metadata Encoder 111.
[0099] As mentioned above, in some embodiments, the audio object metadata encoder 111 is configured to encode the ISM ratio. In some embodiments, the direction (i.e., the azimuth and altitude angles for each object) is forwarded to the encoder 111 and encoded.
[0100] In some embodiments, the audio object metadata encoder 111 comprises an ISM ratio quantizer 203. The ISM ratio quantizer 203 is configured to receive the ISM ratio values 202 and quantize them. In some embodiments, for each subband and temporal subframe, the indexes can be scalar quantized with nb=3 bits. Quantizing each ratio thus returns a positive integer value between 000 and 111 in binary (or between 0 and 7 in decimal or base 10 format). In other embodiments, the quantization can be performed using any suitable number of bits. However, as mentioned above, in the following examples, a uniform scalar quantizer based on 3 bits for each value is shown. The quantizer could also be a non-uniform scalar quantizer. The distribution of the indexes does not affect the indexing. However, in principle, the distribution of the indexes could be taken into account by observing that some vector indices are more probable than others. Using 3 bits for quantization allows the following examples to use decimal numbers to represent numbers. However, quantizers based on more than 3 bits can also be used. For example, hex or base 16 representation can be used for 4-bit quantized representation, and base 32 for 5-bit.
[0101] The quantized ISM ratio index values 204 may then be passed to an ISM ratio vector generator 205 .
[0102] In some embodiments, the audio object metadata encoder 111 includes an ISM ratio vector generator 205 configured to receive the quantized ISM ratio values and generate a vector representation of the ISM ratios for the subbands. The vector to be indexed is
number
[0103] In the example below, there are three objects, and the ISM ratios are each quantized with 3 bits. Thus, the example using the above has n=3 and K=7.
[0104] Thus, the vectors generated by ISM ratio vector generator 205 in this example can be configured to determine the vectors (which will be indexed in a manner described below) as follows:
[0007] and all different permutations
[0070] ;
[0700]
[0016] and all permutations
[0061] ;
[0160] ;
[0610] ;
[0601] ;
[0106]
[0025] and all permutations
[0052] ;
[0250] ;
[0520] ;
[0502] ;
[0205]
[0115] and all the different permutations
[0151] ;
[0511]
[0034] and all permutations
[0043] ;
[0340] ;
[0430] ;
[0403] ;
[0304]
[0124] and all permutations
[0421] ;
[0241] ;
[0142] ;
[0412] ;
[0214] ...
[0105] The vector quantized ISM rate index values 206 may then be passed to an ISM vector index generator 207 .
[0106] The audio object metadata encoder 111 may, in some embodiments, include an ISM vector index generator 207. The ISM vector index generator 207 may be configured to generate appropriate index values for the subbands representing the vectors and pass them as coded ISM ratio index values 208, which may be passed to a bitstream generator 209, for inclusion in the bitstream 118, for example.
[0107] With reference to FIG. 3, a flow diagram summarizing the operation of the exemplary audio object analyzer 107 and the exemplary audio object metadata encoder 111 shown in FIG. 2 is shown.
[0108] The initial operation is one of receiving / acquiring an independent stream with metadata, as shown in FIG. 3 at step 301 .
[0109] Next, in step 303, as shown in FIG. 3, the following operation is performed to generate an ISM ratio value from the independent stream with metadata.
[0110] Once the ISM ratio values are determined, they may be quantized to generate quantized ISM ratio values as shown in FIG. 3 by step 305 .
[0111] From the quantized ISM ratios, the next operation is to generate vectors from the quantized ISM ratio values as shown in FIG. 3 by step 307 .
[0112] From that vector, an ISM vector index is then generated as shown in FIG. 3 by step 309 .
[0113] The ISM vector index may then be output for inclusion in the bitstream, as shown in FIG. 3 by step 311.
[0114] With reference to FIG. 4, the ISM vector index generator 207 is shown in further detail.
[0115] In this example, the ISM vector index generator 207 includes a vector component selector 401 configured to receive the vector quantized ISM ratio index values 206 , select vector components, and pass them to a numeric value generator 403 .
[0116] By definition, for each subband and subframe, the sum of the ISM ratios across all objects is 1. Since the ISM ratios sum to 1, there is a corresponding relationship between the quantization indexes, which sum to 2. nb -1 (=K=7 (as above)). This reduces the number of values or vector components that need to be encoded and transmitted.
[0117] Thus, in some embodiments, the vector component selector is configured to select and forward the first N-1 components of an N-length vector.
[0118] For example, in the N=3 example above, if the vector quantization ISM rate index value 206 is 0 in a subband, the selected output vector components are 0 and 4 (3 is not selected). In this example, the first N-1 components are selected, but it should be understood that any suitable criteria can be implemented for selecting the N-1 components, such as selecting the "last N-1" components. The selection or "dropping" of components can be implemented based on any suitable selection method. For example, in some embodiments, the first N-1 values are selected (or in other words, the selection always "drops" the last value, since the entire permutation space is used). In some embodiments, a histogram of the resulting indices is estimated, and the indices (corresponding to the first components) can be encoded using a variable bit rate.
[0119] Additionally, in some embodiments, a complexity reduction may be used by favoring the lowest value at the start (in a while loop).
[0120] The ISM vector index generator 207 further comprises a number generator 403 configured to receive the output of the vector component selector 401 and generate a number from the selected vector component.
[0121] In some embodiments, numeric value generator 403 is configured to generate a decimal numeric value from concatenating the decimal representations of each of the selected components. Thus, in the above example where the selected components are 0 and 4, a decimal value of 4 is generated by numeric value generator 403.
[0122] The numeric value generator can then pass the generated numeric value to the numeric value pair index generator 405 .
[0123] In some embodiments, the ISM vector index generator 207 includes a number-to-index generator 405. The mapping of the number generated by the number generator to the appropriate index is not a matter of taking the number as the index, because when numbers from 0 onwards are considered, not all numbers can be generated by the number generator (in the example where K=7, the maximum value of each digit in the number is 7), so there will be some numbers that correspond to valid vectors from the set we want to enumerate and some numbers that are not generated.
[0124] For example, the correspondence relationship for an example where N=3 and K=7 is as follows: [Table 1]
[0125] In this way, we can obtain indices for all 36 possible vectors of a pyramidal layer of norm 7 from the lattice Z3.
[0126] The same idea applies to n=2, n=4, or even more.
[0127] The indexing function can be defined by the following pseudocode implementation: Input: A vector of n positive integer values whose sum is equal to K,y. Output: Enumeration index 1. Get the first n-1 vector components of vector y 2. Form Number:
number
[0128] The function valid() verifies whether a given number corresponds to a valid (n-1)-dimensional integer array, i.e., whether its Laplacian norm is less than or equal to K. In other words, a valid vector can be either: the sum of the vector element values is less than or equal to seven, in which case, in the 3-bit quantization example, the number of valid vector loop iterations would be 7 for two objects (since only one value is checked), 70 for three objects (since one value, two values are checked), or 700 for four objects (since three values are checked). Thus, for two objects there are 8 valid vectors (0, 1, 2, 3, 7), for three objects there are 36 valid vectors (as shown in the table above), and for four objects there are 120 valid vectors.
[0129] The valid() function is defined as follows: int16_tvalid(int16_tindex,int16_tK,int16_tlen) { int16_tout; int16_ti,sum,elem; int16_tbase[4]; / *Maximum spatial dimension is assumed to be 4* / sum=0; set_s(base,1,len); / *set all values to 1* / for(i=1;i <len;i++) { base[i]=base[i-1]*10; } sum=0; for(i=len-1;i>=0;i--) { elem=index / base[i]; sum+=elem; index-=elem*base[i]; } if(sum<=K) { out=1; } else { out=0; } returnout; }
[0130] The encoded ISM rate index value 208 may then be output.
[0131] With reference to FIG. 5, a flow diagram summarizing the operation of the exemplary ISM vector index generator 207 shown in FIG. 4 is shown.
[0132] The initial operation is one of receiving / obtaining vector quantized ISM rate index values as shown in FIG. 5 by step 501.
[0133] Next, the next operation is performed as shown in FIG. 5 by step 503, which is to select vector components (eg, select all but the last component, or in other words, remove the last component).
[0134] Once the selected vector components have been determined, they can be used to generate a single numeric value from the selected component values, for example, by appending the selected vector component values to a single numeric value that is a decimal or base 10 representation of the numeric value, as shown in FIG. 5 by step 505.
[0135] The next operation is then one of generating an index value from that number as shown in FIG. 5 by step 507.
[0136] The operation of generating an index value from this number can be shown as the following loop: In step 571, as shown in FIG. 5, the loop index is set to 0 to start the loop, and the vector index is initialized to 0. Step 573, as shown in FIG. 5, checks whether the vector corresponding to the loop index value is valid, and if so, increments the vector index. Step 575 increments the loop index as shown in FIG. 5, stopping if the loop index is the input number, the index is the vector index value, otherwise a new check is performed on the incremented loop index value.
[0137] Next, in step 511, the ISM vector index value is output as shown in FIG.
[0138] With respect to Figure 6, the decoder shown in Figure 1 is shown in further detail with respect to decoding and generating decoded ISM ratio values. Additionally, Figures 7 and 8 show flow diagrams of exemplary decoder operations shown in Figure 6, according to some embodiments.
[0139] In some embodiments, the decoder includes a bitstream demultiplexer 601 configured to demultiplex the bitstream 118 and extract the coded ISM ratio index values 602. These coded ISM ratio index values 602 may be passed to a metadata decoder, which outputs decoded ISM ratio values and uses them to generate the spatial audio signal.
[0140] In other words, the method includes receiving / obtaining a bitstream as shown in FIG. 7 by step 701 .
[0141] The coded ISM rate index values are then demultiplexed from the bitstream as shown in FIG.
[0142] In some embodiments, the metadata decoder 603 includes an ISM vector index-to-vector generator 605. The ISM vector index-to-vector generator 605 is configured to generate decoded ISM vector values 606 from the encoded ISM ratio index values 602.
[0143] In some embodiments, this may be implemented using opposite indexing for vector value determination as described above.
[0144] Therefore, step 705 performs the operation of generating ISM vector values from ISM ratio index values, as shown in FIG.
[0145] This operation is further described with respect to FIG. 8 and the operation of the exemplary ISM vector index to vector generator shown in FIG. 6, where encoded vector quantized ISM rate index values are received or obtained by step 801.
[0146] Next, in step 803, a loop is started as shown in FIG. 8, with the index value=ISM ratio index value and variable J set to 0.
[0147] Next, in a loop as shown in FIG. 8 by step 805, if the vector corresponding to J is valid, the index value is decremented (by 1).
[0148] The value of J is incremented (by 1) as shown in FIG. 8 by step 807 .
[0149] A check is then made on the index value to determine if it is zero, as shown in Figure 8 by step 809. If the value is zero, the loop returns to step 805;
[0150] Next, the J numbers are assigned as the first n-1 components of the ISM vector as shown in FIG. 8 by step 811.
[0151] Furthermore, the final component is generated based on the difference between the sum of the n-1 components and the Laplacian norm value K, as shown in FIG.
[0152] Additionally, a deindexing function may be defined by the following pseudocode: Input: index Output: An array of integers with Laplacian norm equal to K 1. j=0 2. While(index>0) a. If the vector corresponding to j is valid i. index = index-1 b. End if c. j=j+1 3. End while 4. Assign the jth number to the first n-1 elements of the output array 5. Calculate the value of the last element as K minus the sum of the first (n-1) elements. This can further be expressed by a corresponding C language function. static void decode_index_slice( int16_tindex, int16_t*ratio_idx_ism, int16_tn, int16_tK) { int16_ti,j,sum,base[MAX_NUM_OBJECTS],elem; switch(n) { case 2: ratio_idx_ism[0]=index; ratio_idx_ism[1]=K-ratio_idx_ism[0]; break; Case 3: Case 4: { j=0; while(index>0) { if(valid(j,K,n-1)) { index--; } j++; } base[0]=1; for(i=1;i <n-1;i++) { base[i]=base[i-1]*10; } sum=0; for(i=n-2;i>=0;i--) { elem=j / base[i]; ratio_idx_ism[ni-2]=elem; sum+=elem; j-=elem*base[i]; } ratio_idx_ism[n-1]=K-sum; } } default: break; } }
[0153] The decoded ISM vector values 606 may then be passed to an ISM ratio generator 607 .
[0154] The metadata decoder 603, in some embodiments, may include an ISM ratio generator 607 configured to receive the decoded ISM vector values 606 and generate the decoded ISM ratios 608 in a manner that uses the opposite methodology to that described above.
[0155] The operation of generating ISM ratios from ISM vector values is illustrated in FIG.
[0156] The ISM ratio value may then be output as shown in FIG. 7 by step 709.
[0157] 9 is an exemplary electronic device that may be used as any of the apparatus portions of the systems described above. The device may be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc. The device may be configured to implement, for example, the encoder / analyzer portion and / or decoder portion shown in FIG. 1 or any of the functional blocks described above.
[0158] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. Processor 1407 can be configured to execute various program code, such as the methods described herein.
[0159] In some embodiments, device 1400 includes at least one memory 1411. In some embodiments, at least one processor 1407 is coupled to memory 1411. Memory 1411 may be any suitable storage means. In some embodiments, memory 1411 includes program code sections for storing program code implementable on processor 1407. Additionally, in some embodiments, memory 1411 may further include a storage data section for storing data, e.g., data that has been processed or will be processed in accordance with embodiments described herein. The implemented program code stored in the program code sections and the data stored in the storage data section may be retrieved by processor 1407 whenever needed via the memory-processor coupling.
[0160] In some embodiments, device 1400 includes a user interface 1405. User interface 1405 may be coupled to processor 1407 in some embodiments. In some embodiments, processor 1407 may control the operation of user interface 1405 and receive input from user interface 1405. In some embodiments, user interface 1405 may allow a user to input commands to device 1400, for example, via a keypad. In some embodiments, user interface 1405 may allow a user to obtain information from device 1400. For example, user interface 1405 may include a display configured to display information from device 1400 to the user. User interface 1405, in some embodiments, may include a touch screen or touch interface that can both allow information to be entered into device 1400 and also display information to the user of device 1400. In some embodiments, user interface 1405 may be a user interface for communication.
[0161] In some embodiments, device 1400 includes an input / output port 1409. In some embodiments, input / output port 1409 includes a transceiver. The transceiver in such embodiments may be coupled to processor 1407 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means may, in some embodiments, be configured to communicate with other electronic devices or apparatuses via a wire or wire coupling.
[0162] The transceiver may communicate with the further device via any suitable known communication protocol. For example, in some embodiments, the transceiver may use a suitable radio access architecture based on Long Term Evolution Advanced (LTE-Advanced, LTE-A) or New Radio (NR) (which may also be referred to as 5G), Universal Mobile Telecommunications System (UMTS) Radio Access Network (UTRAN or E-UTRAN), Long Term Evolution (same as LTE, E-UTRA), 2G networks (legacy network technologies), Wireless Local Area Networks (WLAN or Wi-Fi), Worldwide Interoperability for Microwave Access (WiMAX), Bluetooth®, Personal Communications Services (PCS), ZigBee®, Wideband Code Division Multiple Access (WCDMA), systems using Ultra Wideband (UWB) technology, sensor networks, Mobile Ad Hoc Networks (MANET), Cellular Internet of Things (IoT) RAN, and Internet Protocol Multimedia Subsystem (IMS), any other suitable options, and / or any combination thereof.
[0163] The input / output port 1409 of the transceiver may be configured to receive a signal.
[0164] In some embodiments, device 1400 may be used as at least part of a synthesis device. Input / output port 1409 may be coupled to headphones (which may be head-tracked or non-head-tracked headphones) or the like, and loudspeakers.
[0165] In general, various embodiments of the present invention may be implemented in hardware or special purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. While various aspects of the present invention may be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it will be appreciated that these blocks, apparatus, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing device, or some combination thereof.
[0166] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device, such as within a processor entity, or by hardware, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that any block of logic flow in the diagrams may represent program steps or interconnected logic circuits, blocks, and functions, or combinations of program steps and logic circuits, blocks, and functions. Software may be stored on physical media, such as memory chips or blocks implemented within a processor, magnetic media, such as hard disks or floppy disks, and optical media, such as DVDs and their data variants, CDs.
[0167] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technology environment and may include, by way of non-limiting example, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), gate-level circuitry, and a processor based on a multi-core processor architecture.
[0168] Embodiments of the present invention may be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert logic-level designs into semiconductor circuit designs ready to be etched into semiconductor substrates.
[0169] Programs offered by companies such as Synopsys, Inc. of Mountain View, California, and Cadence Design of San Jose, California, use established design rules and a library of pre-stored design modules to automatically route conductors and place components on semiconductor chips. Once the design of a semiconductor circuit is complete, the resulting design can be sent in a standardized electronic format (e.g., Opus, GDSII, etc.) to a semiconductor manufacturing facility or "fab" for fabrication.
[0170] As used in this application, the term "circuitry" may refer to one or more or all of the following: (a) Hardware-only circuit implementation (e.g., implementation using only analog and / or digital circuits), (b) For example (where applicable), the following combinations of hardware circuitry and software: (i) a combination of analog and / or digital hardware circuitry(s) and software / firmware; (ii) any portion of the hardware processor(s) using software (including digital signal processor(s)), software, and memory(s) that cooperate to cause a device, such as a mobile phone or server, to perform various functions; and hardware circuit(s) and / or processor(s), such as microprocessor(s) or portions of microprocessor(s), that require software (e.g., firmware) to operate, but may not have software present if not necessary for operation.
[0171] This definition of circuit applies to all uses of the term in this application, including any claims. As a further example, as used in this application, the term circuit also encompasses implementations of simply a hardware circuit or processor(s), or portions of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuit also encompasses, for example, baseband or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, other computing devices, or other network devices, where applicable to certain claim elements.
[0172] As used herein, the term "non-transitory" is not a limitation regarding the permanence of the data storage (eg, RAM vs. ROM), but rather a limitation of the medium itself (ie, tangible rather than signal).
[0173] As used herein, "at least one of: " and "at least one of " and similar phrases where a list of two or more elements is joined by "and" or "or" mean at least any one element, or at least any two or more elements, or at least all of the elements.
[0174] The foregoing description provides a full and informative description of exemplary embodiments of the present invention, by way of illustrative and non-limiting example. However, various modifications and adaptations may become apparent to those skilled in the art in light of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still be included within the scope of the present invention, as defined by the appended claims.
Claims
1. 1. An apparatus for encoding audio object parameters, comprising: obtaining a proportion parameter associated with each audio object in an audio environment, the audio environment including at least two audio objects, the proportion parameter being configured to identify a distribution of the respective object within the object portion of the overall audio environment; quantizing the ratio parameter for the audio object using a first number of bits; generating a vector from said quantization rate parameter selection; generating integer values based on indexing from the vector, the generated integer values representing the ratio parameters of the at least two audio objects; The apparatus comprising:
2. the means for generating the integer values based on the indexing from the vector, the generated integer values representing the ratio parameters of the at least two audio objects, generating a single number by appending elements from said vector; generating the index from the number by executing an iterative loop from zero iterations up to and including the number of iterations of the number, and sequentially associating index values with iteration numbers of the iterative loop having a valid vector, the integer value being the highest index value reached at the end of the iterative loop; 10. The device of claim 1, wherein the device is a means for:
3. 3. Apparatus according to claim 1 or claim 2, wherein the means for generating the vectors from the selections of the quantization rate parameters comprises means for generating the vectors from the selections of all but one of the quantization rate parameters.
4. The means for generating the vector from the selection of all but one of the quantization rate parameters comprises: generating a full vector from the quantization rate parameters of the audio object; generating the vector from a selection of all but one of the quantization rate parameters of the audio object; 4. The device of claim 3, which is a means for:
5. 5. The apparatus according to claim 1, wherein the means for quantizing the ratio parameter for the audio object using the first number of bits is means for scalar quantizing the ratio parameter for the audio object using the first number of bits.
6. 6. The apparatus of claim 5, wherein the first number of bits is three and the integer value is a decimal integer value.
7. When dependent on claim 2, the effective vector is The combination of vector element values is seven or less, or No element of said vector has a value greater than seven, and the combination of vector element values is seven or less.
7. The device of claim 6, wherein the device is one of:
8. 1. An apparatus for decoding a ratio parameter of an audio object, comprising: obtaining an integer value representing a ratio parameter of said audio object; converting the integer values into vectors representing selections of quantization rate parameters based on indexing of the vectors; regenerating at least one further quantization rate parameter from said vector selection of said quantization rate parameters; - dequantizing the quantization ratio parameters to obtain ratio parameters of the audio objects, the ratio parameters being configured to identify a distribution of a particular object within the object portion of the overall audio environment; The apparatus comprising:
9. The means for converting the integer values into the vector representing the selection of the quantization rate parameter based on the indexing of the vector comprises: generating the number from the integer value by executing an iterative loop including zero iterations through a single iteration and sequentially associating index values with iteration numbers of the iterative loop having a valid vector, the integer value being the highest index value; dividing the number into vector component values to generate the vector; 9. The device of claim 8, which is a means for:
10. 10. The apparatus of claim 8, wherein the means for regenerating at least one further quantization ratio parameter from the vector selection of quantization ratio parameters is means for generating at least one further quantization ratio parameter based on a value of a sum element of the vector subtracted from an expected sum value.
11. 11. The apparatus according to claim 8, wherein the means for dequantizing is means for scalar dequantizing the ratio parameters for the audio objects using a first number of bits, and for dequantizing the quantization ratio parameters to obtain ratio parameters for the audio objects, the ratio parameters being configured to identify the distribution of the particular objects within an object portion of the overall audio environment.
12. 12. The apparatus of claim 11, when dependent on claim 10, wherein the first number of bits is three, the expected total value is seven, and the integer value is a decimal integer value.
13. 1. A method for encoding audio object parameters, comprising: obtaining a proportion parameter associated with each audio object in an audio environment, the audio environment including at least two audio objects, the proportion parameter being configured to identify a distribution of the respective object within the object portion of the overall audio environment; quantizing the ratio parameter for the audio object using a first number of bits; generating a vector from said quantization rate parameter selection; generating integer values based on indexing from the vector, the generated integer values representing the ratio parameters of the at least two audio objects; The method comprising:
14. generating the integer values based on the indexing from the vector, the generated integer values representing the ratio parameters of the at least two audio objects; generating a single number by appending elements from said vector; generating the index from the number by executing an iterative loop from zero iterations up to and including the number of iterations of the number, and sequentially associating index values with iteration numbers of the iterative loop having a valid vector, the integer value being the highest index value reached at the end of the iterative loop; 14. The method of claim 13, comprising:
15. 15. A method according to claim 13 or claim 14, wherein generating the vector from the selection of the quantization rate parameters comprises generating the vector from the selection of all but one of the quantization rate parameters.
16. generating the vector from the selection of all but one of the quantization rate parameters generating a full vector from the quantization rate parameters of the audio object; generating the vector from a selection of all but one of the quantization rate parameters of the audio object; 16. The method of claim 15, comprising:
17. 17. The method of claim 13, wherein quantizing the ratio parameter for the audio object using the first number of bits comprises scalar quantizing the ratio parameter for the audio object using the first number of bits.
18. 1. A method for decoding a ratio parameter of an audio object, comprising: obtaining an integer value representing a ratio parameter of said audio object; converting the integer values into vectors representing selections of quantization rate parameters based on indexing of the vectors; regenerating at least one further quantization rate parameter from said vector selection of said quantization rate parameters; - dequantizing the quantization ratio parameters to obtain ratio parameters of the audio objects, the ratio parameters being configured to identify a distribution of a particular object within the object portion of the overall audio environment; The method comprising:
19. converting the integer values to the vector representing a selection of the quantization rate parameter based on the indexing of the vector, generating the number from the integer value by executing an iterative loop including zero iterations through a single iteration and sequentially associating index values with iteration numbers of the iterative loop having a valid vector, the integer value being the highest index value; dividing the number into vector component values to generate the vector; 20. The method of claim 18, comprising:
20. 20. A method according to claim 18 or claim 19, wherein regenerating at least one further quantization ratio parameter from the vector selection of quantization ratio parameters comprises generating at least one further quantization ratio parameter based on a value of a sum element of the vector subtracted from an expected sum value.
Citation Information
Patent Citations
Audio signal encoder
JP2017504829A
Audio and speech coding device, audio and speech decoding device, method for coding audio and speech, and method for decoding audio and speech
WO2013118476A1
Combining spatial audio streams
WO2022200666A1
Cited By
Parametric Spatial Audio Encoding
JP2026500131A