CODIFICAÇÃO DE ÁUDIO ESPACIAL PARAMÉTRICA DE BAIXA TAXA DE CODIFICAÇÃO
Patent Information
- Application Number
- BR112025020175
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-24
- Filing Date
- 2024-02-13
- Publication Date
- 2026-08-04
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
1 / 48 Low-Rate Parametric Spatial Audio Coding Field of Invention
[0001] This application relates to apparatus and methods for spatial audio representation and encoding, but not exclusively for audio representation for an audio encoder. Fundamentals of the Invention
[0002] Parametric spatial audio processing is a field of audio signal processing where the spatial aspect of sound is described using a set of parameters. For example, in parametric spatial audio capture from microphone arrays, it is a typical and effective choice to estimate from the microphone array sounds a set of parameters, such as sound directions in frequency bands, and the ratios between the directional and non-directional parts of the captured sound in frequency bands. These parameters are known to describe well the perceptual spatial properties of the captured sound at the microphone array position. These parameters can be used in spatial sound synthesis accordingly, for binaural headphones, loudspeakers, or other formats, such as ambisonics.
[0003] The directions and ratios of direct to total energy in the frequency bands are therefore a parameterization that is particularly effective for capturing spatial audio.
[0004] A set of parameters consisting of a directionality parameter in frequency bands and a power ratio parameter in frequency bands (indicating the directionality of the sound) can also be used as the spatial metadata (which may also include other parameters such as surround coherence, propagation coherence, number of directions, distance, etc.) for an audio codec. For example, these parameters can be estimated from audio signals captured from a microphone array, and, for example, a stereo or mono signal can be generated from Petition 870250085402, dated 09 / 22 / 2025, pages 225 / 283 2 / 48 of the microphone array signals to be transmitted with spatial metadata. The stereo signal can be encoded, for example, with an AAC encoder, and the mono signal can be encoded with an EVS encoder. A decoder can decode the audio signals into PCM signals and process the sound into frequency bands (using spatial metadata) to obtain the spatial output, for example, a binaural output.
[0005] Immersive audio codecs are being implemented supporting a multitude of operational points ranging from low bitrate operation to transparency. One example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, which is being designed to be suitable for use in a communications network, such as a 3GPP 4G / 5G network, including use in immersive services such as, for example, immersive voice and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and generic audio. Furthermore, it is expected to support channel-based audio and scene-based audio inputs, including spatial information about the sound field and sound sources. The codec should also operate with low latency to enable conversational services, as well as support high error robustness under various transmission conditions.
[0006] The above-mentioned immersive audio codecs are particularly suitable for encoding spatial sound captured from microphone arrays (e.g., in mobile phones, VR cameras, independent microphone arrays). However, such an encoder may have other types of input, e.g., speaker signals, audio object signals, ambisonic signals. Summary
[0007] According to a first aspect, an apparatus is provided comprising means configured to: receive a direction value for a period of time from an audio object; compare, as a first comparison, a bit allocation with a threshold bit allocation value; Petition 870250085402, dated 09 / 22 / 2025, pages 226 / 283 3 / 48 depending on the first comparison, quantize the direction value for the audio object's time period with a quantizer according to the bit allocation or compare, as a second comparison, the direction value for the audio object's time period with a direction value for a previous time period of the audio object; and, depending on the second comparison, quantize the direction value for the time period with the quantizer according to the bit allocation and signal the second comparison.
[0008] The means configured to, depending on the second comparison, quantize the direction value for the time period with the quantizer according to the bit allocation and signal the comparison can be configured to: quantize the direction value for the time period with the quantizer according to the bit allocation when the second comparison indicates that the direction value for the time period of the audio object is different from the direction value for the previous time period of the audio object and set a signal flag to indicate that the direction value for the previous time period differs from the direction value for the previous time period;and set a signal flag to indicate that the direction value for the time period is the same as the direction value for the previous time period when the second comparison indicates that the direction value for the audio object's time period is the same as the direction value for the previous time period of the audio object.
[0009] The means configured to, depending on the first comparison, quantize the direction value for the audio object's time period with a quantizer according to the bit allocation or compare, as a second comparison, the direction value for the audio object's time period with a direction value for a previous time period of the audio object can be configured to: quantize the direction value for the audio object's time period with a quantizer according to the bit allocation when the bit allocation for quantizing the direction value is not less than the threshold bit allocation value; and compare the direction value Petition 870250085402, dated 09 / 22 / 2025, pages 227 / 283 4 / 48 for the audio object's time period with a direction value for a previous time period of the audio object as the second comparison when the bit allocation for the direction value quantization is less than the threshold bit allocation value.
[0010] The quantizer can be a spherical lattice quantizer in which the spherical lattice is formed by covering the sphere with smaller spheres, where the smaller spheres define the points of the spherical lattice.
[0011] The direction value may comprise an azimuth value and an elevation value.
[0012] According to a second aspect, an apparatus is provided comprising means configured for: comparing a bit allocation for a direction value index for a time period of an audio object with a threshold bit allocation value; depending on the comparison, decoding the direction value index for the time period of the audio object with a dequantizer according to the bit allocation or reading a state of a signal flag associated with the direction value index for the time period of the audio object; and depending on the state of the signal flag;To define a quantized direction value for the audio object's time span as a quantized direction value for a previous time span of the audio object, or to decode the direction value index for the audio object's time span with a dequantizer according to the bit allocation, generating a quantized direction value for the audio object's time span, and to adjust an azimuth value of the quantized direction value for the audio object's frame time span according to a quantization resolution of the azimuth value.
[0013] The means configured to adjust an azimuth value of the quantized direction value for the time period of the audio object according to a quantization resolution of the azimuth value can be configured to: compare the product of the azimuth value of the quantized direction value for the time period and an azimuth value of the quantized direction value Petition 870250085402, dated 09 / 22 / 2025, pp. 228 / 283 5 / 48 for the previous time period of the audio object and when the product is greater than zero: determine a difference between two consecutive azimuth quantization values of the dequantizer; when a difference between the azimuth value of the direction value quantized for the time period and the azimuth value of the direction value quantized for the previous time period of the audio object is greater than half the difference between two consecutive azimuth quantization values of the dequantizer, subtract half the difference between two consecutive azimuth quantization values of the dequantizer from the azimuth value of the direction value quantized for the time period of the audio object;and when a difference between the azimuth value of the quantized direction value for the previous time period and the azimuth value of the quantized direction value for the audio object's time period is greater than half the difference between two consecutive azimuth quantization values of the dequantizer, add half the difference between two consecutive azimuth quantization values of the dequantizer to the azimuth value of the quantized direction value for the audio object's time period.
[0014] The means configured to determine a difference between two consecutive azimuth quantization values of the dequantizer can be configured to: divide the circumference of a circle by a number of azimuth values, where the number of azimuth values is determined by the elevation value of the quantized direction value for the time period.
[0015] The state of the signal flag indicates one of: a direction value for the audio object's time period differs from a direction value for the previous time period of the audio object; or a direction value for the audio object's time period is the same as a direction value for the previous time period of the audio object;
[0016] The dequantizer is a spherical lattice quantizer in which the spherical lattice is formed by covering the sphere with smaller spheres, where the smaller spheres define the points of the spherical lattice. Petition 870250085402, dated 09 / 22 / 2025, pages 229 / 283 6 / 48
[0017] According to a third aspect, a method is provided comprising: receiving a direction value for a time period of an audio object; comparing, as a first comparison, a bit allocation with a threshold bit allocation value; depending on the first comparison, quantizing the direction value for the time period of the audio object with a quantizer according to the bit allocation or comparing, as a second comparison, the direction value for the time period of the audio object with a direction value for a previous time period of the audio object; and, depending on the second comparison, quantizing the direction value for the time period with the quantizer according to the bit allocation and signaling the second comparison.
[0018] When depending on the second comparison, quantizing the direction value for the time period with the quantizer according to the bit allocation and signaling the comparison may comprise: quantizing the direction value for the time period with the quantizer according to the bit allocation when the second comparison indicates that the direction value for the time period of the audio object is different from the direction value for the previous time period of the audio object and setting a signal flag to indicate that the direction value for the previous time period differs from the direction value for the previous time period; and setting a signal flag to indicate that the direction value for the time period is the same as the direction value for the previous time period when the second comparison indicates that the direction value for the time period of the audio object is the same as the direction value for the previous time period of the audio object.
[0019] When, depending on the first comparison, quantizing the direction value for the audio object's time period with a quantizer according to bit allocation, or comparing, as a second comparison, the direction value for the audio object's time period with a direction value for a previous time period of the audio object, it may comprise: quantizing the direction value for the audio object's time period of Petition 870250085402, dated 09 / 22 / 2025, pages 230 / 283 7 / 48 audio with a quantizer according to bit allocation when the bit allocation for quantizing the direction value is not less than the threshold bit allocation value; and compare the direction value for the audio object's time period with a direction value for a previous time period of the audio object as the second comparison when the bit allocation for quantizing the direction value is less than the threshold bit allocation value.
[0020] The quantizer can be a spherical lattice quantizer in which the spherical lattice is formed by covering the sphere with smaller spheres, where the smaller spheres define the points of the spherical lattice.
[0021] The direction value may comprise an azimuth value and an elevation value.
[0022] According to a fourth aspect, a method is provided comprising: comparing a bit allocation for a direction value index for a time period of an audio object with a threshold bit allocation value; depending on the comparison, decoding the direction value index for the time period of the audio object with a dequantizer according to the bit allocation or reading a state of a signal flag associated with the direction value index for the time period of the audio object; and depending on the state of the signal flag;To define a quantized direction value for the audio object's time span as a quantized direction value for a previous time span of the audio object, or to decode the direction value index for the audio object's time span with a dequantizer according to the bit allocation, generating a quantized direction value for the audio object's time span, and to adjust an azimuth value of the quantized direction value for the audio object's frame time span according to a quantization resolution of the azimuth value.
[0023] Adjust an azimuth value of the quantized direction value for the audio object's time period according to a resolution of Petition 870250085402, dated 09 / 22 / 2025, pages 231 / 283 8 / 48 Azimuth value quantization may involve: comparing the product of the azimuth value of the direction value quantized for the time period and an azimuth value of the direction value quantized for the previous time period of the audio object, and when the product is greater than zero: determining a difference between two consecutive azimuth quantization values of the dequantizer; when a difference between the azimuth value of the direction value quantized for the time period and the azimuth value of the direction value quantized for the previous time period of the audio object is greater than half the difference between two consecutive azimuth quantization values of the dequantizer, subtracting half the difference between two consecutive azimuth quantization values of the dequantizer from the azimuth value of the direction value quantized for the time period of the audio object;and when a difference between the azimuth value of the quantized direction value for the previous time period and the azimuth value of the quantized direction value for the audio object's time period is greater than half the difference between two consecutive azimuth quantization values of the dequantizer, add half the difference between two consecutive azimuth quantization values of the dequantizer to the azimuth value of the quantized direction value for the audio object's time period.
[0024] Determining a difference between two consecutive azimuth quantization values of the dequantizer can involve: dividing the circumference of a circle by a number of azimuth values, where the number of azimuth values is determined by the elevation value of the quantized direction value for the time period.
[0025] The state of the signal flag can indicate one of: a direction value for the audio object's time period differs from a direction value for the previous time period of the audio object; or a direction value for the audio object's time period is the same as a direction value for the previous time period of the audio object;
[0026] The dequantizer can be a spherical lattice quantizer in Petition 870250085402, dated 09 / 22 / 2025, pages 232 / 283 9 / 48 that the spherical grid is formed by covering the sphere with smaller spheres, where the smaller spheres define the points of the spherical grid.
[0027] According to a fifth aspect, an apparatus is provided comprising at least one processor and at least one memory storing instructions which, when executed by at least one processor, cause the system to at least: receive a direction value for a time period of an audio object; compare, as a first comparison, a bit allocation with a threshold bit allocation value; depending on the first comparison, quantize the direction value for the time period of the audio object with a quantizer according to the bit allocation or compare, as a second comparison, the direction value for the time period of the audio object with a direction value for a previous time period of the audio object; and, depending on the second comparison, quantize the direction value for the time period with the quantizer according to the bit allocation and signal the second comparison.
[0028] According to a sixth aspect, an apparatus is provided comprising at least one processor and at least one memory storing instructions which, when executed by at least one processor, cause the system to at least: compare a bit allocation for a direction value index for a time span of an audio object with a threshold bit allocation value; depending on the comparison, decode the direction value index for the time span of the audio object with a dequantizer according to the bit allocation or read a state of a signal flag associated with the direction value index for the time span of the audio object; and depending on the state of the signal flag;Define a quantized direction value for the audio object's time span as a quantized direction value for a previous time span of the audio object, or decode the direction value index for the audio object's time span with a dequantizer according to the bit allocation, generating a quantized direction value for the object's time span. Petition 870250085402, dated 09 / 22 / 2025, pages 233 / 283 10 / 48 audio and adjust an azimuth value of the quantized direction value for the audio object's time frame period according to a quantization resolution of the azimuth value.
[0029] An apparatus comprising means for carrying out the actions of the method as described above.
[0030] A device configured to perform the actions of the method as described above.
[0031] A computer program comprising program instructions for inducing a computer to perform the method as described above.
[0032] A computer program product stored on a medium may cause an apparatus to perform the method as described in this document.
[0033] An electronic device may comprise an apparatus as described in this document.
[0034] A set of chips may comprise an apparatus as described in this document.
[0035] The modalities of this application aim to address problems associated with the state of the art. Summary of Figures
[0036] For a better understanding of the present application, reference will now be made, by way of example, to the attached figures, in which: Figure 1 schematically shows a suitable apparatus system for implementing some modalities; Figure 2 schematically shows an example of a coding mode selector as shown in the device system as shown in Figure 1 according to some modes; Figure 3 shows a flow diagram of the operation of the encoding mode selector example shown in Figure 2 according to some modalities; Figure 4 shows a flow diagram of the operation of the example of a Petition 870250085402, dated 09 / 22 / 2025, pages 234 / 283 11 / 48 first, lower or single MASA bit rate encoding mode shown in Figure 4 according to some embodiments; Figure 5 shows a flow diagram of the operation of the example of a second, lower, or object information encoding mode shown in Figure 4 according to some modalities; Figure 6 shows a flow diagram of the operation of the example of a third, superior, or single-object encoding mode shown in Figure 4 according to some modalities; Figure 7 shows a flow diagram of the operation of the example of a fourth, upper, or multi-input and independent object encoding mode shown in Figure 4 according to some modalities; Figure 8 schematically shows an example of an audio object metadata encoder as shown in Figure 1 according to some modalities; Figure 9 shows a flow diagram of the operation of the audio object metadata encoder encoding mode selector example shown in Figure 8 according to some modalities; Figure 10 shows a flow diagram of the operation of the encoder determinator example and the low-rate encoder in Figure 8 according to some embodiments; Figure 11 schematically shows an example of an audio object direction value decoder according to some modalities; Figure 12 shows a flow diagram of an operation of the low-rate decoder example in Figure 11 according to some modes; and Figure 13 shows an example of a suitable device for implementing the apparatus shown in the previous figures. Order Types
[0037] The following describes in more detail suitable apparatus and possible mechanisms for encoding spatial audio signals. Petition 870250085402, dated 09 / 22 / 2025, pages 235 / 283 12 / 48 parametric codes comprising transport audio signals and spatial metadata. As indicated above, immersive audio codecs (such as 3GPP IVAS) are being planned that support a multitude of operating points ranging from low bitrate operation to transparency. It is expected to support channel-based audio and scene-based audio inputs, including spatial information about the sound field and sound sources. The following codec example is configured to be able to receive multiple input formats. In particular, the codec is configured to obtain or receive a multi-audio signal (e.g., received from a microphone array, or as a multi-channel audio format input, an ambisonic format input) and one or more audio object signals (these may also be called an independent stream with metadata - ISM format).Furthermore, in some situations, the codec is configured to handle more than one input format at a time. This combined (input) format mode can, for example, allow the simultaneous encoding of two different audio input formats. An example of two different audio input formats currently being considered is the combination of the MASA format with the audio object format. Metadata-assisted spatial audio (MASA) is an example of a parametric spatial audio format and a suitable representation as an input format for IVAS.
[0038] It can be considered as an audio representation consisting of 'N channels + spatial metadata'. It is a scene-based audio format particularly suited for capturing spatial audio on practical devices such as smartphones. The idea is to describe the sound scene in terms of sound source directions that vary in time and frequency and, for example, energy ratios. Sound energy that is not defined (described) by the directions is described as diffuse (coming from all directions).
[0039] As discussed above, the spatial metadata associated with audio signals can comprise multiple parameters (such as Petition 870250085402, dated 09 / 22 / 2025, pages 236 / 283 13 / 48 multiple directions and associated with each direction (or directional value) a direct-to-total energy ratio, propagation coherence, distance, etc.) per time-frequency window. Spatial metadata may also comprise other parameters or may be associated with other parameters that are considered non-directional (such as surround coherence, diffuse-to-total energy ratio, residual-to-total energy ratio), but when combined with directional parameters can be used to define the characteristics of the audio scene. For example, a reasonable design choice that is capable of producing good quality output is one in which spatial metadata comprising one or more directions for each time-frequency subframe (and associated with each direct-to-total direction ratio, propagation coherence, distance values, etc.) are determined.
[0040] The concept, as discussed in more detail in this document, is the definition of a multi-rate encoding model that provides encoding of a combined format at multiple bitrates. This encoding model allows for parametric encoding of the audio object input that includes variable rate encoding for an object direction parameter based on a given object priority and the available bitrate. Thus, it comprises modalities, as discussed in more detail in this document, of devices and methods for defining a quantization resolution of the object metadata as a function of the MASA-to-total power ratios and ISM ratios. Based on these parameters, a priority value is calculated for each object and is used to alter the quantization resolution of the object metadata.Furthermore, if the multi-rate encoding model determines a lower encoding rate for encoding the object's direction parameter, it includes modalities that can exploit the stability of the object's direction parameter over time.
[0041] As described above, the parametric representation of spatial metadata can use multiple simultaneous spatial directions. With Petition 870250085402, dated 09 / 22 / 2025, pages 237 / 283 In the 14 / 48 MASA model, the maximum proposed number of simultaneous directions is two. For each simultaneous direction, there may be associated parameters, such as: Direction Index; Direct-to-Total Ratio; Propagation Coherence; and Distance. In some modes, other parameters, such as Diffuse-to-Total Energy Ratio; Surround Coherence; and Remaining-to-Total Energy Ratio are defined.
[0042] In this regard, Figure 1 represents an example of apparatus 100 and system for implementing application modalities. The system is shown with an 'analysis' part. The 'analysis' part is the part that receives the multi-channel signals up to an encoding of the metadata and downmix signal.
[0043] The input for the 'analysis' part of the system is the 102 multi-channel audio signals. In the following examples, a microphone channel signal input is described; however, any suitable input format (or synthetic multi-channel) can be implemented in other embodiments. For example, in some embodiments, the spatial analyzer and spatial analysis can be implemented externally to the encoder. For example, in some embodiments, the spatial metadata (MASA) associated with the audio signals can be provided to an encoder as a separate bitstream. In some embodiments, the spatial metadata (MASA) can be provided as a set of spatial index (direction) values.
[0044] In addition, Figure 1 also depicts multiple audio objects 104 as an additional input for the analysis portion. As mentioned above, these multiple audio objects (or audio object stream) 104 can represent various sound sources within a physical space. Each audio object can be characterized by an audio signal (object) and accompanying metadata comprising directional data (in the form of azimuth and elevation values) that indicate the position or direction of the audio object within a physical space on an audio frame basis.
[0045] The multi-channel signals 102 are passed to an analyzer and encoder 101, and specifically a signal generator of Petition 870250085402, dated 09 / 22 / 2025, pages 238 / 283 15 / 48 transport 105 and to a metadata generator 103.
[0046] In some embodiments, the metadata generator 103 is also configured to receive the multichannel signals and analyze the signals to produce metadata 104 associated with the multichannel signals and thus associated with the transport signals 106. The analysis processor 103 can be configured to generate metadata that may comprise, for each time-frequency analysis interval, a direction parameter and a power ratio parameter and a coherence parameter (and in some embodiments a diffusion parameter). The direction, power ratio and coherence parameters may, in some embodiments, be considered spatial audio MASA parameters (or MASA metadata). In other words, spatial audio parameters comprise parameters that aim to characterize the sound field created / captured by the multichannel signals (or two or more audio signals in general).
[0047] In some embodiments, the generated parameters may differ from frequency band to frequency band. Thus, for example, in band X all parameters are generated and transmitted, while in band Y only one of the parameters is generated and transmitted and, furthermore, in band Z no parameters are generated or transmitted. A practical example of this may be that, for some frequency bands, such as the highest band, some of the parameters are not required for perception reasons. The transport signals 106 and the metadata 104 may be passed to a combined encoder core 109.
[0048] In some embodiments, the transport signal generator 105 is configured to receive the multi-channel signals and generate a suitable transport signal comprising a specified number of channels and output the transport signals 106 (MASA transport audio signals). For example, the transport signal generator 105 may be configured to generate a 2-channel audio downmix of the multi-channel signals. The specified number of channels may be any suitable number of channels. The Petition 870250085402, dated 09 / 22 / 2025, pp. 239 / 283 In some modes, a 16 / 48 transport signal generator is configured to select or combine, for example by beamforming techniques, the input audio signals for a specified number of channels and output them as transport signals.
[0049] In some embodiments, the transport signal generator 105 is optional and the multi-channel signals are passed unprocessed to a combined encoder core 109 in the same way as the transport signal is in this example.
[0050] Audio objects 104 can be passed to audio object analyzer 107 for processing. In some embodiments, audio object analyzer 107 analyzes the audio input stream of object 104 in order to produce suitable audio object transport signals 128 and audio object metadata 108. For example, audio object analyzer 107 can be configured to produce audio object transport signals 12 by downmixing the audio signals of audio objects 104 into a stereo channel together using amplitude panning based on the associated audio object directions. Furthermore, the audio object analyzer can also be configured to produce audio object metadata 108 associated with the audio object input stream 104. The audio object metadata can comprise direction values that are applicable to all sub-bands. So, if there are 4 objects, there are 4 directions.In the examples described in this document, the direction values also apply to all sub-frames of the frame, but in some modes the temporal resolution of the direction values may differ, and the direction values may apply to one or more sub-frames of the frame. Additionally, power ratios (or ISM ratios) can be determined for each object. The power ratio (ISM ratio) defines the object's contribution within the object's share of the total audio environment. In the following examples, the power ratios (or ISM ratios) are for each time-frequency window for each object.
[0051] In some modes, the audio object analyzer 107 Petition 870250085402, dated 09 / 22 / 2025, pages 240 / 283 17 / 48 may be located elsewhere, and the input audio objects 104 for the analyzer and encoder 101 are audio object transport signals and audio object metadata.
[0052] The analyzer and encoder 101 may comprise a combined encoder core 109 that is configured to receive the transport audio signals (e.g., downmix) 106 and audio object transport signals 128 in order to generate a suitable encoding of these audio signals.
[0053] The parser and encoder 101 may also comprise an audio object metadata encoder 111 which is similarly configured to receive audio object metadata 108 and output an encoded or compressed form of the input information as encoded audio object metadata 112.
[0054] In some embodiments, the combined encoder core can be configured to implement a stream separation metadata determiner and encoder that can be configured to determine the relative contributory proportions of the 102 multichannel signals (MASA audio signals) and 104 audio objects to the overall audio scene. This proportionality measure produced by the stream separation metadata determiner and encoder can be used to determine the proportion of quantization and encoding effort spent on the input 102 multichannel signals and 104 audio objects. In other words, the stream separation metadata determiner and encoder can produce a metric that quantifies the proportion of encoding effort spent on the 102 multichannel audio signals compared to the encoding effort spent on the 104 audio objects.This metric can be used to trigger the encoding of audio object metadata 108. Furthermore, the metric as determined by the metadata separation determiner and encoder can also be used as an influencing factor in the encoding process of transport audio signals 106 and audio object transport audio signal 128 performed by the combined encoder core 109. The output metric. Petition 870250085402, dated 09 / 22 / 2025, pages 241 / 283 18 / 48 of the stream separation metadata determinant and encoder can, in addition, be represented as encoded stream separation metadata and be combined into the encoded metadata stream of the combined encoder core 109.
[0055] In some embodiments, the analyzer and encoder 101 comprise a bitstream generator 113 configured to obtain the encoded metadata 116, the encoded transport audio signals 138 and the encoded audio object metadata 112 and generate the bitstream 118 for potential transmission or storage.
[0056] In some embodiments, the analyzer and encoder 101 comprise an encoder controller 115. The encoder controller 115 may, in some embodiments, control the encoding implemented by the audio object metadata encoder 111 and the combined encoder core 109. In some embodiments, the encoder controller 115 is configured to determine the bitrate for the bitstream 118 and based on the bitrate control of the encoding. In some embodiments, the encoder controller 115 is additionally configured to control at least one of the audio object analyzer 107, transport signal generator 105, and metadata generator 103 in the parameter generation.
[0057] The analyzer and encoder 101 may, in some embodiments, be a computer or mobile device (running suitable software stored in memory and on at least one processor) or, alternatively, a specific device using, for example, FPGAs or ASICs. Encoding may be implemented using any suitable scheme. In some embodiments, the encoder 107 may additionally interleave, multiplex into a single data stream, or incorporate the encoded MASA metadata, audio object metadata, and stream separation metadata within the encoded (downmixed) transport audio signals prior to transmission or storage, as shown in Figure 1 by the dashed line. Multiplexing may be implemented using any suitable scheme. Petition 870250085402, dated 09 / 22 / 2025, pages 242 / 283 19 / 48
[0058] Furthermore, with regard to Figure 1, an associated decoder and renderer 129 is shown which are configured to obtain the bitstream 118 comprising encoded metadata 116, encoded transport audio signals 138 and encoded audio object metadata 112 and from these generate suitable spatial audio output signals. The decoding and processing of such audio signals are known in principle and are not discussed in detail hereafter, except for the decoding of the encoded ISM ratio metadata.
[0059] In relation to Figure 2, the encoder controller 115 is shown in greater detail according to some modalities.
[0060] In this example, the encoder controller 115 comprises a bitrate determiner / monitor 201 configured to determine and / or monitor the bitrate available for the bandwidth for the encoded audio and metadata. This can be determined based on an estimate of the transmission path bandwidth (and, for example, based on an estimated signal strength) or a determination of bandwidth storage to keep the file below a required size for a specified time or by any other suitable method.
[0061] The bitrate determiner / monitor 201 can, in addition, be configured to control an encoding mode selector 203. The encoder controller 115 can comprise an encoding mode selector 203 configured to select an encoding mode, for example, based on the determined bandwidth or bitrate, and then control the encoders, for example, the combined encoder core 109 and the audio object metadata encoder 111.
[0062] With regard to Figure 3, a flow diagram of an example operation of the encoder controller 115 shown in Figure 2 is shown. In this example, there is an initial operation to receive or obtain or otherwise determine the bit rate or bandwidth for encoded parameters and audio data as shown in Figure 3 by step 301. Petition 870250085402, dated 09 / 22 / 2025, pages 243 / 283 20 / 48
[0063] Having obtained the available bandwidth or bit rate, a check can then be made to determine if the bit rate is below a first threshold limit (or lower or minimum object), as shown in Figure 3 by step 303.
[0064] When the available bandwidth or bit rate is below the first threshold limit (or lower or minimum object), then the encoders can be controlled to encode the transport channels and only the MASA metadata (also shown as Mode A), as shown in Figure 3 by step 304.
[0065] When the available bandwidth or bit rate is above the first threshold limit (or lower or minimum object limit), then an additional check can be made to determine if the bit rate is below a second threshold limit (or lower or single object limit), as shown in Figure 3 by step 305.
[0066] When the available bandwidth or bit rate is below the second threshold limit (or lower or single object), then the encoders can be controlled to encode transport channels, MASA metadata, ISM metadata (all objects), MASA to total ratios, ISM ratios (also shown as Mode B), as shown in Figure 3 by step 306.
[0067] When the available bandwidth or bit rate is above the second threshold limit (either lower or single object), then an additional check can be made to determine if the bit rate is below a third threshold limit, either upper or total object, as shown in Figure 3 by step 307.
[0068] When the available bandwidth or bit rate is below the third threshold limit, or above or above the total object limit, then the encoders can be controlled to encode transport channels, MASA metadata, ISM metadata (all objects), MASA ratios to total, ISM ratios and 1 object audio data, with 1 object identifier (also Petition 870250085402, dated 09 / 22 / 2025, pages 244 / 283 21 / 48 shown as Mode C), as shown in Figure 3 by step 308.
[0069] When the available bandwidth or bit rate is above the third threshold limit, or higher or total object, then the encoders can be controlled to encode transport channels, MASA metadata, ISM metadata (all objects), audio data from all objects (also shown as Mode D), as shown in Figure 3 by step 310.
[0070] The encoding modes can, for example, be summarized by the following table. Mode | Bitrate / Track Format | Encoded Parameters | A - 32kbps - Transport channels - MASA metadata | B 48-80kbps ISM_MASA_MODE_PARAM - Transport channels - MASA metadata - ISM metadata - MASA-to-total ratios - ISM ratios | C 96-128kbps ISM_MASA_MODE_ONE_OBJ - Transport channels - MASA metadata - ISM metadata - MASA-to-total ratios - ISM ratios - Audio data of 1 object - Identifier of 1 object | D 160 - ISM_MASA_MODE_DISC - MASA transport channels - MASA metadata - ISM metadata - Audio data of all objects Table 1 Petition 870250085402, dated 09 / 22 / 2025, pages 245 / 283 22 / 48
[0071] The bit rates shown in this document are examples and it would be understood that they may be other specific values.
[0072] With regard to Figures 4 to 7, flow diagrams are shown depicting a first encoding mode (either minimum or combined), as shown in Figure 3 by step 304, a second encoding mode (either lower or object metadata), as shown in Figure 3 by step 306, a third encoding mode (either upper or single object), as shown in Figure 3 by step 308, and a fourth encoding mode (either maximum or all objects), as shown in Figure 3 by step 310, respectively.
[0073] For example, Figure 4, the A-mode encoding method, shows the first encoding mode (either minimal or combined), as shown in Figure 3 by step 304 in more detail. Thus, for very low total bit rates (e.g., less than or equal to 32kbps), all encoding is implemented using a MASA representation.
[0074] Thus, for example, there is an operation to receive / obtain the object-based streams (independent streams with metadata) and multi-channel transport audio signals (MASA stream) and metadata as shown in Figure 4 by step 401.
[0075] So, as shown in Figure 4 by step 403, there is an operation to generate an object-based MASA flow from the object flows (independent flows with metadata). This object-based MASA flow can, in some modes, be created from the object flow using, for example, the methods presented in WO2019086757A1.
[0076] After that, as shown in Figure 4 by step 405, the object-based MASA stream and the multi-channel-based MASA stream are combined. In some embodiments, the original MASA stream and the MASA stream created from the objects can be combined using the method presented in GB2574238. The decoder obtains the objects and the MASA audio content in MASA format. Petition 870250085402, dated 09 / 22 / 2025, pages 246 / 283 23 / 48
[0077] Then, the combined stream is output as shown in Figure 4 by step 407. In such modes, the audio content of the object (along with the MASA audio content) is present in the decoded audio scene, but the objects cannot be edited or separated from the scene in the decoder.
[0078] Figure 5 shows the encoding method of mode B, the second encoding mode (or lower or object metadata), as shown in Figure 3 by step 306. Thus, for low bitrates (e.g., between 48kbps and 80kbps) and since there are more bits available, there is the possibility of parameterizing the audio scene by sending a downmix of common audio data, the MASA metadata, the ISM metadata, and additional parameter sets indicating for each time-frequency window how much of the signal corresponds to the MASA component of the total audio scene (in other words, this can be presented or indicated by the MASA-to-total power ratios) and ratios indicating how the audio scene corresponding to the objects is distributed among the ISMs (in other words, this can be presented or indicated by the ISM ratios).
[0079] Thus, for example, there is a method step of receiving / obtaining object-based streams (independent streams with metadata) and multi-channel transport audio signals (MASA stream) and metadata as shown in Figure 5 by step 501.
[0080] Then, as shown by step 503 in Figure 5, generate combined audio signals from MASA and object-based downmix (channel pair element). In other words, the audio content from MASA and objects is downmixed to 2 channels (CPE channel pair element).
[0081] The MASA to total ratios and the ISM ratios can be determined as shown in Figure 5 by step 505.
[0082] The MASA to total ratios and the ISM ratios can then be encoded based on any suitable coding method. For example, the ISM ratios can be encoded using a coding of Petition 870250085402, dated 09 / 22 / 2025, pages 247 / 283 24 / 48 network or other vector quantization method or the MASA to total ratios encoded by DCT transformation followed by entropy encoding (e.g., as described in WO2022 / 200666). The encoding of the MASA to total ratios and ISM ratios is shown in Figure 5 by step 507.
[0083] Furthermore, MASA metadata can then be encoded based on any suitable MASA metadata encoding method, as shown in Figure 5 by step 509.
[0084] The combined audio signals can then be encoded based on any suitable audio signal encoding method, as shown in Figure 5 by step 511.
[0085] The encoder can then output encoded MASA metadata, MASA-to-total ratios, ISM ratios, and combined transport audio signals as shown in Figure 5 by step 513.
[0086] Figure 6 shows the encoding method of mode C, the third (or higher or single-object) encoding mode, as shown in Figure 3 by step 308. Thus, at medium or higher bitrates (e.g., bitrates greater than or equal to 96kbps and less than 160kbps), the audio content of an object is separated and sent independently. Furthermore, the downmix formed from the MASA transport channels and the remaining objects are sent in MASA format with the additional parameters of MASA-to-total-energy ratios and ISM ratios. Additionally, ISM metadata is sent and an identifier describing which object was separated. In each frame, it is decided which object should be separated. The decision can, for example, be based on the relative level of the objects in relation to other objects (e.g., separating the loudest object). This is explained in detail in WO2022 / 214730.
[0087] Thus, for example, there is a step in the method of receiving / obtaining object-based streams (independent streams with metadata) and multi-channel transport audio signals (MASA stream) and metadata as shown in Figure 6 by step 601. Petition 870250085402, dated 09 / 22 / 2025, pages 248 / 283 25 / 48
[0088] Then, as shown by step 603 in Figure 6, an audio object is selected and an object identifier is generated based on the selected audio object. Additionally, the audio signal associated with the selected audio object is encoded. Any suitable audio signal encoder can be used to encode the audio signal of the selected object. For example, the same or similar audio signal encoder used to encode the MASA audio signals can be employed. Then, a combined MASA and remaining (or unselected) object-based transport audio signals (or downmix) are generated as shown in Figure 6 by step 605. The object transport signals can be created in the same way as presented in the previous mode, mode B, with the difference that the selected or separate object is not included in the mix.For example, multichannel audio signals or MASA and object transport signals (unselected) can be summed to generate combined transport audio signals.
[0089] The MASA to total ratios and the ISM ratios can be determined as shown in Figure 6 by step 607.
[0090] The object identifier, MASA metadata, object metadata for all objects, MASA to total ratios, and ISM ratios can then be encoded based on any suitable encoding method, as shown in Figure 6 by step 609. For example, the encoding can employ any scalar or vector quantizer followed or not by entropy encoding. The encoding of the MASA to total energy ratio can be implemented in the manner described in WO2022 / 200666. The encoding of the ISM ratios is described later in further detail.
[0091] The combined audio signals can then be encoded based on any suitable MASA audio signal coding method, as shown in Figure 6 by step 611. The encoding of the combined transport audio signals can employ any suitable transport audio signal coding, for example, the MASA encoder. Petition 870250085402, dated 09 / 22 / 2025, pages 249 / 283 26 / 48
[0092] In other words, the separate object is determined, separated and coded as described in WO2022 / 214730, and for the remaining objects and the MASA flow the processing works as described in WO2022 / 200666.
[0093] The encoder can then output the encoded object identifier, MASA metadata, MASA-to-total ratios, ISM ratios, object metadata (for all objects), selected single object audio signal, and combined transport audio signals as shown in Figure 6 by step 613.
[0094] Figure 7 shows the D-mode encoding method, the fourth (or higher, or all-object) encoding mode, as shown in Figure 3 by step 310. Thus, at higher bit rates (e.g., bit rates above or equal to 160kbps), the two input audio formats, MASA and ISM, are encoded and transmitted independently.
[0095] Thus, for example, there is a method step of receiving / obtaining object-based streams (independent streams with metadata) and multichannel-based transport audio signals (MASA stream) and metadata as shown in Figure 7 by step 701.
[0096] Then, as shown by step 703 in Figure 7, the multi-channel based transport audio signals (MASA stream) and metadata are encoded based on any suitable MASA encoding method.
[0097] The object (independent streams with metadata) and the associated metadata can, in addition, be encoded as shown in Figure 7 by step 705. Any suitable mono encoder (as part of the main encoder) can be employed to implement the encoding, for example, an EVS-based mono encoder.
[0098] The encoder can then output the independently encoded object (independent streams with metadata) and associated metadata, and independently encoded multi-channel based transport audio signals (MASA stream) and metadata as shown. Petition 870250085402, dated 09 / 22 / 2025, pages 250 / 283 27 / 48 in Figure 7 by step 707.
[0099] With regard to the following, the generation and encoding of ISM ratio values, as determined and encoded within encoding modes B and C, is described in more detail.
[0100] With regard to the following encoding of the direction metadata parameter associated with objects (as determined from the ISM metadata), it is described and, in some embodiments, the encoding of the direction metadata parameter associated with objects when the encoder is operating in encoding modes B and C (the second or third). However, it would be appreciated that, in some embodiments, the following could also be applied to any encoding mode in which audio object metadata (and specifically direction metadata) is encoded. In embodiments where MASA-to-total ratios and ISM ratios are not sent, extra information about the encoding details (bit allocation) is sent. Furthermore, in some embodiments, the ratios could, in principle, be calculated from the encoded signals, both in the encoder and in the decoder.
[0101] Thus, with regard to Figure 8, the audio object metadata encoder 111 is shown in greater detail according to some embodiments. In the following examples, the audio object metadata encoder 111 is configured to receive as input the ISM metadata and specifically the MASA-to-total energy ratio m2t 812, the ISM ratios r 814 and the direction values 802. In other words, the ISM metadata comprises directional information (elevation and azimuth) for each object in each frame. There is a standard resolution of 11 bits per elevation-azimuth pair. In addition, the MASA-to-total ratios and the ISM ratios can be determined or otherwise obtained from the ISM and also from multi-channel audio signals (e.g., as described in WO2022 / 200666).
[0102] In some modalities, the ISM ratios may, for example, be Petition 870250085402, dated 09 / 22 / 2025, pages 251 / 283 28 / 48 obtained as follows.
[0103] First, the audio signals of object underJ(t, i) are transformed into time-frequency domain underj(b, n) (where t is the time sample index, b is the frequency bin index, n is the time frame index and i is the object index). The time-frequency domain signals can, for example, be obtained through short-duration Fourier transform (STFT) or banks of quadrature filters modulated by complexes (QMF) (or low-delay variants thereof).
[0104] So, the energies of objects are computed in frequency bands. Eobi(kni)=Σbb^\Sobi(b,n,i')\2(1) where bbbaixa is the smallest and baaU is the largest bin in the frequency band k. Then, the ISM ratios ξ(k,n> i) can be calculated as ξ (k ,n,Q Eo bj(k Λΐ) Σΐ-1Εο1)μ ,ni) (2) where I is the number of objects.
[0105] In some embodiments, the temporal resolution of the ISM ratios may be different from the temporal resolution of the audio signals in the time-frequency domain S'obj(b, n, i) (i.e., the temporal resolution of the spatial metadata may be different from the temporal resolution of the time-frequency transformation). In such cases, the calculation (of energy and / or ISM ratios) may include summing across multiple time frames of the audio signals in the time-frequency domain and / or the energy values.
[0106] ISM ratios are numbers between 0 and 1 and correspond to the fraction with which an object is active within the audio scene created by all objects. For each object, there is one ISM ratio per frequency sub-band and time sub-frame. As discussed above, ISM ratios are passed to the audio object metadata encoder 111.
[0107] In some modes, the object metadata encoder Petition 870250085402, dated 09 / 22 / 2025, pages 252 / 283 Audio 111 29 / 48 is configured to encode MASA-to-total ratios and ISM ratios; the specific encoding of ISM ratios and MASA-to-total ratios is not described in this document in more detail. For example, WO2022 / 200666 describes a suitable MASA-to-total ratio encoding method, and GB applications 2217884.2 and 2217905.5 describe a suitable ISM ratio encoding method.
[0108] In some modes, there is an 803 object priority (audio) determinant. The 803 object priority determinant is configured to generate a priority value for objects in a time-frequency window. In some modes, the 803 object priority determinant is configured to obtain the MASA-to-total ratios 812 and the ISM ratios 814. For each time-frequency (TF) window (i.e., for each sub-band combination subframe) there is a MASA-to-total ratio and N ISM ratios, where N is the number of objects.
[0109] The priority value can, in addition, in some modes, be defined as the maximum in all time-frequency windows of the object contribution ratios.
[0110] For example, a priority value can be generated based on p(i) = max^ ((1 — m2t(b, k))r(i, b, k)), i = 0: N — 1 (3) k=0:M-1
[0111] In the formula above, there are N objects, B sub-bands, and M time subframes. m2t(b,k) represents the quantized MASA-to-total energy ratio for sub-band b and subframe k. r(i, b,k) is the quantized ISM ratio of object i, for sub-band b and subframe k. Subsequently, the quantized versions of the MASA-to-total and ISM ratios are used to have the same values available in the decoder as well.
[0112] In some embodiments, some other parameter (besides the mass-to-total-energy ratio) may be employed. For example, an object-to-total-energy ratio, which could be something like o2t(b,k) = 1-m2t(b,k), and Petition 870250085402, dated 09 / 22 / 2025, pages 253 / 283 30 / 48 would effectively carry the same information.
[0113] In such modalities where the alternative parameter is employed, then the equation above is updated accordingly. Thus, in the above object-to-total energy ratio example, the object-to-total o2t(b,k) would simply replace the part (1-m2t(b,k)).
[0114] Alternatively, in some modes, a (weighted) average of the object contributions in the TF windows may also be considered to select the priority. Furthermore, in some other modes, the max. operator may be replaced by a second or third max. value (or any similar value). In this way, the largest contribution in a single TF window would not alone determine the priority.
[0115] The 804 object priority values can then be passed to an 805 bit determiner.
[0116] In some embodiments, the audio object metadata encoder 111 comprises an 805 bit determiner. The 805 bit determiner is configured to obtain or receive the 804 object priorities and based on these object priority values determine the number of bits that can be used to encode the direction parameter. In other words, define the number of bits that define the quantization grid used to encode the 802 direction values.
[0117] In some modes, bit determiner 805 is configured to determine or assign fewer bits to objects with lower priority.
[0118] As an example, based on the object's priority, the number of bits allocated for each directional metadata of the object can be calculated as: bits(í) = 11 — [(1 — p(í)) * 6], i = 0:N — 1 (4)
[0119] The [ ] operator means rounding to the nearest integer operation. The maximum number of bits per object direction is 11 and the minimum number of bits is 4. The maximum and minimum number of bits may be other values based on implementation details and, Petition 870250085402, dated 09 / 22 / 2025, pages 254 / 283 31 / 48 Thus, the value 7 (the difference between the maximum number of bits and the minimum number of bits per object direction) can also change in some other modes. Although this example shows a linear scale, the formula described above, providing the number of bits, can be replaced by any other linear or non-linear increasing function of the object priority and ensuring that the number of bits is within a [4,11] or similar domain.
[0120] Figure 8 also shows a low bit-rate directional encoder 820 that can be applied to encoding directional values of audio objects when the encoder controller 115 determines a lower encoding rate, such as Mode B and Mode C encoding modes. This low bit-rate directional encoder 820 can be applied, instead of the spherical grid determiner and encoder 807, for encoding a directional value of an audio object. To this end, Figure 8 also depicts an encoder determiner 819 that is arranged to select between the low bit-rate encoder 820 and the spherical grid determiner and encoder 807. The selected encoder is then used to encode a directional value of an audio object.
[0121] The selection between the two different encoding schemes may be dependent on bit allocation 806 to quantize the direction value. For example, the encoder determinant 819 may be arranged to receive bit allocation 806 and test it against a predetermined threshold bit value. The test result may then be used to route audio object direction value 802 to the low-rate encoder 820 or to the spherical grid encoder and 807. Figure 8 depicts the deployment of encoder determinant 819 as being arranged to receive both bit allocation 806 and audio object direction value 802 and output audio object direction value 802 to either the low-rate encoder 820 or the spherical grid encoder and 807.
[0122] In modalities, the 819 encoder determinant may have the Petition 870250085402, dated 09 / 22 / 2025, pages 255 / 283 32 / 48 following functionality. Initially, bit allocation 806 can be inspected to determine if the number of bits allocated to quantize a particular audio object direction value is below a predetermined threshold bit value. If the number allocated is below the predetermined threshold, then the audio object direction value 802 is routed to the low-rate encoder 820 along the signal feed 832 for encoding. Otherwise, the audio object direction value 802 is routed along the signal feed 830 to the spherical grid determiner and encoder 807 for encoding. The selection between the low-rate encoder 820 and the spherical grid determiner and encoder 807 can be performed per audio object.
[0123] The functionality above the 819 encoder determiner can be exemplified by the pseudocode below. 1. For each object i a. If blts(l) < 8 I. Send non-quantized direction values to the 820 low-rate encoder. ii. End b. Otherwise L Send non-quantized direction values to the spherical grating determinant and 807 ii encoder. End c. End 2. End for
[0124] In the pseudocode above, it can be seen that the default threshold bit value has been set to 8 bits. However, it should be noted that other threshold values can be used and that these values can be obtained through experimentation.
[0125] Returning to Figure 8, when the low-rate encoder 820 is selected, the encoder is set to receive the direction value of Petition 870250085402, dated 09 / 22 / 2025, pages 256 / 283 33 / 48 audio object 802 along signal feed 832. The low-rate encoder 820 can then be arranged to compare the audio object direction value of the current frame (Azimuth and Elevation) with an audio object direction value for a previous frame through the same audio object. In the modes, this comparison can be performed using the non-quantized audio object direction values. If the result of the comparison determines that the audio object direction value of the previous frame is the same as the audio object direction value of the current frame, then the low-rate encoder 820 is arranged to signal this as a single bit to indicate that there is no change in the direction value of the previous frame. This is represented in Figure 8 as the signal bit LR_enc 824. In other words, in this case, the audio object direction value encoded for the audio object is a single bit.
[0126] Note that the comparison is performed on a direction value component basis. That is, the azimuth value of the current frame is compared to the azimuth value of the previous frame, and the elevation value of the current frame is compared to the elevation value of the previous frame. If the comparison does not register any change (or difference) for either component, then the result of the comparison determines that the audio object direction value of the previous frame is the same as the audio object direction value of the current frame.
[0127] Another arrangement can be made to have an LR_enc signal comprising a plurality of bits; such arrangements allow the difference to be determined between the components of the direction values from a previous frame to a current frame. For example, a single bit can be used to indicate whether one or both of the azimuth and elevation values of the previous frame are the same as one or both of the azimuth and elevation values of the current frame. An additional bit can be used to indicate whether one of the azimuth or elevation values is the same, or whether both azimuth and elevation values are the same from the previous frame to the current frame. If one of the azimuth or elevation values is the same, then another bit can be used to Petition 870250085402, dated 09 / 22 / 2025, pages 257 / 283 34 / 48 identify whether the azimuth or elevation value is the same from the previous frame to the current frame. This additional mode can then be used to quantify, encode, and index the angle (azimuth or elevation) that changed from the previous frame to the current frame.
[0128] If the result of the above comparison for an audio object indicates that the current frame direction value is not the same as the previous frame direction value, the low-rate encoder 820 directs the audio object direction value 802 to the spherical grid determinant and encoder 807 along the signal feed 834 for encoding. This condition can be signaled using the opposite state of the LR_enc signal bit above 824.
[0129] In terms of pseudocode, the functionality of the 820 low-rate encoder can take the following form If the unquantized elevation and azimuth are the same as for the previous table. 1. Send a bit (1) to signal that the directions are the same ii, Otherwise 1. Send a bit (0) to signal that the directions are not the same. 2. Direct audio object direction value for 807 encoder for iiLFim encoding
[0130] Figure 10 shows the processing steps of the 819 encoder determiner and the 820 low-rate encoder.
[0131] Initially, bit allocation 806 for an audio object is received by encoder determiner 819 and then checked against the threshold bit allocation value. This is shown as processing step 1001 in Figure 10. When processing step 1001 determines that the bit allocation for the audio object is less than the threshold value, the Petition 870250085402, dated 09 / 22 / 2025, pages 258 / 283 Encoder determinant 819 determines that the low-rate encoder 820 should be used to encode the audio object direction value. This is depicted in Figure 10 as the progression to processing step 1005. However, when processing step 1001 determines that the bit allocation is not less than the threshold value, encoder determinant 819 will determine that the spherical grating quantizer 807 should be used directly to encode the audio object direction value 802. This is depicted in Figure 10 as the progression to processing step 1003, resulting in the audio object direction value 802 being sent directly to the spherical grating quantizer 807 for encoding.
[0132] In processing step 1005, the low-rate encoder 820 compares the current-frame audio object direction value with a previous-frame audio object direction value in order to determine whether the audio object direction value 802 should be sent to the spherical grid quantizer 807 for encoding. If the comparison indicates that the current-frame audio object direction value is the same as the previous-frame audio object direction value, the audio object direction parameter is not sent to the spherical grid quantizer 807 for encoding. Instead, the current-frame audio object direction parameter 802 is not encoded by itself; instead, the state of the LR_enc signal bit 824 is set to a state that indicates that the current-frame audio object direction value 802 is the same as the audio object direction value for the previous frame.This is reflected in Figure 10 as the progression to step 1009.
[0133] If the comparison in step 1005 indicates that the two direction values are not the same, the low-rate encoder 820 determines that the current frame audio object direction value is sent to the spherical grating quantizer 807 for encoding. This is depicted in Figure 10 as the transition from processing step 1005 to processing step 1003 via step 1007. In step 1007, the low-rate encoder 820 is set to set the state of the LR_enc signal 824 to Petition 870250085402, dated 09 / 22 / 2025, pages 259 / 283 36 / 48 indicates that the audio object direction values of the current frame and the previous frame differ.
[0134] The low-rate encoder 820 can be configured to output the LR_Enc signal bit 824 to the bit stream generator 113, in order to be included in the bit stream 118.
[0135] As indicated above, the audio object metadata encoder 111 comprises a spherical grid determiner and encoder 807 configured to receive bit allocation 806 and direction values 802. The spherical grid determiner and encoder 807 are then configured to generate an output index value based on the nearest point relative to the direction values 802 on a given spherical grid defined by bit allocation 806 for the audio object. The encoded direction index values 808 for the audio object can then be passed to the bitstream generator 113 for inclusion in the bitstream 113.
[0136] Other modes can be configured to quantize the direction of each object, with the corresponding number of bits, and output an elevation index and an azimuth index. Object metadata can be encoded independently for each object or the directional metadata of objects can be co-encoded using, for example, some weights to correspond to the different bit allocations.
[0137] The spherical grid uses the idea of covering a sphere with smaller spheres and considering the centers of the smaller spheres as points that define a grid of nearly equidistant directions. Each point on the grid is defined by pairing an azimuth value and an elevation value.
[0138] The highest resolution spherical grid (11 bits) used for direction quantization ensures a quantization resolution of 5 degrees. The structure of the spherical grid is the same as that used for the quantization of MASA directional metadata (and the methods for defining the grid and encoding the index relative to the values of MASA directional metadata are known as discussed in PCT / EP2017 / 078948, GB1811071.8). Petition 870250085402, dated 09 / 22 / 2025, pages 260 / 283 37 / 48
[0139] In some modes where the priority of an object is exactly zero, the number of bits for that object is set to zero and no directional metadata is sent for that object.
[0140] In relation to Figure 9, a flow diagram is shown summarizing the operations of the audio object metadata encoder example 111 shown in Figure 8.
[0141] The initial operation is to receive / obtain the independent flows with metadata and determined (quantized) MASA-to-total-energy ratio (m2t) and ISM ratios (r) as shown in Figure 9 by step 901.
[0142] Next, the following operation is to determine the priority of the object based on the values of the ISM ratio (quantized) (r) and the MASA-to-total energy ratio (m2t) (quantized) as shown in Figure 9 by step 903.
[0143] Having determined the object priority, then determine the bit allocations for the direction parameter based on the determined object priority, as shown in Figure 9 by step 905.
[0144] From the bit allocations, then, determine a spherical grid to encode direction parameters for the objects, as shown in Figure 9 by step 907.
[0145] Then, generate the encoded direction index within the determined spherical grids, as shown in Figure 9 by step 909.
[0146] The encoded direction index values can then be output for inclusion in the bitstream, as shown in Figure 9 by step 911.
[0147] With regard to a decoder, the decoding of the encoded directional information can be implemented by determining a similar priority ordering and determining the associated quantization grid. Thus, for example, a pseudocode representation of encoding and decoding could be: Object metadata encoding Petition 870250085402, dated 09 / 22 / 2025, pages 261 / 283 38 / 48 1. For each object a. calculate the priority p(i) b. calculate the number of bits bits(i) 2. End for 3. For each object a. Encode the directional metadata in the spherical grid corresponding to the number of bits allocated to the object. 4. End for Decoding object metadata 1. For each object a. calculate the priority p(i) b. calculate the number of bits bits(i) 2. End for 3. For each object a. Decode the directional metadata in the spherical grid corresponding to the number of bits allocated to the object. 4. End for
[0148] With regard to Figure 11, an audio direction value decoder 1101 and a bitstream receiver and demultiplexer 1113 are shown. The bitstream receiver and demultiplexer 1113 are arranged to receive and demultiplex the encoded bitstream 118 into several encoded parameter signal streams, of which Figure 11 represents the encoded streams that are pertinent to the audio direction value decoder 1101.
[0149] The audio direction value decoder 1101 is shown as comprising a decoder determiner and spherical grid decoder 1119 and a low-rate decoder 1120. The decoder determiner and spherical grid decoder 1119 are arranged to receive the encoded direction index values 808 and the bit allocation 806 corresponding to an audio object. Petition 870250085402, dated 09 / 22 / 2025, pages 262 / 283 39 / 48
[0150] The allocation of bit 806 for an audio object i can be determined locally in the decoder by setting the priority value p(0) for the audio object. The priority value can be determined from the ISM ratio for the audio object. The ISM ratio is sent to the decoder as part of the bit stream 118.
[0151] The decoder determiner and spherical grid decoder 1119 are initially deployed (for each audio object) to compare the bit allocation 806 for the audio object with the predetermined threshold bit value in order to determine whether the decoded direction values 1108 are obtained by directly decoding the encoded direction index or by using the functionality of the low-rate decoder 1120. In other words, when the bit allocation 806 is below the threshold, the low-rate decoder 1120 is selected to form the decoded direction values 1108. This is depicted in Figure 11 as the selected signal feed LR decoder 1118. Conversely, when the bit allocation 806 is not below the threshold, the decoded direction values 1108 are generated directly by the decoder determiner and spherical grid decoder 1119.
[0152] The above functionality of the 1119 spherical grid encoder and decoder determinator can be exemplified by the pseudocode below. For each object i If bits{i) < 8 i. Select low-rate decoder ii. End Otherwise i. Decode the encoded direction index value ii. End The end if Petition 870250085402, dated 09 / 22 / 2025, pages 263 / 283 40 / 48
[0153] When the 1120 low-rate decoder is selected to decode audio direction values for an audio object, the 1120 low-rate decoder is set to initially inspect the value of the LR_enc_signal bit 824. In the case where the state of the LR_enc_signal bit 824 indicates that there is no change in the direction value [for the audio object] between the previous audio frame and the current audio frame (LR_enc_signal = 1), the 1120 low-rate decoder will simply output the direction value of the previous frame [for the audio object] as the direction value of the audio object for the current frame.
[0154] It should be noted that the spherical grid decoder located in the spherical grid decoder and decoder determinator 1119 is arranged to decode the encoded direction index values of the audio object according to the 806 bit allocation for the audio object.
[0155] On the other hand, in the case where bit LR_enc_signal 824 indicates that the direction values for the audio object between the previous frame and the current frame differ, the low-rate decoder 1120 is arranged to operate in a different way and, in this respect, Figure 12 represents the processing performed by the low-rate decoder 1120.
[0156] First, the low-rate decoder 1120 is configured to receive decoded direction values 0 and Θ for the current frame of the decoder determinant and spherical grid decoder 1119. This is shown as processing step 1201.
[0157] The low-rate decoder 1120 then checks if the multiple of the current frame quantized azimuth value (direction)Θ and the previous frame quantized azimuth value 0 are above zero. This is shown as processing step 1203 in Figure 12.
[0158] When the check in processing step 1203 indicates that the product of φ $previous > 0, the low-rate encoder 1120 is then configured to move to processing step 1205 where the Petition 870250085402, dated 09 / 22 / 2025, pages 264 / 283 41 / 48 azimuth quantization resolution is determined. This can be determined as the distance between azimuth points around the circumference of a circle on the spherical grid, where the circle is given by the elevation value Θ. The azimuth value resolution can be found as ΔΦ = 360 / ηθ (5)
[0159] Where ηθ is the number of azimuth points around the circumference of the circle for elevation Θ and coincides with the number of azimuth values in the codebook corresponding to elevation Θ.
[0160] Once the resolution is found, the process performs an additional check in step 1207 to determine if the difference between the current frame quantized azimuth value (direction)Θ and the previous frame quantized azimuth value 0previous is greater than half the azimuth resolution Δφ / 2 (φ - $previous > Δφ / 2)). If this is the case, the current frame azimuth value is defined as the difference between the current frame azimuth value and half the azimuth resolution, φ = φ - Δφ / 2. This is shown as processing step 1209 in Figure 12.
[0161] However, if the verification condition in processing step 1207 is not met, then the process proceeds to an additional verification in processing step 1211. In this step, the difference between the previous frame quantized azimuth value (direction) a>anteOior and the current frame quantized azimuth value Θ is taken and the result is checked to determine if it is greater than half the azimuth resolution Δφ / 2 (0previous - φ > Δφ / 2). If this condition is met, the current frame azimuth value is defined as the sum of the current frame azimuth value and half the azimuth resolution, = φ + Δφ / 2. This is shown as processing step 1213 in Figure 12.
[0162] In summary. The overall effect of the processing steps in Figure 12 is that, in the case where the state of the LR_enc_signal bit indicates that there is a Petition 870250085402, dated 09 / 22 / 2025, pages 265 / 283 42 / 48 change in direction values [for the audio object] between the previous audio frame and the current audio frame (LR_enc_signal = 0, Figure 12), the 1120 low-rate decoder will produce as decoded output direction values comprising the decoded elevation value Θ and the decoded azimuth value φ that has been changed by adding the value Δψ / 2 or subtracting the value Δφ / 2. When the LR_enc_signal bit state indicates that there is no change in direction values [for the audio object] between the previous audio frame and the current audio frame (LR_enc_signal = 1), the 1120 low-rate decoder will produce as decoded output direction values comprising the decoded direction values of the previous audio frame.
[0163] The function of the 1120 low-rate decoder can also be exemplified by the following pseudocode. i, Read one bit, LR_enc_signal bit, from bit stream ií. If LR_enc_sígnal == 1 1. The direction is the same as for the previous frame. Otherwise 1. Receive direction index 2. Decode the direction index into φ and  3, If φ > 0 a. Δψ= 360 / ¾ b. If i - Φ = Φ - 4ψ / 2 c. Otherwise if $n,líf,í0J- φ > i- Φ = Φ + Δψ / 2 d. End 4. Close if iv Close if
[0164] With regard to Figure 10, an electronic device is shown of Petition 870250085402, dated 09 / 22 / 2025, pages 266 / 283 43 / 48 example that can be used as any of the parts of the system apparatus, as described above. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc. The device can, for example, be configured to implement the encoder / analyzer part and / or the decoder part as shown in Figure 1 or any functional block as described above.
[0165] In some embodiments, the 1400 device comprises at least one 1407 processor or central processing unit. The 1407 processor can be configured to execute various program codes, such as the methods described in this document.
[0166] In some embodiments, the 1400 device comprises at least one 1411 memory. In some embodiments, at least one 1407 processor is coupled to the 1411 memory. The 1411 memory may be any suitable storage medium. In some embodiments, the 1411 memory comprises a program code section for storing program code implementable in the 1407 processor. In addition, in some embodiments, the 1411 memory may further comprise a stored data section for storing data, for example, data that has been processed or is to be processed in accordance with the embodiments described in this document. The implemented program code stored in the program code section and the data stored in the stored data section may be retrieved by the 1407 processor whenever necessary through memory processor coupling.
[0167] In some embodiments, the 1400 device comprises a 1405 user interface. The 1405 user interface may be coupled in some embodiments to the 1407 processor. In some embodiments, the 1407 processor may control the operation of the 1405 user interface and receive inputs from the 1405 user interface. In some embodiments, the Petition 870250085402, dated 09 / 22 / 2025, pages 267 / 283 44 / 48 User interface 1405 may allow a user to input commands to device 1400, for example, via a keyboard. In some embodiments, user interface 1405 may allow the user to obtain information from device 1400. For example, user interface 1405 may comprise a screen configured to display information from device 1400 to the user. User interface 1405 may, in some embodiments, comprise a touch-sensitive screen or touch interface capable of allowing information to be entered into device 1400 and displaying further information to the user from device 1400. In some embodiments, user interface 1405 may be the user interface for communication.
[0168] In some embodiments, the device 1400 comprises an input / output port 1409. The input / output port 1409 in some embodiments comprises a transceiver. The transceiver in such embodiments may be coupled to the processor 1407 and configured to allow communication with other electronic devices or apparatus, for example, via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiving medium may, in some embodiments, be configured to communicate with other electronic devices or apparatus by wired or wireless coupling.
[0169] The transceiver can communicate with another device by any suitable known communication protocol. For example, in some embodiments, the transceiver may use a suitable radio access architecture based on long term evolution advanced (LTE Advanced, LTE-A) or new radio (NR) (or may be referred to as 5G), universal mobile telecommunications system (UMTS) radio access network (UTRAN or E-UTRAN), long term evolution (LTE, the same as E-UTRA), 2G networks (legacy network technology), wireless local area network (WLAN or Wi-Fi), worldwide interoperability for microwave access (WiMAX), Bluetooth®, personal communication services (PCS), ZigBee®, wideband code division multiple access (WCDMA), systems using ultra-wideband (UWB) technology, sensor networks, Petition 870250085402, dated 09 / 22 / 2025, pages 268 / 283 45 / 48 mobile ad-hoc networks (MANETs), cellular internet of things (IoT) RAN and internet protocol multimedia subsystem (IMS), any other suitable option and / or any combination thereof.
[0170] The input / output port of the 1409 transceiver can be configured to receive signals.
[0171] In some embodiments, the 1400 device may be employed as at least part of the synthesis device. The 1409 input / output port may be coupled to headphones (which may be headphones with or without tracking) or similar and loudspeakers.
[0172] In general, the various embodiments of the invention can be implemented in special-purpose hardware or circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device, although the invention is not limited to these. Although various aspects of the invention may be illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it is well understood that these blocks, devices, systems, techniques, or methods described in this document may be implemented, as non-limiting examples, in special-purpose hardware, software, firmware, circuits or logic, general-purpose hardware or controller, or other computing devices, or some combination thereof.
[0173] The embodiments of this invention can be implemented by computer software executable by a mobile device data processor, such as in the processor entity, or by hardware, or by a combination of software and hardware. Furthermore, in this respect, it should be noted that any blocks of the logic flow, for example, as in the Figures, can represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits. Petition 870250085402, dated 09 / 22 / 2025, pages 269 / 283 46 / 48 blocks and functions. Software can be stored on physical media such as memory chips or memory blocks implemented within the processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and data variants thereof, CDs.
[0174] Memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. Data processors may be of any type suitable to the local technical environment and may include one or more general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits, and multi-core processor architecture-based processors, as non-limiting examples.
[0175] The embodiments of the inventions can be practiced in various components, such as integrated circuit modules. The design of integrated circuits is, in general, a highly automated process. Complex and powerful software tools are available to convert a logic-level design into a semiconductor circuit design ready to be etched and formed onto a semiconductor substrate.
[0176] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design of San Jose, California, automatically route leads and locate components on a semiconductor chip using well-established design rules as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resulting design, in a standardized electronic format (e.g., Opus, GDSII, or similar), can be transmitted to a semiconductor fabrication facility or fab for manufacturing. Petition 870250085402, dated 09 / 22 / 2025, pages 270 / 283 47 / 48
[0177] As used in this application, the term circuit may refer to one, more, or all of the following: (a) hardware-only circuit implementations (such as analog-only and / or digital-only circuit implementations) and (b) combinations of hardware and software circuits, such as (as applicable): (i) a combination of analog and / or digital hardware circuits with software / firmware and (ii) any portions of hardware processors with software (including digital signal processors), software and memories that work together to enable a device, such as a mobile phone or server, to perform various functions) and hardware circuits and / or processors, such as microprocessors or a portion of microprocessors, that require software (e.g., firmware) for operation, but the software may not be present when it is not required for operation.
[0178] This definition of circuit applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuit also encompasses an implementation of merely a hardware circuit or processor (or multiple processors) or part of a hardware circuit or processor and its accompanying software and / or firmware. The term circuit also encompasses, for example and if applicable to the specific claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device or other computing or networking device.
[0179] The term non-transient, as used in this document, is a limitation of the medium itself (i.e., tangible, not a signal) rather than a limitation in the persistence of data storage (e.g., RAM vs. ROM).
[0180] As used in this document, at least one of the following: Petition 870250085402, dated 09 / 22 / 2025, pp. 271 / 283 48 / 48 and at least one of and similar wording, in which the list of two or more elements is joined by "and" or "or," means at least any one of the elements, or at least any two or more of the elements, or at least all the elements.
[0181] The foregoing description has provided, by way of exemplary and non-limiting examples, a complete and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant art in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. Nevertheless, all such and similar modifications of the teachings of this invention will still be within the scope of this invention as defined in the appended claims. Petition 870250085402, dated 09 / 22 / 2025, pp. 272 / 283
Claims
1 / 11 CLAIMS 1. Apparatus characterized in that it comprises means configured to: receive a direction value for a time period of an audio object; compare, as a first comparison, a bit allocation with a threshold bit allocation value; depending on the first comparison, quantize the direction value for the time period of the audio object with a quantizer according to the bit allocation or compare, as a second comparison, the direction value for the time period of the audio object with a direction value for a previous time period of the audio object; and depending on the second comparison, quantize the direction value for the time period with the quantizer according to the bit allocation and signal the second comparison.
2. Device according to Claim 1, characterized in that the means configured to, depending on the second comparison, quantize the direction value for the time period with the quantizer according to the bit allocation and signal that the comparison are configured to: quantize the direction value for the time period with the quantizer according to the bit allocation when the second comparison indicates that the direction value for the time period of the audio object is different from the direction value for the previous time period of the audio object and set a signal flag to indicate that the direction value for the time period differs from the direction value for the previous time period;and define a signal flag to indicate that the direction value for the time period is the same as the direction value for the previous time period when the second comparison indicates that the direction value for the audio object's time period is the same as the direction value for the previous time period of the audio object.
3. Apparatus, according to any one of Claims 1 or 2, characterized in that the means configured to, depending on the first comparison, quantize the direction value for the time period of the audio object with a quantizer according to the bit allocation or compare, as a second comparison, the direction value for the time period of the audio object with a direction value for a previous time period of the audio object is configured to: quantize the direction value for the time period of the audio object with a quantizer according to the bit allocation when the bit allocation for quantizing the direction value is not less than the threshold bit allocation value;and compare the direction value for the audio object's time period with a direction value for a previous time period of the audio object as the second comparison when the bit allocation for the direction value quantization is less than the threshold bit allocation value.
4. Apparatus, according to any one of Claims 1 to 3, characterized in that the quantizer is a spherical lattice quantizer wherein the spherical lattice is formed by covering the sphere with smaller spheres, wherein the smaller spheres define the points of the spherical lattice.
5. Device according to any one of Claims 1 to 4, characterized in that the direction value comprises an azimuth value and an elevation value.
6. Device, according to Claim 5, characterized in that the second comparison is performed on a direction value component basis, comprising that the azimuth value of the direction value for the time period of the audio object is compared to an azimuth value of the direction value for the previous time period of the audio object and that the elevation value of the direction value for the time period of the audio object is compared to an elevation value of the direction value for the previous time period of the audio object.
7. Device according to any one of Claims 1 to 6, characterized in that the means is additionally configured to: generate a priority value for the audio object; determine bit allocation based on the priority value for the audio object.
8. Apparatus, according to Claim 7, characterized in that the means configured to determine bit allocation based on the priority value for the audio object are configured to: determine the bit allocation bits (i) based on the priority value p(i) for the audio object i according to an increasing linear or non-linear function of the priority value p(i), ensuring that the bit allocation is within a fixed domain between a minimum number of bits and a maximum number of bits.
9. Device according to any one of Claims 1 to 8, characterized in that the threshold bit allocation value is predetermined to be 8 bits.
10. Apparatus, according to any one of Claims 1 to 9, characterized in that the bit allocation is a number of bits that defines the quantization grid used to encode the direction value.
11. Apparatus characterized in that it comprises means configured to: compare a bit allocation for a direction value index for a time period of an audio object with a threshold bit allocation value; depending on the comparison, decode the direction value index for the time period of the audio object with a dequantizer according to the bit allocation or read a state of a signal flag associated with the direction value index for the time period of the audio object; and depending on the state of the signal flag; Petition 870250085402, dated 09 / 22 / 2025, p.275 / 283 4 / 11 define a quantized direction value for the audio object's time period to be a quantized direction value for a previous time period of the audio object, or decode the direction value index for the audio object's time period with a dequantizer according to bit allocation, generating a quantized direction value for the audio object's time period, and adjust an azimuth value of the quantized direction value for the audio object's time period according to a quantization resolution of the azimuth value.
12. Apparatus, according to Claim 11, characterized in that the means configured to adjust an azimuth value of the quantized direction value for the time period of the audio object according to a quantization resolution of the azimuth value are configured to: compare the product of the azimuth value of the quantized direction value for the time period and an azimuth value of the quantized direction value for the previous time period of the audio object and when the product is greater than zero: determine a difference between two consecutive azimuth quantization values of the dequantizer;When a difference between the azimuth value of the quantized direction value for the time period and the azimuth value of the quantized direction value for the previous time period of the audio object is greater than half the difference between two consecutive azimuth quantization values of the dequantizer, subtract half the difference between two consecutive azimuth quantization values of the dequantizer from the azimuth value of the quantized direction value for the time period of the audio object;and when a difference between the azimuth value of the quantized direction value for the previous time period and the azimuth value of the quantized direction value for the audio object's time period is greater than half the difference between two consecutive azimuth quantization values of the dequantizer, add half the difference between two consecutive azimuth quantization values of the dequantizer to the azimuth value of the quantized direction value for the audio object's time period.
13. Apparatus, according to Claim 12, characterized in that the means configured to determine a difference between two consecutive azimuth quantization values of the dequantizer are configured to: divide the circumference of a circle by a number of azimuth values, where the number of azimuth values is determined by the elevation value of the quantized direction value for the time period.
14. Apparatus, according to any one of Claims 11 to 13, characterized in that the state of the signal indicator shows one of: a direction value for the time period of the audio object differs from a direction value for the previous time period of the audio object; or a direction value for the time period of the audio object is the same as a direction value for the previous time period of the audio object.
15. Apparatus, according to any one of Claims 11 to 14, characterized in that the dequantizer is a spherical lattice quantizer wherein the spherical lattice is formed by covering the sphere with smaller spheres, wherein the smaller spheres define the points of the spherical lattice.
16. Device according to any one of Claims 11 to 15, characterized in that the quantized direction value comprises an azimuth value and an elevation value.
17. Device according to any one of Claims 11 to 16, characterized in that the means is additionally configured to: generate a priority value for the audio object; determine bit allocation based on the priority value for the audio object.
18. Device, according to Claim 17, characterized by the fact Petition 870250085402, dated 09 / 22 / 2025, pp. 277 / 283 6 / 11 that the means configured to determine bit allocation based on the priority value for the audio object are configured to: determine the bit allocation bits (i) based on the priority value p(i) for the audio object i according to an increasing linear or non-linear function of the priority value p(i), ensuring that the bit allocation is within a fixed domain between a minimum number of bits and a maximum number of bits.
19. Device according to any one of Claims 11 to 18, characterized in that the threshold bit allocation value is predetermined to be 8 bits.
20. Apparatus, according to any one of Claims 11 to 19, characterized in that the bit allocation is a number of bits that defines the quantization grid used to decode the direction value index.
21. A method characterized by the fact that it comprises: receiving a direction value for a time period of an audio object; comparing, as a first comparison, a bit allocation with a threshold bit allocation value; depending on the first comparison, quantizing the direction value for the time period of the audio object with a quantizer according to the bit allocation or comparing, as a second comparison, the direction value for the time period of the audio object with a direction value for a previous time period of the audio object; and depending on the second comparison, quantizing the direction value for the time period with the quantizer according to the bit allocation and signaling the second comparison.
22. Method, according to Claim 21, characterized by the fact that depending on the second comparison, quantize the direction value for the time period with the quantizer according to the bit allocation and sign of Petition 870250085402, dated 09 / 22 / 2025, p.278 / 283 7 / 11 comparison includes: quantizing the direction value for the time period with the quantizer according to the bit allocation when the second comparison indicates that the direction value for the time period of the audio object is different from the direction value for the previous time period of the audio object and setting a signal flag to indicate that the direction value for the time period differs from the direction value for the previous time period; and setting a signal flag to indicate that the direction value for the time period is the same as the direction value for the previous time period when the second comparison indicates that the direction value for the time period of the audio object is the same as the direction value for the previous time period of the audio object.
23. A method according to either of Claims 21 or 22, characterized in that, depending on the first comparison, quantizing the direction value for the time period of the audio object with a quantizer according to bit allocation or comparing, as a second comparison, the direction value for the time period of the audio object with a direction value for a previous time period of the audio object comprises: quantizing the direction value for the time period of the audio object with a quantizer according to bit allocation when the bit allocation for quantizing the direction value is not less than the threshold bit allocation value; and comparing the direction value for the time period of the audio object with a direction value for a previous time period of the audio object as the second comparison when the bit allocation for quantizing the direction value is less than the threshold bit allocation value.
24. Method, according to any one of Claims 21 to 23, characterized in that the quantizer is a spherical lattice quantizer in which the spherical lattice is formed by covering the sphere with smaller spheres, in Petition 870250085402, dated 09 / 22 / 2025, pp. 279 / 283 8 / 11, where the smaller spheres define the points of the spherical lattice.
25. A method, according to any one of Claims 21 to 24, characterized in that the direction value comprises an azimuth value and an elevation value.
26. Method, according to Claim 25, characterized in that the second comparison is performed on a direction value component basis, comprising that the azimuth value of the direction value for the time period of the audio object is compared to an azimuth value of the direction value for the previous time period of the audio object and that the elevation value of the direction value for the time period of the audio object is compared to an elevation value of the direction value for the previous time period of the audio object.
27. A method, according to any one of Claims 21 to 26, characterized in that it comprises: generating a priority value for the audio object; determining bit allocation based on the priority value for the audio object.
28. A method according to Claim 27, characterized in that determining bit allocation based on priority value for the audio object comprises: determining the bit allocation bits (i) based on priority value p(i) for audio object i according to an increasing linear or nonlinear function of priority value p(i), ensuring that the bit allocation is within a fixed domain between a minimum number of bits and a maximum number of bits.
29. A method, according to any one of Claims 21 to 28, characterized in that the threshold bit allocation value is predetermined to be 8 bits.
30. Method, according to any of Claims 21 to 29, characterized in that the bit allocation is a number of bits that defines Petition 870250085402, dated 09 / 22 / 2025, pp. 280 / 283 9 / 11 the quantization grid used to encode the direction value.
31. A bitstream characterized in that it comprises, for a period of time, for each of a plurality of audio objects, a respective encoded direction value, the respective direction value being encoded according to the method of any one of Claims 21 to 30.
32. Bitstream, according to Claim 31, characterized in that the bitstream comprises for the time period: - an encoded object identifier of an audio object selected from the plurality of audio objects; - an encoded audio signal associated with the selected audio object; - an encoded downmix consisting of metadata-assisted spatial audio, MASA, transport audio signals and transport audio signals associated with unselected objects from the plurality of objects; - encoded MASA metadata; - encoded MASA-total ratios; - an independent stream encoded with metadata, ISM, ratios determined for each object in the plurality of objects; and - encoded object metadata for each object in the plurality of audio objects, wherein the encoded object metadata comprises the encoded direction values for the time period.
33. A method characterized in that it comprises: comparing a bit allocation for a direction value index for a time period of an audio object with a threshold bit allocation value; depending on the comparison, decoding the direction value index for the time period of the audio object with a dequantizer according to the bit allocation or reading a state of a signal flag associated with the direction value index for the time period of the audio object; and depending on the state of the signal flag; Petition 870250085402, dated 09 / 22 / 2025, p.281 / 283 10 / 11 define a quantized direction value for the audio object's time period to be a quantized direction value for a previous time period of the audio object, or decode the direction value index for the audio object's time period with a dequantizer according to bit allocation, generating a quantized direction value for the audio object's time period, and adjust an azimuth value of the quantized direction value for the audio object's time period according to a quantization resolution of the azimuth value.
34. A method according to Claim 33, characterized in that adjusting an azimuth value of the quantized direction value for the time period of the audio object according to a quantization resolution of the azimuth value comprises: comparing the product of the azimuth value of the quantized direction value for the time period and an azimuth value of the quantized direction value for the previous time period of the audio object and when the product is greater than zero: determining a difference between two consecutive azimuth quantization values of the dequantizer;When a difference between the azimuth value of the quantized direction value for the time period and the azimuth value of the quantized direction value for the previous time period of the audio object is greater than half the difference between two consecutive azimuth quantization values of the dequantizer, subtract half the difference between two consecutive azimuth quantization values of the dequantizer from the azimuth value of the quantized direction value for the time period of the audio object;and when a difference between the azimuth value of the quantized direction value for the previous time period and the azimuth value of the quantized direction value for the audio object's time period is greater than half the difference between two consecutive azimuth quantization values of the dequantizer (Petition 870250085402, 09 / 22 / 2025, pp. 282 / 283 11 / 11) add half the difference between two consecutive azimuth quantization values of the dequantizer to the azimuth value of the quantized direction value for the audio object's time period.
35. A method, according to Claim 34, characterized in that determining a difference between two consecutive azimuth quantization values of the dequantizer comprises: dividing the circumference of a circle by a number of azimuth values, where the number of azimuth values is determined by the elevation value of the quantized direction value for the time period.
36. A method, according to Claims 33 to 35, characterized in that the state of the signal flag indicates one of: a direction value for the time period of the audio object differs from a direction value for the previous time period of the audio object; or a direction value for the time period of the audio object is the same as a direction value for the previous time period of the audio object.
37. Method, according to Claims 33 to 36, characterized in that the quantizer is a spherical lattice quantizer wherein the spherical lattice is formed by covering the sphere with smaller spheres, wherein the smaller spheres define the points of the spherical lattice. Petition 870250085402, dated 09 / 22 / 2025, pp. 283 / 283