Low coding rate parameter space audio coding

The multirate coding model with spherical grid quantizers optimizes spatial audio encoding for immersive codecs, addressing bitrate variability and format complexity to maintain audio quality.

JP2026511173APending Publication Date: 2026-04-10NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2024-02-13
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing immersive audio codecs face challenges in efficiently encoding spatial audio data at varying bitrates, particularly in handling multiple audio formats and maintaining audio quality under different transmission conditions.

Method used

A multirate coding model is employed to adapt quantization resolution based on available bitrates, using spherical grid quantizers to encode spatial metadata, and a device that adjusts encoding based on bit allocation thresholds and signal flags to optimize encoding of audio object orientations.

Benefits of technology

This approach enables efficient encoding of spatial audio data across varying bitrates, supporting multiple audio formats and maintaining audio quality by prioritizing encoding based on available bitrates and object stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026511173000001_ABST
    Figure 2026511173000001_ABST
Patent Text Reader

Abstract

A device configured to: receive a direction value for an audio object with respect to a time frame; compare a bit assignment with a bit assignment threshold as a first comparison; quantize the direction value for an audio object with respect to a time frame using a quantizer according to the bit assignment in response to the first comparison, or compare the direction value for an audio object with respect to a previous time frame as a second comparison; quantize the direction value for a time frame using a quantizer according to the bit assignment in response to the second comparison, and signal the second comparison.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to an apparatus and method for spatial audio representation and encoding, but is not limited to audio representation of an audio encoder. [Background technology]

[0002] Parameter-space audio processing is a field of audio signal processing in which the spatial aspect of sound is described using a set of parameters. For example, in parameter-space audio capture from a microphone array, it is a common and effective choice to estimate a set of parameters from the microphone array signal, such as the direction of the sound within a frequency band and the ratio between the directional and omnidirectional parts of the captured sound within that frequency band. These parameters are known to well describe the perceptual spatial characteristics of the sound captured at the microphone array's position. Accordingly, these parameters can be used for spatial sound synthesis, binaural for headphones, loudspeakers, or other formats such as ambisonics.

[0003] Therefore, the direction and the direct-to-total energy ratio within the frequency band are particularly effective parameterizations for capturing spatial audio.

[0004] A parameter set consisting of directional parameters within the frequency band and energy ratio parameters within the frequency band (indicating sound directivity) can also be used as spatial metadata for an audio codec (which may also include other parameters such as surround coherence, spread coherence, number of directions, and distance). For example, these parameters can be estimated from audio signals captured by a microphone array, and stereo or mono signals can be generated from the microphone array signal and transmitted along with the spatial metadata. Stereo signals can be encoded using, for example, an AAC encoder, and mono signals can be encoded using an EVS encoder. A decoder can decode the audio signal into a PCM signal and process the sound within the frequency band (using the spatial metadata) to obtain a spatial output, such as a binaural output.

[0005] Immersive audio codecs are implemented to support a wide range of operating points, from low-bitrate operation to transparency. An example of such a codec is the Immersive Speech and Audio Services (IVAS) codec, which is designed for use over communication networks such as 3GPP 4G / 5G networks, including use in immersive services such as immersive speech and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general-purpose audio. Furthermore, it is expected to support channel-based audio and scene-based audio inputs, including spatial information about the sound field and sound source. The codec is also expected to operate with low latency to enable conversational services and to support high error tolerance under various transmission conditions.

[0006] The aforementioned immersive audio codecs are particularly well-suited for encoding spatial audio captured from microphone arrays (e.g., mobile phones, VR cameras, standalone microphone arrays). However, such encoders may also accept other input types, such as loudspeaker signals, audio object signals, and ambisonic signals. [Overview of the project]

[0007] According to a first embodiment, there is a device provided that includes means configured to: receive a direction value for an audio object with respect to a time frame; compare a bit assignment with a bit assignment threshold as a first comparison; quantize the direction value for an audio object with respect to a time frame using a quantizer according to the bit assignment in response to the first comparison, or compare the direction value for an audio object with respect to a previous time frame as a second comparison; and quantize the direction value for a time frame using a quantizer according to the bit assignment in response to the second comparison and signal the second comparison.

[0008] A means configured to signal a comparison by quantizing the direction value for a time frame using a quantizer according to bit assignments in response to a second comparison may be configured to: if the second comparison indicates that the direction value for a time frame of the audio object is different from the direction value for the previous time frame of the audio object, quantize the direction value for a time frame using a quantizer according to bit assignments and set a signal flag to indicate that the direction value for a time frame is different from the direction value for the previous time frame; and if the second comparison indicates that the direction value for a time frame of the audio object is the same as the direction value for the previous time frame of the audio object, set a signal flag to indicate that the direction value for a time frame is the same as the direction value for the previous time frame.

[0009] A means configured to perform, in response to a first comparison, quantize the direction value of an audio object with respect to a time frame using a quantizer according to bit allocation, or, as a second comparison, compare the direction value of an audio object with respect to a previous time frame of the audio object, may be configured to: if the bit allocation for quantizing the direction value is greater than or equal to a bit allocation threshold, quantize the direction value of an audio object with respect to a time frame using a quantizer according to bit allocation; and if the bit allocation for quantizing the direction value is less than a bit allocation threshold, as a second comparison, compare the direction value of an audio object with respect to a previous time frame of the audio object.

[0010] The quantizer may be a spherical grid quantizer, which is formed by covering a sphere with smaller spheres, and the smaller spheres define points on the spherical grid.

[0011] Directional values ​​may include azimuth and elevation values.

[0012] According to a second embodiment, there is provided an apparatus comprising means configured to: compare the bit assignment of the direction value index for an audio object with a bit assignment threshold; in response to the comparison, decode the direction value index for an audio object with an inverse quantizer according to the bit assignment, or read the start of a signal flag associated with the direction value index for an audio object with a time frame; and, depending on the state of the signal flag, set the quantized direction value for an audio object with a time frame to the quantized direction value for the previous time frame of the audio object, or decode the direction value index for an audio object with an inverse quantizer according to the bit assignment to provide a quantized direction value for an audio object with a time frame, and adjust the azimuth value of the quantized direction value for an audio object with respect to the azimuth value according to the quantization resolution of the azimuth value.

[0013] Means configured to adjust the azimuth value of the quantized direction value of an audio object for a time frame according to the quantization resolution of the azimuth value include: comparing the product of the azimuth value of the quantized direction value of the audio object for a time frame and the azimuth value of the quantized direction value for the previous time frame, and when the product is greater than zero: determining the difference between two consecutive azimuth quantized values ​​of the inverse quantizer; and determining that the difference between the azimuth value of the quantized direction value of the audio object for a time frame and the azimuth value of the quantized direction value for the previous time frame is the difference between two consecutive azimuth quantized values ​​of the inverse quantizer The inverse quantizer may be configured to: subtract half the difference between two consecutive azimuth quantized values ​​of the inverse quantizer from the azimuth value of the quantized direction value for the time frame of the audio object when the difference between the azimuth value of the quantized direction value for the previous time frame and the azimuth value of the quantized direction value for the time frame of the audio object when the difference between the azimuth value of the quantized direction value for the time frame of the audio object is greater than half the difference between two consecutive azimuth quantized values ​​of the inverse quantizer; and add half the difference between two consecutive azimuth quantized values ​​of the inverse quantizer to the azimuth value of the quantized direction value for the time frame of the audio object.

[0014] A means configured to determine the difference between two consecutive azimuth quantized values ​​of an inverse quantizer may be configured to divide the circumference by a number of azimuth values, where the number of azimuth values ​​is determined by the elevation value of the quantized directional value with respect to the time frame.

[0015] The signal flag status indicates one of the following: the direction value of the audio object relative to the time frame is different from the direction value of the audio object relative to the previous time frame; or the direction value of the audio object relative to the time frame is the same as the direction value of the audio object relative to the previous time frame.

[0016] The inverse quantizer is a spherical grid quantizer, where the spherical grid is formed by covering a sphere with smaller spheres, and the smaller spheres define the points of the spherical grid.

[0017] A third embodiment provides a method comprising: receiving a direction value for an audio object with respect to a time frame; comparing a bit allocation with a bit allocation threshold as a first comparison; quantizing the direction value for an audio object with respect to a time frame using a quantizer according to the bit allocation in response to the first comparison, or comparing the direction value for an audio object with respect to a previous time frame of the audio object as a second comparison; and quantizing the direction value for a time frame using a quantizer according to the bit allocation in response to the second comparison and signaling the second comparison.

[0018] In response to the second comparison, signaling the comparison by quantizing the direction value for a time frame using a quantizer according to the bit assignment includes: when the second comparison indicates that the direction value for a time frame of the audio object is different from the direction value for the previous time frame of the audio object, quantizing the direction value for a time frame using a quantizer according to the bit assignment and setting a signal flag to indicate that the direction value for a time frame is different from the direction value for the previous time frame; and when the second comparison indicates that the direction value for a time frame of the audio object is the same as the direction value for the previous time frame of the audio object, setting a signal flag to indicate that the direction value for a time frame is the same as the direction value for the previous time frame.

[0019] In response to the first comparison, quantizing the direction value of the audio object with respect to the time frame using a quantizer according to the bit allocation, or as a second comparison, comparing the direction value of the audio object with respect to the direction value of the audio object with respect to the previous time frame, may include: quantizing the direction value of the audio object with respect to the time frame using a quantizer according to the bit allocation when the bit allocation for quantizing the direction value is greater than or equal to the bit allocation threshold; and as a second comparison, comparing the direction value of the audio object with respect to the direction value of the time frame of the audio object with respect to the previous time frame of the audio object when the bit allocation for quantizing the direction value is less than the bit allocation threshold.

[0020] The quantizer may be a spherical grid quantizer, which is formed by covering a sphere with smaller spheres, and the smaller spheres define points on the spherical grid.

[0021] Directional values ​​may include azimuth and elevation values.

[0022] A fourth aspect of the invention provides a method comprising: comparing the bit assignment of an audio object's orientation index to a time frame with a bit assignment threshold; in response to the comparison, decoding the audio object's orientation index to a time frame using an inverse quantizer according to the bit assignment, or reading the start of a signal flag associated with the audio object's orientation index to a time frame; and, depending on the state of the signal flag, setting the quantized orientation value of the audio object to a time frame to the quantized orientation value of the previous time frame of the audio object, or decoding the audio object's orientation index to a time frame using an inverse quantizer according to the bit assignment, to provide a quantized orientation value of the audio object to a time frame, and adjusting the azimuth value of the quantized orientation value of the audio object to a time frame according to the quantization resolution of the azimuth value.

[0023] Adjusting the azimuth value of the quantized direction value for the time frame of the audio object according to the quantization resolution of the azimuth value may include: comparing the product of the azimuth value of the quantized direction value for the time frame of the audio object and the azimuth value of the quantized direction value for the previous time frame; and when the product is greater than zero: determining the difference between two consecutive azimuth quantization values of the inverse quantizer; and when the difference between the azimuth value of the quantized direction value for the time frame of the audio object and the azimuth value of the quantized direction value for the previous time frame is greater than half of the difference between two consecutive azimuth quantization values of the inverse quantizer, subtracting half of the difference between two consecutive azimuth quantization values of the inverse quantizer from the azimuth value of the quantized direction value for the time frame of the audio object; and when the difference between the azimuth value of the quantized direction value for the previous time frame of the audio object and the azimuth value of the quantized direction value for the time frame is greater than half of the difference between two consecutive azimuth quantization values of the inverse quantizer, adding half of the difference between two consecutive azimuth quantization values of the inverse quantizer to the azimuth value of the quantized direction value for the time frame of the audio object.

[0024] Determining the difference between two consecutive azimuth quantization values of the inverse quantizer includes dividing the circumference by the number of azimuth values, and the number of azimuth values is determined by the elevation value of the quantized direction value for the time frame.

[0025] The state of the signal flag may indicate one of: the direction value for the time frame of the audio object is different from the direction value for the previous time frame of the audio object; or the direction value for the time frame of the audio object is the same as the direction value for the previous time frame of the audio object.

[0026] The inverse quantizer may be a spherical grid quantizer, and the spherical grid is formed by covering the sphere with smaller spheres, and the smaller spheres define the points of the spherical grid.

[0027] According to a fifth aspect, there is provided an apparatus including at least one processor and at least one memory storing instructions, which, when executed by the at least one processor, cause the system to at least: receive a direction value for a time frame of an audio object; as a first comparison, compare a bit allocation with a bit allocation threshold; in response to the first comparison, quantize the direction value for the time frame of the audio object using a quantizer according to the bit allocation, or as a second comparison, compare the direction value for the time frame of the audio object with the direction value for the previous time frame of the audio object; and in response to the second comparison, quantize the direction value for the time frame according to the bit allocation using a quantizer and signal the second comparison.

[0028] According to a sixth aspect, there is provided an apparatus including at least one processor and at least one memory storing instructions, which, when executed by the at least one processor, cause the system to at least: compare a bit allocation of a direction value index for a time frame of an audio object with a bit allocation threshold; in response to the comparison, decode the direction value index for the time frame of the audio object using an inverse quantizer according to the bit allocation or read a start of a signal flag associated with the direction value index for the time frame of the audio object; according to the state of the signal flag, set the quantized direction value for the time frame of the audio object to the quantized direction value for the previous time frame of the audio object, or decode the direction value index for the time frame of the audio object using an inverse quantizer according to the bit allocation to provide the quantized direction value for the time frame of the audio object and adjust the azimuth value of the quantized direction value for the time frame of the audio object according to the quantization resolution of the azimuth value.

[0029] A device comprising means for performing the actions of the above method.

[0030] A device configured to perform the actions described above.

[0031] A computer program that contains program instructions to cause a computer to perform the above method.

[0032] A computer program product stored on a medium can cause the device to execute the methods described herein.

[0033] Electronic devices may include the apparatus described herein.

[0034] The chipset may include the device described herein.

[0035] The embodiments of this application aim to address problems related to the current state of the technology.

[0036] To better understand this application, references to the attached drawings are made here as an example. [Brief explanation of the drawing]

[0037] [Figure 1] A schematic diagram of a system of apparatus suitable for implementing several embodiments is provided. [Figure 2] A schematic representation of an exemplary coding mode selector, as shown within the system of the apparatus in Figure 1, is shown according to several embodiments. [Figure 3] Figure 2 shows a flowchart illustrating the operation of an exemplary coding mode selector in several embodiments. [Figure 4] Figure 4 shows a flowchart illustrating the operation of an exemplary first, lowest, or only MASA bitrate coding mode in several embodiments. [Figure 5]Figure 4 shows a flowchart illustrating the operation of an exemplary second, low, or object-information coding mode in several embodiments. [Figure 6] Figure 4 shows a flowchart illustrating the operation of an exemplary third, high, or single-object coding mode in several embodiments. [Figure 7] Figure 4 shows flowcharts illustrating the operation of an exemplary fourth, best, or independent object and multiple input coding mode according to several embodiments. [Figure 8] A schematic representation of an exemplary audio object metadata encoder, shown in Figure 1, is provided for several embodiments. [Figure 9] Figure 8 shows a flowchart illustrating the operation of the encoding mode selector of an exemplary audio object metadata encoder, as shown in several embodiments. [Figure 10] Figure 8 shows flowcharts illustrating the operation of the encoder determination unit and low-rate encoder in several embodiments. [Figure 11] The following schematic diagrams illustrate exemplary audio object direction value decoders according to several embodiments. [Figure 12] Figure 11 shows a flowchart illustrating the operation of an exemplary low-rate decoder in several embodiments. [Figure 13] An exemplary device suitable for implementing the device shown in the aforementioned diagram is shown. [Modes for carrying out the invention]

[0038] The following describes in more detail suitable devices and possible mechanisms for encoding transport audio signals and parameter spatial audio signals, including spatial metadata. As mentioned above, immersive audio codecs (such as 3GPP IVAS) are planned to support a large number of operating points ranging from low bitrate operation to transparency. They are expected to support channel-based audio inputs and scene-based audio inputs, including spatial information about the sound field and sound source. In the following, exemplary codecs are configured to receive multiple input formats. Specifically, the codec is configured to acquire or receive multi-audio signals (e.g., received from a microphone array, or as a multi-channel audio format input, ambisonics format input) and one or more audio object signals (these can also be called independent streams with metadata (ISM) formats). Furthermore, in some situations, the codec is configured to process multiple input formats at once. This combined (input) format mode allows, for example, simultaneous encoding of two different audio input formats. An example of two different audio input formats currently under consideration is a combination of the MASA format and the audio object format. Metadata-assisted spatial audio (MASA) is an example of a parameter-space audio format and representation suitable as an input format for IVAS.

[0039] This can be considered an audio representation consisting of "N channels + spatial metadata." This is a scene-based audio format particularly well-suited for spatial audio capture on practical devices such as smartphones. The idea is to describe the acoustic scene in terms of the direction of sound sources, which continuously changes both temporally and frequency, and, for example, the energy ratio. Sound energy not defined (described) by direction is described as diffuse (coming from all directions).

[0040] As mentioned above, spatial metadata associated with an audio signal may include multiple parameters for each time-frequency tile (such as multiple directions, and the direct-to-total energy ratio, spread coherence, distance, etc. associated with each direction (or direction value)). Spatial metadata may also include other parameters or be associated with other parameters considered omnidirectional (such as surround coherence, spread-to-total energy ratio, remainder-total energy ratio), but spatial metadata can be used in combination with directional parameters to define the characteristics of an audio scene. For example, a reasonable design choice that can produce a good quality output is one in which the spatial metadata includes determining one or more directions (and the direct-to-total ratio, spread coherence, distance value, etc. associated with each direction) for each time-frequency subframe.

[0041] A concept further described herein is the definition of a multirate coding model that provides coding of combined formats at various bitrates. This coding model enables parametric coding of audio object inputs, including variable rate coding of object orientation parameters, based on the determined priority of the objects and the available bitrates. Thus, embodiments of apparatus and methods for defining the quantization resolution of object metadata as a function of both the MASA vs. Total Energy Ratio and the ISM Ratio, which are further described herein, are included. Based on these parameters, a priority value for each object is calculated and used to modify the quantization resolution of the object metadata. Furthermore, embodiments are included in which the stability of object orientation parameters over time can be utilized when the multirate coding model determines a lower coding rate for coding the object orientation parameters.

[0042] As described above, the parameter space metadata representation can use multiple concurrent spatial directions. In the case of MASA, the maximum number of proposed concurrent directions is two. For each concurrent direction, there may be associated parameters such as direction index; direct vs. total ratio; spread coherence; and distance. In some embodiments, other parameters such as diffusion vs. total energy ratio; surround coherence; and remainder vs. total energy ratio are defined.

[0043] In this regard, Figure 1 shows an exemplary apparatus 100 and system for implementing an embodiment of the present application. A system with an “analysis” component is shown. The “analysis” component is the part from receiving a multichannel signal to encoding metadata and a downmix signal.

[0044] The input to the "analysis" portion of the system is a multi-channel audio signal 102. In the following example, a microphone channel signal input is described, but in other embodiments, any suitable input (or composite multi-channel) format may be implemented. For example, in some embodiments, the spatial analyzer and spatial analysis may be implemented outside the encoder. For example, in some embodiments, spatial (MASA) metadata associated with the audio signal may be provided to the encoder as a separate bitstream. In some embodiments, the spatial (MASA) metadata may be provided as a set of spatial (directional) index values.

[0045] Furthermore, Figure 1 also shows multiple audio objects 104 as additional inputs to the analysis section. As described above, these multiple audio objects (or audio object streams) 104 may represent various sound sources in physical space. Each audio object may be characterized by an audio (object) signal and accompanying metadata including directional data (in the form of azimuth and elevation values), where the directional data indicates the position or orientation of the audio object in physical space for each audio frame.

[0046] The multi-channel signal 102 is passed to the analyzer and encoder 101, specifically to the transport signal generation unit 105 and the metadata generation unit 103.

[0047] In some embodiments, the metadata generation unit 103 is also configured to receive a multichannel signal, analyze the signal, and create metadata 104 related to the multichannel signal and, consequently, to the transport signal 106. The analysis processor 103 may be configured to generate metadata that may include directional parameters, energy ratio parameters, and coherence parameters (and, in some embodiments, diffusion parameters) for each time-frequency analysis interval. The directional, energy ratio, and coherence parameters may, in some embodiments, be considered MASA spatial audio parameters (or MASA metadata). In other words, spatial audio parameters include parameters intended to characterize the sound field created / captured by the multichannel signal (or generally two or more audio signals).

[0048] In some embodiments, the generated parameters may differ for each frequency band. For example, in band X, all parameters may be generated and transmitted, but in band Y, only one of the parameters may be generated and transmitted, and furthermore, in band Z, no parameters may be generated or transmitted. A real-world example of this is that in some frequency bands, such as the highest band, some parameters may not be needed for perceptual reasons. The transport signal 106 and metadata 104 may be passed to the coupled encoder core 109.

[0049] In some embodiments, the transport signal generation unit 105 is configured to receive a multichannel signal, generate an appropriate transport signal including a determined number of channels, and output a transport signal 106 (MASA transport audio signal). For example, the transport signal generation unit 105 may be configured to generate a downmix of two audio channels of the multichannel signal. The determined number of channels may be any appropriate number of channels. In some embodiments, the transport signal generation unit is configured to select the input audio signal by other means, such as beamforming technology, or to combine the input audio signal into a determined number of channels and output these as transport signals.

[0050] In some embodiments, the transport signal generation unit 105 is optional, and the multi-channel signal is passed to the coupled encoder core 109 in the same manner as the transport signal, without being processed in this example.

[0051] The audio object 104 may be passed to an audio object analyzer 107 for processing. In some embodiments, the audio object analyzer 107 analyzes the object audio input stream 104 to create appropriate audio object transport signals 128 and audio object metadata 108. For example, the audio object analyzer 107 may be configured to create an audio object transport signal 128 by combining the audio signals of the audio object 104 and downmixing them into a stereo channel using amplitude panning based on the associated audio object direction. Furthermore, the audio object analyzer may also be configured to create audio object metadata 108 associated with the audio object input stream 104. The audio object metadata may include direction values ​​applicable to all subbands. Thus, if there are four objects, there are four directions. In the examples described herein, the direction values ​​are also applied across all subframes of the frame, although in some embodiments, the time resolution of the direction values ​​may differ, and the direction values ​​are applied to one or more subframes of the frame. Furthermore, an energy ratio (or ISM ratio) may be determined for each object. The energy ratio (ISM ratio) defines the contribution of the object within the object portion of the overall audio environment. In the following example, the energy ratio (or ISM ratio) is for each time-frequency tile of each object.

[0052] In some embodiments, the audio object analyzer 107 may be located elsewhere, and the audio objects 104 input to the analyzer and encoder 101 are audio object transport signals and audio object metadata.

[0053] The analyzer and encoder 101 may include a coupled audio encoder core 109, which is configured to receive transport audio (e.g., downmix) signals 106 and audio object transport signals 128 in order to produce appropriate encoding of these audio signals.

[0054] The analyzer and encoder 101 may also include an audio object metadata encoder 111, which is configured to receive audio object metadata 108 and output the input information in an encoded or compressed format as encoded audio object metadata 112.

[0055] In some embodiments, the combined encoder core can be configured to implement a stream separation metadata determination unit and encoder, which can be configured to determine the relative contribution of the multi-channel signal 102 (MASA audio signal) and audio object 104 to the overall audio scene. This proportional measure created by the stream separation metadata determination unit and encoder may be used to determine the proportion of quantization and encoding “effort” spent on the input multi-channel signal 102 and audio object 104. In other words, the stream separation metadata determination unit and encoder may create a metric that quantizes the proportion of encoding effort spent on the multi-channel audio signal 102 compared to the encoding effort spent on audio object 104. This metric may be used to drive the encoding of audio object metadata 108. Furthermore, the metric determined by the separation metadata determination unit and encoder may also be used as an influencing factor in the process of encoding the transport audio signal 106 and audio object transport audio signal 128 performed by the combined encoder core 109. Furthermore, the output metrics from the stream isolation metadata determination unit and encoder can be represented as encoded stream isolation metadata and may be coupled to an encoded metadata stream from the coupled encoder core 109. In some embodiments, the analyzer and encoder 101 include a bitstream generation unit 113 configured to acquire encoded metadata 116, encoded transport audio signals 138, and encoded audio object metadata 112, and generate a bitstream 118 for potential transmission or storage.

[0056] In some embodiments, the analyzer and encoder 101 include an encoder controller 115. In some embodiments, the encoder controller 115 can control the encoding implemented by the audio object metadata encoder 111 and the coupled encoder core 109. In some embodiments, the encoder controller 115 is configured to determine the bitrate of the bitstream 118 and to control the encoding based on that bitrate. In some embodiments, the encoder controller 115 is further configured to control at least one of the audio object analyzer 107, the transport signal generation unit 105, and the metadata generation unit 103 when generating parameters.

[0057] The analyzer and encoder 101 may, in some embodiments, be a computer or mobile device (running appropriate software stored in memory and at least one processor), or alternatively, a specific device utilizing, for example, an FPGA or ASIC. Encoding may be performed using any suitable scheme. In some embodiments, the encoder 107 may further embed encoded MASA metadata, audio object metadata, and stream isolation metadata within the encoded (downmixed) transport audio signal before interleaving, multiplexing, or transmitting or storing in a single data stream, as shown by the dashed lines in Figure 1. Multiplexing may be performed using any suitable scheme.

[0058] Furthermore, with respect to Figure 1, the relevant decoder and renderer 129 is shown, which is configured to acquire a bitstream 118 containing encoded metadata 116, encoded transport audio signal 138, and encoded audio object metadata 112, and from there generate a suitable spatial audio output signal. The decoding and processing of such audio signals are known in principle and, apart from the decoding of encoded ISM ratio metadata, will not be described in detail further herein.

[0059] With respect to Figure 2, several embodiments of the encoder controller 115 are shown in more detail.

[0060] In this example, the encoder controller 115 includes a bitrate determination unit / monitor 201 configured to determine and / or monitor the bitrate available for the bandwidth of the encoded audio and metadata. This can be determined based on an estimate of the transmission path bandwidth (and, for example, based on estimated signal strength) or based on a bandwidth storage determination, or by any suitable method, in order to keep the file below the required size for a determined period of time.

[0061] Furthermore, the bitrate determination unit / monitor 201 can be configured to control the encoding mode selector 203. The encoder controller 115 may include an encoding mode selector 203 configured to select an encoding mode based, for example, the determined bandwidth or bitrate, and then control encoders, such as a combined encoder core 109 and an audio object metadata encoder 111.

[0062] Regarding Figure 3, a flowchart illustrating the exemplary operation of the encoder controller 115 shown in Figure 2 is provided. In this example, there is an initial operation to receive, acquire, or otherwise determine the encoded parameters and the bitrate or bandwidth of the audio data, as shown by step 301 in Figure 3.

[0063] Once the available bandwidth or bitrate is obtained, a check can then be performed to determine whether the bitrate is below a first (or minimum, or object minimum) threshold, as shown in step 303 of Figure 3.

[0064] If the available bandwidth or bitrate is below a first (or minimum, or minimum for an object) threshold, the encoder can be controlled to encode only the transport channel and MASA metadata, as shown by step 304 in Figure 3 (also shown as Mode A).

[0065] If the available bandwidth or bitrate exceeds the first (or lowest, or minimum for an object) threshold, further checks can be performed to determine whether the bitrate is below the second (or lowest, or single object) threshold, as shown by step 305 in Figure 3.

[0066] If the available bandwidth or bitrate is below a second (or lower, or one object) threshold, the encoder may be controlled to encode the transport channel, MASA metadata, ISM metadata (all objects), MASA to overall ratio, and ISM ratio, as shown by step 306 in Figure 3 (also shown as Mode B).

[0067] If the available bandwidth or bitrate exceeds the second (or lower, or single-object) threshold, further checks can be performed to determine whether the bitrate is below the third, higher, or full-object threshold, as shown by step 307 in Figure 3.

[0068] If the available bandwidth or bitrate is below a third, high, or full object threshold, the encoder may be controlled to encode the transport channel, MASA metadata, ISM metadata (all objects), MASA to total ratio, ISM ratio, and audio data for one object, along with one object identifier, as shown by step 308 in Figure 3 (also shown as Mode C).

[0069] If the available bandwidth or bitrate exceeds a third, high, or full object threshold, the encoder may be controlled to encode the transport channel, MASA metadata, ISM metadata (all objects), and audio data for all objects, as shown by step 310 in Figure 3 (also shown as Mode D).

[0070] The encoding modes can be summarized, for example, by the following table. [Table 1] It should be understood that the bitrates shown herein are examples and may be other specific values.

[0071] Figures 4-7 show flowcharts illustrating the first (or lowest, or combined) encoding mode shown by step 304 in Figure 3, the second (or low, or object metadata) encoding mode shown by step 306 in Figure 3, the third (or high, or one object) encoding mode shown by step 308 in Figure 3, and the fourth (or highest, or all objects) encoding mode shown by step 310 in Figure 3, respectively.

[0072] For example, the encoding method for Mode A in Figure 4 further illustrates the first (or lowest, or combined) encoding mode shown by step 304 in Figure 3. Therefore, for very low total bitrates (e.g., 32kbps or less), all encoding is performed using MASA representation.

[0073] Therefore, for example, there are operations to receive / acquire object-based streams (independent streams with metadata) and multi-channel-based (MASA streams) transport audio signals and metadata, as shown by step 401 in Figure 4.

[0074] Next, as shown by step 403 in Figure 4, there is an operation to generate an object-based MASA stream from an object stream (a standalone stream with metadata). In some embodiments, this object-based MASA stream can be created from the object stream using the method presented, for example, in WO2019086757A1.

[0075] Subsequently, the object-based MASA stream and the multi-channel-based MASA stream are combined, as shown by step 405 in Figure 4. In some embodiments, the original MASA stream and the MASA stream created from the objects can be combined using the method presented in GB2574238. The decoder obtains the object and MASA audio content in MASA format.

[0076] Next, the combined stream is output as shown by step 407 in Figure 4. In such an embodiment, the object audio content is present in the decoded audio scene (along with the MASA audio content), but the objects cannot be edited, nor can the decoder separate the objects from the scene.

[0077] Figure 5 illustrates a mode B coding method, a second (or lower, or object metadata) coding mode as shown by step 306 in Figure 3. Therefore, for low bitrates (e.g., between 48kbps and 80kbps), there is a greater number of bits available, and thus the possibility exists to parameterize the audio scene by transmitting, for each common audio data downmix, MASA metadata, ISM metadata, and time-frequency tile, an additional set of parameters indicating the amount of signal corresponding to MASA components outside the range of the entire audio scene (in other words, this can be presented or shown by the MASA to total energy ratio) and a ratio indicating how the audio scene corresponding to an object is distributed across the ISMs (in other words, this can be presented or shown by the ISM ratio).

[0078] Therefore, as shown by step 501 in Figure 5, for example, there are steps for receiving / acquiring object-based streams (independent streams with metadata) and multi-channel-based (MASA streams) transport audio signals and metadata.

[0079] Next, as shown in step 503 of Figure 5, a combined MASA and object-based downmix (channel-to-element) audio signal is generated. In other words, the audio content of the MASA and objects is downmixed to 2 channels (channel-to-element CPE).

[0080] The MASA to total ratio and the ISM ratio can be determined as shown in step 505 of Figure 5.

[0081] Next, the MASA-to-whole ratio and the ISM ratio can be encoded based on any suitable encoding method. For example, the ISM ratio can be encoded using lattice coding or other vector quantization methods, or the MASA-to-whole ratio can be encoded by DCT transformation followed by entropy coding (e.g., as described in WO2022 / 200666). The encoding of the MASA-to-whole ratio and the ISM ratio is shown by step 507 in Figure 5.

[0082] Furthermore, the MASA metadata can then be encoded based on any suitable MASA metadata encoding method shown by step 509 in Figure 5.

[0083] The combined audio signals can then be encoded based on any suitable audio signal coding method, as shown by step 511 in Figure 5.

[0084] Next, the encoder can output the encoded MASA metadata, MASA to whole ratio, ISM ratio, and combined transport audio signal, as shown by step 513 in Figure 5.

[0085] Figure 6 shows the Mode C coding method, a third (or higher, or single-object) coding mode as shown by step 308 in Figure 3. Thus, at medium to high bitrates (e.g., bitrates between 96kbps and 160kbps), the audio content of a single object is isolated and transmitted independently. Furthermore, the downmix formed from the MASA transport channel and the rest of the object are transmitted in MASA format, along with additional parameters of the MASA-total energy ratio and the ISM ratio. In addition, ISM metadata is transmitted, along with an identifier describing which object was isolated. In each frame, it is determined which object to isolate. This determination may be based, for example, on the relative level of the object to other objects (e.g., isolate the loudest volume object). This is described in detail in WO2022 / 214730.

[0086] Therefore, there are steps for receiving / acquiring object-based streams (independent streams with metadata) and multi-channel-based (MASA streams) transport audio signals and metadata, for example, as shown by step 601 in Figure 6.

[0087] Next, as shown by step 603 in Figure 6, one audio object is selected, and an object identifier is generated based on the selected audio object. Furthermore, the audio signal associated with the selected audio object is encoded. Any suitable audio signal encoder may be used to encode the audio signal of the selected object. For example, the same or similar audio signal encoder used to encode the MASA audio signal(s) can be employed. Then, the combined MASA and the remaining (or unselected) object-based transport audio signal (or downmix) is generated as shown by step 605 in Figure 6. The object transport signal can be created in the same way as presented in the previous mode, mode B, the difference being that the selected or separated objects are not included in the mix. For example, the multi-channel or MASA audio signal and the (unselected) object transport signal can be combined and summed to produce a combined transport audio signal.

[0088] The MASA to total ratio and the ISM ratio can be determined as shown in step 607 of Figure 6.

[0089] Next, as shown by step 609 in Figure 6, the object identifier, MASA metadata, object metadata for all objects, MASA to total ratio, and ISM ratio can be encoded based on any suitable encoding method. For example, the encoding can employ any scalar or vector quantizer, followed by or without entropy encoding. The encoding of the MASA to total energy ratio can be carried out in the manner described in WO2022 / 200666. The encoding of the ISM ratio will be described in further detail later.

[0090] The combined audio signals can then be encoded based on any suitable MASA audio signal encoding method, as shown by step 611 in Figure 6. Encoding of the combined transport audio signals can employ any suitable transport audio signal encoding method, such as a MASA encoder.

[0091] In other words, the separated objects are determined, separated, and encoded as described in WO2022 / 214730, and the remaining objects and MASA streams are processed as described in WO2022 / 200666.

[0092] The encoder can then output the encoded object identifier, MASA metadata, MASA to total ratio, ISM ratio, object metadata (for all objects), selected single object audio signal, and combined transport audio signal, as shown by step 613 in Figure 6.

[0093] Figure 7 shows the Mode D coding method, the fourth (or best, or all object) coding mode, as shown by step 310 in Figure 3. Thus, at higher bitrates (e.g., 160kbps or higher), the two input audio formats, MASA and ISM, are coded and transmitted independently.

[0094] Therefore, as shown by step 701 in Figure 7, there are steps for receiving / acquiring, for example, object-based streams (independent streams with metadata) and multi-channel-based (MASA streams) transport audio signals and metadata.

[0095] Next, as shown by step 703 in Figure 7, the multi-channel-based (MASA stream) transport audio signal and metadata are encoded based on any suitable MASA encoding method.

[0096] The object (independent stream with metadata) and associated metadata can be further encoded, as shown in step 705 of Figure 7. For example, encoding can be performed using any suitable mono encoder (as part of the main encoder), such as an EVS-based mono encoder.

[0097] Next, as shown by step 707 in Figure 7, the encoder can output independently encoded objects (independent streams with metadata) and associated metadata, as well as independently encoded multi-channel-based (MASA stream) transport audio signals and metadata.

[0098] The generation and encoding of ISM ratio values, which are determined and encoded within encoding modes B and C, will be described in more detail below.

[0099] The following describes the encoding of directional metadata parameters related to objects (determined from ISM metadata), and in some embodiments, the encoding of directional metadata parameters related to objects when the encoder is operating in Mode B and Mode C (second or third) encoding modes. However, in some embodiments, it will be understood that the following can also be applied to any encoding mode in which audio object metadata (and specifically directional metadata) is encoded. In embodiments where the MASA to overall ratio and ISM ratio are not transmitted, further information regarding the encoding details (bit allocation) is then transmitted. Furthermore, in some embodiments, the ratio can, in principle, be calculated from the encoded signal in both the encoder and the decoder.

[0100] Accordingly, with respect to Figure 8, the audio object metadata encoder 111 in several embodiments is shown in more detail. In the following examples, the audio object metadata encoder 111 is configured to receive ISM metadata, specifically the MASA to total energy ratio m2t812, the ISM ratio r814, and the direction value 802, as input. In other words, the ISM metadata includes the direction information (elevation and azimuth) of each object in each frame. There is a standard resolution of 11 bits for each elevation-azimuth pair. Furthermore, the MASA to total ratio and the ISM ratio can also be determined or obtained from the ISM and from the multi-channel audio signal (as described, for example, in WO2022 / 200666).

[0101] In some embodiments, the ISM ratio can be obtained, for example, as follows:

[0102] First, the object audio signal S obj (t,i) is in the time-frequency domain S obj The signal is converted to (b,n,i) (where t is the time sample index, b is the frequency bin index, n is the time frame index, and i is the object index). The time-frequency domain signal can be obtained, for example, via the Short-Time Fourier Transform (STFT) or a Complex Modulated Quadrature Filter Bank (QMF) (or their low-latency variants).

[0103] Next, the object's energy is calculated in the following frequency band.

number

number

[0104] In some embodiments, the time resolution of the ISM ratio is defined as the time-frequency domain audio signal S obj The time resolution of (b,n,i) may differ (i.e., the time resolution of spatial metadata may differ from the time resolution of the time-frequency conversion). In such cases, the calculation (of the energy ratio and / or ISM ratio) may include the sum of the time-frequency domain audio signal and / or energy values ​​over multiple time frames.

[0105] The ISM ratio is a number between 0 and 1, corresponding to the proportion of time an object is active within the audio scene created by all objects. For each object, there is one ISM ratio for each frequency subband and time subframe. As described above, the ISM ratio is passed to the audio object metadata encoder 111.

[0106] In some embodiments, the audio object metadata encoder 111 is configured to encode the MASA to whole ratio and the ISM ratio, and specific encodings of the ISM ratio and the MASA to whole ratio are not described in further detail herein. For example, WO2022 / 200666 describes a suitable MASA to whole ratio encoding method, and GB applications 2217884.2 and 2217905.5 describe a suitable ISM ratio encoding method.

[0107] In some embodiments, an (audio) object priority determination unit 803 is present. The object priority determination unit 803 is configured to generate priority values ​​for objects within a time-frequency tile. In some embodiments, the object priority determination unit 803 is configured to obtain MASA-total ratio 812 and ISM ratio 814. For each time-frequency (TF) tile (i.e., each subband-subframe combination), there is one MASA-total energy ratio and N ISM ratios, where N is the number of objects.

[0108] Furthermore, in some embodiments, the priority value can be defined as the maximum value of the object contribution rate across all time-frequency tiles.

[0109] For example, priority values ​​can be generated based on the following:

number

[0110] In the above equation, there are N objects, B subbands, and M time subframes. M2t(b,k) represents the quantized MASA to total energy ratio for subband b and subframe k. r(i,b,k) is the quantized ISM ratio for object i in subband b and subframe k. Below, quantized versions of the MASA to total ratio and ISM ratio are used so that the decoder can also use the same values.

[0111] In some embodiments, several other parameters (other than the MASA to total energy ratio) can be adopted. For example, the object to total energy ratio may be something like o2t(b,k) = 1 - m2t(b,k), and thus the same information can be effectively conveyed.

[0112] In such embodiments where alternative parameters are adopted, the above equation is updated accordingly. Thus, in the above example of the object-to-total energy ratio, the object-to-total o2t(b,k) may simply replace the (1-m2t(b,k)) part.

[0113] Alternatively, in some embodiments, the (weighted) average of object contributions within a TF tile can also be considered to select priority. Furthermore, in some other embodiments, the maximum operator can be replaced with a second or third maximum value (or any similar value). This ensures that the highest contribution in a single TF tile does not determine priority on its own.

[0114] Next, the object priority value 804 may be passed to the bit determination unit 805.

[0115] In some embodiments, the audio object metadata encoder 111 includes a bit determination unit 805. The bit determination unit 805 is configured to acquire or receive object priorities 804 and, based on those object priority values, determine the number of bits available to encode the direction parameters. In other words, it sets the number of bits that define the quantization grid used to encode the direction values ​​802.

[0116] In some embodiments, the bit determination unit 805 is configured to determine or assign fewer bits to lower-priority objects.

[0117] As an example, the number of bits allocated to each object direction metadata based on object priority can be calculated as follows: bits(i)=11-[(1-p(i))*6],i=0:N-1

[0118] The [ ] operator represents rounding to the nearest integer operation. The maximum number of bits per object direction is 11, and the minimum number of bits is 4. The maximum and minimum number of bits may be other values ​​depending on the details of the embodiment, and therefore the value 7 (the difference between the maximum and minimum number of bits per object direction) may vary in some other embodiments. Although this example shows a linear scale, the above formula providing the number of bits can be replaced with any other increasing linear or nonlinear function of object priority, ensuring that the number of bits is within or similar to the region [4,11].

[0119] Also shown in Figure 8 is a low-bitrate directional encoder 820 that can be applied to encoding the directional values ​​of audio objects when the encoder controller 115 determines a lower coding rate, such as the coding modes of Mode B and Mode C. This low-bitrate directional encoder 820 may be applied to encoding the directional values ​​of audio objects instead of the spherical grid determination unit and encoder 807. For that purpose, Figure 8 also shows an encoder determination unit 819 configured to select between the low-bitrate encoder 820 and the spherical grid determination unit and encoder 807. The selected encoder is then used to encode the directional values ​​of audio objects.

[0120] The choice between two different encoding schemes may depend on the bit allocation 806 for quantizing the direction value. For example, the encoder determination unit 819 may be configured to receive the bit allocation 806 and test it against a predetermined threshold bit value. The result of the test can then be used to direct the audio object direction value 802 to either the low-rate encoder 820 or the spherical grid determination unit and encoder 807. Figure 8 shows the configuration of the encoder determination unit 819, which receives both the bit allocation 806 and the audio object direction value 802 and outputs the audio object direction value 802 to either the low-rate encoder 820 or the spherical grid determination unit and encoder 807.

[0121] In the embodiment, the encoder determination unit 819 may have the following functions. First, it may examine the bit allocation 806 to determine whether the number of bits allocated to quantize the direction value of a particular audio object is less than a predetermined threshold bit value. If the allocated number is less than the predetermined threshold, the audio object direction value 802 is directed along the signal feed 832 to the low-rate encoder 820 for encoding. Otherwise, the audio object direction value 802 is directed along the signal feed 830 to the spherical grid determination unit and encoder 807 for encoding. The selection between the low-bitrate encoder 820 and the spherical grid determination unit and encoder 807 may be made on a per-audio-object basis.

[0122] The above functions of the encoder determination unit 819 can be illustrated by the following pseudocode. 1. For each object i a. If bit (i) < 8 i. Send the unquantized direction values ​​to the low-rate encoder 820. ii. End b.Else i. Transmit the unquantized direction values ​​to the spherical grid determination unit and encoder 807. ii. End c.End 2. End for

[0123] The pseudocode above shows that the specified threshold bit value is set to 8 bits. However, please understand that other threshold values ​​may be used, and that these values ​​can be obtained through experimentation.

[0124] Returning to Figure 8, when the low-rate encoder 820 is selected, the encoder is configured to receive audio object direction values ​​802 along the signal feed 832. The low-rate encoder 820 may then be configured to compare the audio object direction values ​​(azimuth and elevation) of the current frame with the audio object direction values ​​of the same audio object from the previous frame. In embodiments, this comparison may be performed using unquantized audio object direction values. If the comparison determines that the audio object direction value from the previous frame is the same as the audio object direction value from the current frame, the low-rate encoder 820 is configured to signal this as a single bit to indicate that there has been no change in the direction value from the previous frame. This is shown as the LR_enc signal bit 824 in Figure 8. In other words, in this instance, the encoded audio object direction value for an audio object is a single bit.

[0125] Note that the comparison is based on the directional value component. That is, the azimuth value of the current frame is compared to the azimuth value of the previous frame, and the elevation value of the current frame is compared to the elevation value of the previous frame. If the comparison does not register a change (or difference) in either component, the result of the comparison is that the audio object directional value of the previous frame is the same as the audio object directional value of the current frame.

[0126] Further embodiments may be configured to have an LR_enc signal containing multiple bits, which allows for the determination of a difference between the directional value components from the previous frame to the current frame. For example, a single bit may be employed to indicate whether either or both of the azimuth and elevation values ​​of the previous frame are the same as either or both of the azimuth and elevation values ​​of the current frame. Additional bits may be used to indicate whether either the azimuth or elevation value is the same from the previous frame to the current frame, or whether both the azimuth and elevation values ​​are the same. If either the azimuth or elevation value is the same from the previous frame to the current frame, other bits may be used to identify whether it is the azimuth value or the elevation value. This further embodiment may then be configured to quantize, encode, and index the angle (azimuth or elevation) that has changed from the previous frame to the current frame.

[0127] If the above comparison of audio objects indicates that the current frame direction value is not the same as the previous frame direction value, the low-rate encoder 820 directs the audio object direction value 802 along the signal feed 834 to the spherical grid determination unit and encoder 807 for encoding. This condition can be signaled using the opposite state of the LR_enc signal bit 824 described above.

[0128] Regarding pseudocode, the functionality of the low-rate encoder 820 may take the following forms: i. When the unquantized elevation and azimuth angles are the same as in the previous frame. Send 1.1 bits (1) to signal that the direction is the same. ii. Else 1. Send a 1.1 bit (0) to signal that the directions are not the same. 2. Direct the audio object direction value to encoder 807 for encoding. iii. End

[0129] Figure 10 shows the processing steps of the encoder determination unit 819 and the low-rate encoder 820.

[0130] First, the bit allocation 806 of the audio object is received by the encoder determination unit 819 and then checked against the bit allocation threshold. This is shown in Figure 10 as processing step 1001. If processing step 1001 determines that the bit allocation of the audio object is below the threshold, the encoder determination unit 819 determines that the audio object direction value should be encoded using the low-rate encoder 820. This is shown in Figure 10 as proceeding to processing step 1005. However, if processing step 1001 determines that the bit allocation is above the threshold, the encoder determination unit 819 determines that the spherical grid quantizer 807 should be used directly to encode the audio object direction value 802. This is shown in Figure 10 as proceeding to processing step 1003, when the audio object direction value 802 is sent directly to the spherical grid quantizer 807 for encoding.

[0131] In processing step 1005, the low-rate encoder 820 compares the audio object direction value of the current frame with the audio object direction value of the previous frame to determine whether the audio object direction value 802 should be sent to the spherical grid quantizer 807 for encoding. If the comparison indicates that the audio object direction value of the current frame is the same as the audio object direction value of the previous frame, the audio object direction parameter is not sent to the spherical grid determination unit 807 for encoding. Instead, the audio object direction parameter 802 of the current frame is not encoded in essence, and rather, the state of the LR_enc signal bit 824 is set to indicate that the audio object direction value 802 of the current frame is the same as the audio object direction value of the previous frame. This is shown in Figure 10 as the process proceeds to step 1009.

[0132] If the comparison in step 1005 indicates that the two direction values ​​are not the same, the low-rate encoder 820 decides that the audio object direction value for the current frame is sent to the spherical grid quantizer 807 for encoding. This is shown in Figure 10 as a transition from processing step 1005 through step 1007 to processing step 1003. In step 1007, the low-rate encoder 820 is configured to set the state of the LR_enc signal 824 to indicate that the audio object direction values ​​for the current frame and the previous frame are different.

[0133] The low-rate encoder 820 can be configured to output the LR_Enc signal bit 824 to the bitstream generation unit 113 so that it is included in the bitstream 118.

[0134] As described above, the audio object metadata encoder 111 includes a spherical grid determination unit and an encoder 807 configured to receive a bit assignment 806 and a direction value 802. The spherical grid determination unit and encoder 807 are then configured to generate an output index value based on the point closest to the direction value 802 in the determined spherical grid defined by the audio object's bit assignment 806. The encoded direction index value 808 of the audio object may then be passed to the bitstream generation unit 113 for inclusion in the bitstream 113.

[0135] Other embodiments may be configured to quantize the orientation of each object using a corresponding number of bits and output an elevation index and an azimuth index. Object metadata can be encoded independently for each object, or the orientation metadata of objects can be encoded together using, for example, some weights to accommodate different bit assignments.

[0136] The spherical grid employs the idea of ​​covering a sphere with smaller spheres and considering the centers of the smaller spheres as points that define a grid in approximately equidistant directions. Each point on the grid is defined by a pair of azimuth and elevation values.

[0137] The highest resolution (11 bits) spherical grid used for directional quantization guarantees a quantization resolution of 5 degrees. The structure of the spherical grid is the same as that used for quantizing MASA directional metadata (and the method for defining the grid and encoding the index for MASA directional metadata values ​​is known, as described in PCT / EP2017 / 078948, GB1811071.8).

[0138] In some embodiments, when an object's priority is exactly zero, the number of bits for that object is set to zero, and no direction metadata is transmitted for that object.

[0139] Regarding Figure 9, a flowchart summarizing the operation of the exemplary audio object metadata encoder 111 shown in Figure 8 is provided.

[0140] The initial operation is to receive / acquire the metadata-enhanced independent stream, as well as the determined (quantized) MASA versus total energy ratio (m2t) and ISM ratio (r), as shown by step 901 in Figure 9.

[0141] The next step is to determine the object priority based on the (quantized) ISM ratio value (r) and the (quantized) MASA to total energy ratio (m2t), as shown by step 903 in Figure 9.

[0142] Once the object priority is determined, the bit assignment of the direction parameter is then determined based on the determined object priority, as shown in step 905 of Figure 9.

[0143] Next, from the bit allocation, a spherical grid for encoding the object's orientation parameters is determined, as shown by step 907 in Figure 9.

[0144] Next, an encoded directional index is generated within the determined spherical grid, as shown by step 909 in Figure 9.

[0145] Next, the encoded direction index value can be output for inclusion in the bitstream, as shown by step 911 in Figure 9.

[0146] With respect to the decoder, decoding of encoded directional information can be carried out by determining similar prioritizations and the associated quantization grid. Thus, for example, a pseudocode representation of encoding and decoding may be as follows: Encoding of object metadata 1. For each object a. Calculate priority p(i) b. Calculate the number of bits (i) 2. End for 3. For each object a. Encode the spherical grid orientation metadata corresponding to the number of bits assigned to the object. 4. End for Decryption of object metadata 1. For each object a. Calculate priority p(i) b. Calculate the number of bits (i) 2. End for 3. For each object a. Decode the spherical grid orientation metadata corresponding to the number of bits assigned to the object. 4. End for

[0147] Figure 11 shows the audio directional decoder 1101 and the bitstream receiver and demultiplexer 1113. The bitstream receiver and demultiplexer 1113 is configured to receive the encoded bitstream 118 and demultiplex it into various signal streams of the encoded parameters, of which Figure 11 shows the encoded stream related to the audio directional decoder 1101.

[0148] The audio direction value decoder 1101 is shown to include a decoder determination unit and a spherical grid decoder 1119, and a low-rate decoder 1120. The decoder determination unit and the spherical grid decoder 1119 are configured to receive encoded direction index values ​​808 and bit assignments 806 corresponding to audio objects.

[0149] The bit allocation 806 for audio object i can be determined locally by the decoder by determining the priority value p(i) of the audio object. The priority value can be determined from the ISM ratio of the audio object. The ISM ratio is sent to the decoder as part of bitstream 118.

[0150] The decoder determination unit and the spherical grid decoder 1119 are configured to first compare the bit assignment 806 of an audio object with a predetermined threshold bit value (for each audio object) in order to determine whether the decoded direction value 1108 is obtained by directly decoding the encoded direction index or by utilizing the functionality of the low-rate decoder 1120. In other words, when the bit assignment 806 is less than the threshold, the low-rate decoder 1120 is selected to form the decoded direction value 1108. This is shown in Figure 11 as the LR decoder selection 1118 for the signal feed. Conversely, when the bit assignment 806 is greater than or equal to the threshold, the decoded direction value 1108 is generated directly by the decoder determination unit and the spherical grid decoder 1119.

[0151] The above functions of the decoder determination unit and the spherical grid decoder 1119 can be illustrated by the following pseudocode. For each object i, If bit(i) < 8 i. Select a low-rate decoder. ii. End Else i. Decode the encoded direction index value ii. End End if

[0152] When the low-rate decoder 1120 is selected to decode the audio direction value of an audio object, the low-rate decoder 1120 is configured to first check the value of the LR_enc signal bit 824. If the state of the LR_enc_signal bit 824 indicates that there has been no change in the direction value [of the audio object] between the previous audio frame and the current audio frame (LR_enc_signal=1), the low-rate decoder 1120 simply outputs the direction value of the previous frame [of the audio object] as the audio object direction value for the current frame.

[0153] It should be noted that the decoder determination unit and the spherical grid decoder located within the spherical grid decoder 1119 are configured to decode the encoded direction index values ​​of the audio object according to the bit assignment 806 of the audio object.

[0154] Conversely, if the LR_enc_signal bit 824 indicates that the direction values ​​of the audio objects between the previous frame and the current frame are different, the low-rate decoder 1120 is configured to operate in a different way, and in this regard, Figure 12 shows the processing performed by the low-rate decoder 1120.

[0155] Firstly, the low-rate decoder 1120 receives the decoded direction value of the current frame from the decoder determination unit and the spherical grid decoder 1119.

number

[0156] Next, the low-rate decoder 1120 calculates the quantization azimuth angle (direction) value of the current frame.

number

number

[0157] The check in processing step 1203 is cumulative

number

number

number

number

[0158] When resolution is required, the process performs a further check in step 1207 to determine the quantized azimuth (direction) value of the current frame.

number

number

number

number

[0159] However, if the conditions for the check in processing step 1207 are not met, the process proceeds to a further check in processing step 1211. In this step, the quantized azimuth (direction) value of the previous frame is checked.

number

number

number

number

[0160] To summarize. The overall effect of the processing steps in FIG. 12 is that when the state of the LR_enc_signal bit indicates that there is a change in the [direction value of the audio object] between the previous audio frame and the current audio frame (LR_enc_signal = 0, FIG. 12), the low-rate decoder 1120 adds the value Δ φ / 2 or subtracts the value Δ φ / 2 to change the decoded elevation value

Number

Number

[0161] The function of the low-rate decoder 1120 can also be illustrated by the following pseudo-code. i. Read 1 bit, the LR_enc_signal bit, from the bit stream ii. If LR_enc_signal == 1 1. The direction is the same as the previous frame iii. Else 1. Receive the direction index 2. Decode the direction index into

Number

number

number

number

number

number

[0162] With respect to Figure 10, an exemplary electronic device that may be used as one of the device components of the system as described above is shown. The device may be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc. The device may be configured to implement, for example, an encoder / analyzer component and / or decoder component as shown in Figure 1 or any of the functional blocks above.

[0163] In some embodiments, the device 1400 includes at least one processor or central processing unit 1407. The processor 1407 can be configured to execute various program codes, such as those described herein.

[0164] In some embodiments, the device 1400 includes at least one memory 1411. In some embodiments, at least one processor 1407 is coupled to the memory 1411. The memory 1411 may be any suitable storage means. In some embodiments, the memory 1411 includes a program code section for storing program code that can be implemented on the processor 1407. Furthermore, in some embodiments, the memory 1411 may further include a stored data section for storing data, for example, data processed or processed according to the embodiments described herein. The implemented program code stored in the program code section and the data stored in the stored data section can be retrieved by the processor 1407 as needed via the memory-processor coupling.

[0165] In some embodiments, device 1400 includes a user interface 1405. In some embodiments, the user interface 1405 can be coupled to a processor 1407. In some embodiments, the processor 1407 can control the operation of the user interface 1405 and receive input from the user interface 1405. In some embodiments, the user interface 1405 may allow a user to input commands to device 1400, for example, via a keypad. In some embodiments, the user interface 1405 may allow a user to retrieve information from device 1400. For example, the user interface 1405 may include a display configured to display information from device 1400 to the user. In some embodiments, the user interface 1405 may include a touchscreen or touch interface capable of both allowing information to be input to device 1400 and further displaying the information to the user of device 1400. In some embodiments, the user interface 1405 may be a user interface for communication.

[0166] In some embodiments, device 1400 includes an input / output port 1409. In some embodiments, the input / output port 1409 includes a transceiver. The transceiver in such embodiments can be coupled to a processor 1407 and configured to enable communication with other devices or electronic devices, for example, over a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means can, in some embodiments, be configured to communicate with other electronic devices or devices via wires or wire couplings.

[0167] The transceiver can communicate with further devices by any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable radio access architecture based on Long-Term Evolution Advanced (LTE Advanced, LTE-A) or New Radio (NR) (also known as 5G), Universal Mobile Telecommunications System (UMTS) Radio Access Network (UTRAN or E-UTRAN), Long-Term Evolution (LTE, same as E-UTRA), 2G Network (Legacy Network Technology), Wireless Local Area Network (WLAN or Wi-Fi), WiMAX (Worldwide Interoperability for Microwave Access), Bluetooth®, Personal Communication Services (PCS), ZigBee®, Wideband Code Division Multiple Access (WCDMA), Systems using Ultra-Wideband (UWB) Technology, Sensor Networks, Mobile Ad Hoc Networks (MANET), Cellular Internet of Things (IoT) RAN and Internet Protocol Multimedia Subsystem (IMS), any other suitable option, and / or any combination thereof.

[0168] The transceiver's input / output port 1409 may be configured to receive signals.

[0169] In some embodiments, device 1400 may be employed as at least part of a synthesis device. The input / output port 1409 may be coupled to headphones (which may be head-tracking or non-tracking headphones) or similar, as well as speakers.

[0170] Typically, various embodiments of the present invention can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some embodiments may be implemented in hardware, while others may be implemented in firmware or software that can run on a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Various embodiments of the present invention can be illustrated and described using block diagrams, flowcharts, or any other graphical representation, but it will be understood that these blocks, devices, systems, techniques, or methods described herein can be implemented, in non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or any combination thereof.

[0171] Embodiments of the present invention can be implemented by computer software executable by a data processor in a mobile device, such as within a processor entity, or by hardware, or by a combination of software and hardware. Furthermore, it should be noted that in this regard, any block in the logic flow in the figure may represent a program step or an interconnected logic circuit, block, and function, or a combination of a program step and a logic circuit, block, and function. The software can be stored on a physical medium such as a memory chip or memory block implemented within the processor, a magnetic medium such as a hard disk or floppy disk, and an optical medium such as a DVD and its data variant, a CD.

[0172] Memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technical environment and may include, in non-limiting examples, one or more of the following: general-purpose computers, dedicated computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits, and processors based on multicore processor architectures.

[0173] Embodiments of the present invention can be implemented in various components, such as integrated circuit modules. Designing integrated circuits is generally a highly automated process. Complex and powerful software tools can be used to translate logic-level designs into ready-to-etch semiconductor circuit designs for formation on semiconductor substrates.

[0174] Programs such as those provided by Synopsys, Inc. in Mountain View, California, and Cadence Design in San Jose, California, automatically route conductors and identify components on a semiconductor chip using well-established design rules and a library of pre-stored design modules. Once the semiconductor circuit design is complete, the resulting design can be transmitted to a semiconductor manufacturing facility or "fab" for production in a standardized electronic format (e.g., Opus, GDSII, etc.).

[0175] As used in this application, the term “circuit” may refer to one, more, or all of the following: (a) Dedicated circuit implementation for hardware (such as implementation using only analog and / or digital circuits), (b) For example (where applicable), the following combinations of hardware circuits and software: (i) A combination of analog and / or digital hardware circuits(s) and software / firmware, (ii) Any part of a hardware processor(s), software, and memory(s) that use software (including digital signal processors(s)) to cooperate in causing a device such as a mobile phone or server to perform various functions, Hardware circuits(s) that require software (e.g., firmware) to operate, and / or processors(s), such as microprocessors(s) or parts of microprocessors(s), however, the software may be omitted if it is not necessary for operation.

[0176] This definition of circuit applies to all use of the term in this application, including in all claims. Further as used in this application, the term circuit also includes, simply, an implementation of a hardware circuit or processor (or more processors), or a part of a hardware circuit or processor and the software and / or firmware associated therewith. The term circuit also includes, for example, a baseband integrated circuit or processor integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, other computing device, or other network device, where it falls under the element of a particular claim.

[0177] As used herein, the term “non-transient” refers to limitations of the medium itself (i.e., tangible rather than signalistic) rather than limitations on the persistence of data storage (e.g., RAM vs. ROM).

[0178] Where used herein, “at least one of the following: <list of two or more elements>” and “at least one of the <list of two or more elements>” and similar expressions in which lists of two or more elements are joined by “and” or “or” mean at least one of the elements, at least two or more of the elements, or at least all of the elements.

[0179] The foregoing description has provided a complete and useful description of exemplary embodiments of the present invention, as illustrative and non-limiting examples. However, in light of the foregoing description, and in conjunction with the accompanying drawings and claims, various modifications and adaptations may become apparent to those skilled in the art. Nevertheless, all such modifications and similar modifications of the teachings of the present invention remain within the scope of the present invention as defined in the accompanying claims.

Claims

1. It is a device: Receiving the direction value of the audio object relative to the time frame; The first comparison is to compare the bit allocation with the bit allocation threshold; In accordance with the first comparison, the direction value of the audio object with respect to the time frame is quantized using a quantizer according to the bit assignment, or, as a second comparison, the direction value of the audio object with respect to the time frame is compared with the direction value of the audio object with respect to the previous time frame; In response to the second comparison, the direction value for the time frame is quantized using the quantizer according to the bit assignment, and the second comparison is signaled. The apparatus comprising means configured to perform the following.

2. A means configured to signal the second comparison by quantizing the direction value for the time frame using the quantizer according to the bit assignment in response to the second comparison is: When the second comparison indicates that the direction value of the audio object for the time frame is different from the direction value of the audio object for the previous time frame, the direction value for the time frame is quantized using the quantizer according to the bit assignment, and a signal flag is set to indicate that the direction value for the time frame is different from the direction value for the previous time frame; When the second comparison indicates that the direction value of the audio object with respect to the time frame is the same as the direction value of the audio object with respect to the previous time frame, set the signal flag to indicate that the direction value of the time frame is the same as the direction value of the previous time frame. The apparatus according to claim 1, configured to perform the following:

3. A means configured to perform, in response to the first comparison, quantize the direction value of the audio object with respect to the time frame using a quantizer according to the bit assignment, or, as a second comparison, compare the direction value of the audio object with respect to the time frame with respect to the direction value of the audio object with respect to the previous time frame: When the bit allocation for the quantization of the direction value is greater than or equal to the bit allocation threshold, the direction value of the audio object with respect to the time frame is quantized using a quantizer according to the bit allocation; When the bit allocation for the quantization of the direction value is less than the bit allocation threshold, a second comparison is made by comparing the direction value of the audio object for the time frame with the direction value of the audio object for the previous time frame. The apparatus according to claims 1 and 2, configured to perform the following:

4. The apparatus according to claims 1 to 3, wherein the quantizer is a spherical grid quantizer, the spherical grid is formed by covering the sphere with smaller spheres, and the smaller spheres define points on the spherical grid.

5. The apparatus according to claims 1 to 4, wherein the directional value includes an azimuth angle value and an elevation angle value.

6. It is a device: Comparing the bit assignment of the direction value index for the time frame of an audio object with the bit assignment threshold; In accordance with the comparison, the directional index of the audio object to the time frame is decoded using an inverse quantizer according to the bit assignment, or the state of the signal flag associated with the directional index of the audio object to the time frame is read; Depending on the state of the signal flag; Set the quantized direction value of the audio object for the time frame to the quantized direction value of the audio object for the previous time frame, or The directional index of the audio object for the time frame is decoded using an inverse quantizer according to the bit allocation to provide a quantized directional value of the audio object for the time frame, and the azimuth angle value of the quantized directional value of the audio object for the time frame is adjusted according to the quantization resolution of the azimuth angle value. The apparatus comprising means configured to perform the following.

7. The means configured to adjust the azimuth angle value of the quantized direction value of the audio object with respect to the time frame according to the quantization resolution of the azimuth angle value is: The process involves comparing the product of the azimuth angle value of the quantized direction value of the audio object for the time frame with the azimuth angle value of the quantized direction value for the previous time frame, and if the product is greater than zero: Determining the difference between two consecutive azimuth quantization values ​​of the inverse quantizer; When the difference between the azimuth angle value of the quantized direction value of the audio object for the time frame and the azimuth angle value of the quantized direction value for the previous time frame is greater than half the difference between two consecutive azimuth angle quantized values ​​of the inverse quantizer, subtract half the difference between two consecutive azimuth angle quantized values ​​of the inverse quantizer from the azimuth angle value of the quantized direction value of the audio object for the time frame; When the difference between the azimuth angle value of the quantized direction value of the audio object for the previous time frame and the azimuth angle value of the quantized direction value for the time frame is greater than half the difference between two consecutive azimuth angle quantized values ​​of the inverse quantizer, half the difference between two consecutive azimuth angle quantized values ​​of the inverse quantizer is added to the azimuth angle value of the quantized direction value of the audio object for the time frame. The apparatus according to claim 6, configured to perform the following:

8. The means configured to determine the difference between two consecutive azimuth quantization values ​​of the inverse quantizer is: The circumference is configured to be divided by a number of azimuth angle values, the number of azimuth angle values ​​being determined by the elevation value of the quantized direction value with respect to the time frame. The apparatus according to claim 7.

9. The state of the aforementioned signal flag is: The direction value of the audio object with respect to the time frame is different from the direction value of the audio object with respect to the previous time frame; or The direction value of the audio object with respect to the time frame is the same as the direction value of the audio object with respect to the previous time frame; The apparatus according to claims 6 to 8, which shows one of the following.

10. The apparatus according to claims 6 to 9, wherein the inverse quantizer is a spherical grid quantizer, the spherical grid is formed by covering the sphere with smaller spheres, and the smaller spheres define points on the spherical grid.

11. Method: Receiving the direction value of the audio object relative to the time frame; The first comparison is to compare the bit allocation with the bit allocation threshold; In accordance with the first comparison, the direction value of the audio object with respect to the time frame is quantized using a quantizer according to the bit assignment, or, as a second comparison, the direction value of the audio object with respect to the time frame is compared with the direction value of the audio object with respect to the previous time frame; In response to the second comparison, the direction value for the time frame is quantized using the quantizer according to the bit assignment, and the second comparison is signaled. The method, including the method described above.

12. In accordance with the second comparison, the directional value for the time frame is quantized using the quantizer according to the bit assignment, and the comparison is signaled as follows: When the second comparison indicates that the direction value of the audio object for the time frame is different from the direction value of the audio object for the previous time frame, the direction value for the time frame is quantized using the quantizer according to the bit assignment, and a signal flag is set to indicate that the direction value for the time frame is different from the direction value for the previous time frame; When the second comparison indicates that the direction value of the audio object with respect to the time frame is the same as the direction value of the audio object with respect to the previous time frame, set the signal flag to indicate that the direction value of the time frame is the same as the direction value of the previous time frame. The method according to claim 11, including the method described in claim 11.

13. In accordance with the first comparison, the direction value of the audio object for the time frame is quantized using a quantizer according to the bit assignment, or, as a second comparison, the direction value of the audio object for the time frame is compared with the direction value of the audio object for the previous time frame: When the bit allocation for the quantization of the direction value is greater than or equal to the bit allocation threshold, the direction value of the audio object with respect to the time frame is quantized using a quantizer according to the bit allocation; When the bit allocation for the quantization of the direction value is less than the bit allocation threshold, a second comparison is made by comparing the direction value of the audio object for the time frame with the direction value of the audio object for the previous time frame. The method according to claims 11 and 12, including the method according to claims 11 and 12.

14. The method according to claims 11 to 13, wherein the quantizer is a spherical grid quantizer, the spherical grid is formed by covering the sphere with smaller spheres, and the smaller spheres define points on the spherical grid.

15. The method according to claims 11 to 14, wherein the directional value includes an azimuth angle value and an elevation angle value.

16. Method: Comparing the bit assignment of the direction value index for the time frame of an audio object with the bit assignment threshold; In accordance with the comparison, the directional index of the audio object to the time frame is decoded using an inverse quantizer according to the bit assignment, or the state of the signal flag associated with the directional index of the audio object to the time frame is read; Depending on the state of the signal flag; Set the quantized direction value of the audio object for the time frame to the quantized direction value of the audio object for the previous time frame, or The directional index of the audio object for the time frame is decoded using an inverse quantizer according to the bit allocation to provide a quantized directional value of the audio object for the time frame, and the azimuth angle value of the quantized directional value of the audio object for the time frame is adjusted according to the quantization resolution of the azimuth angle value. The method, including the method described above.

17. Adjusting the azimuth angle value of the quantized direction value of the audio object with respect to the time frame according to the quantization resolution of the azimuth angle value is: This includes comparing the product of the azimuth angle value of the quantized direction value of the audio object for the time frame with the azimuth angle value of the quantized direction value for the previous time frame, provided that the product is greater than zero: Determining the difference between two consecutive azimuth quantization values ​​of the inverse quantizer; When the difference between the azimuth angle value of the quantized direction value of the audio object for the time frame and the azimuth angle value of the quantized direction value for the previous time frame is greater than half the difference between two consecutive azimuth angle quantized values ​​of the inverse quantizer, subtract half the difference between two consecutive azimuth angle quantized values ​​of the inverse quantizer from the azimuth angle value of the quantized direction value of the audio object for the time frame; When the difference between the azimuth angle value of the quantized direction value of the audio object for the previous time frame and the azimuth angle value of the quantized direction value for the time frame is greater than half the difference between two consecutive azimuth angle quantized values ​​of the inverse quantizer, half the difference between two consecutive azimuth angle quantized values ​​of the inverse quantizer is added to the azimuth angle value of the quantized direction value of the audio object for the time frame. The method according to claim 16, including the method described in claim 16.

18. Determining the difference between two consecutive azimuth quantization values ​​of the aforementioned inverse quantizer is: The method according to claim 17, comprising dividing the circumference by a number of azimuth angle values, wherein the number of azimuth angle values ​​is determined by the elevation angle value of the quantized direction value with respect to the time frame.

19. The state of the aforementioned signal flag is: The direction value of the audio object with respect to the time frame is different from the direction value of the audio object with respect to the previous time frame; or The direction value of the audio object with respect to the time frame is the same as the direction value of the audio object with respect to the previous time frame; The method according to claims 16 to 18, relating to one of the following.

20. The method according to claims 16 to 19, wherein the inverse quantizer is a spherical grid quantizer, the spherical grid is formed by covering the sphere with smaller spheres, and the smaller spheres define points on the spherical grid.