Spatial audio parameter encoding and related decoding

By using companding and decompressing coding techniques, the directional parameters of multi-channel audio signals are quantized and encoded, solving the problem of poor codec performance at low bit rates and achieving efficient spatial audio signal transmission and reconstruction.

CN116508332BActive Publication Date: 2026-05-19NOKIA TECHNOLOGIES OY
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2021-08-18
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively compress and transmit directional metadata of multi-channel audio signals, especially at low bit rates, resulting in poor codec performance.

Method used

The direction parameter values ​​of multi-channel audio signals are quantized and encoded using companding and decompressing methods. The companding function is used to generate companding azimuth elements, and the encoding process is optimized by quantization grid and energy ratio.

Benefits of technology

It improves codec performance at low bit rates, reduces storage requirements and computational complexity, while maintaining high-quality spatial audio reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116508332B_ABST
    Figure CN116508332B_ABST
Patent Text Reader

Abstract

An apparatus comprising components configured to: obtain a multi-channel audio signal; obtain direction parameter values (301) associated with at least two time-frequency portions of the multi-channel audio signal, the direction parameter values associated with the at least two time-frequency portions comprising elevation elements and azimuth elements associated with the at least two time-frequency portions; and compand the obtained direction parameter values (305), the components configured to compand the obtained direction parameter values further configured to: quantize the elevation elements; determine a companding function based on the quantized elevation elements and / or a format of the multi-channel audio signal; generate companded azimuth elements based on the companding function applied to the azimuth elements; and quantize the companded azimuth elements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to apparatus and methods for encoding sound field related parameters, but not exclusively to apparatus and methods for encoding time-frequency domain directional related parameters for audio encoders and decoders. Background Technology

[0002] Parametric spatial audio processing is a field of audio signal processing where the spatial aspect of sound is described using a set of parameters. For example, in parametric spatial audio capture from a microphone array, estimating a set of directional metadata parameters from the microphone array signal is a typical and effective choice, such as the direction of sound in the frequency band and the ratio between the directional and non-directional portions of the captured sound in the band. These parameters are well-known for describing the perceived spatial characteristics of the sound captured at the location of the microphone array. These parameters can then be used accordingly for the synthesis of spatial sound, in binary form for headphones, for speakers, or for other formats such as high-fidelity stereo reproduction (Ambisonics).

[0003] Therefore, directional metadata such as direction in the frequency band and direction-to-total energy ratios are particularly effective parameterizations for spatial audio capture.

[0004] A set of directional metadata parameters, consisting of one or more directional values ​​for each frequency band and an energy ratio parameter associated with each directional value, can also be used as spatial metadata for an audio codec (which may also include other parameters such as spread coherence, number of directions, distance, etc.). The set of directional metadata parameters may also include other parameters, or may be associated with other parameters considered non-directional, such as surround coherence, diffuse-to-total energy ratio, and remainder-to-total energy ratio. For example, these parameters can be estimated from audio signals captured by a microphone array, and, for example, stereo signals can be generated from microphone array signals to be transmitted along with the spatial metadata.

[0005] Because some codecs are expected to operate at a wide range of bit rates, from very low to relatively high, various strategies are needed to compress spatial metadata to optimize codec performance at each point of operation. The raw bit rate of the encoded parameters (metadata) is relatively high, so especially at lower bit rates, it is expected that only the most important parts of the metadata can be transmitted from the encoder to the decoder.

[0006] The decoder can decode audio signals into PCM signals and process the sound in the frequency band (using spatial metadata) to obtain spatial output, such as binaural output.

[0007] The above solution is particularly well-suited for encoding captured spatial sound from microphone arrays (e.g., in mobile phones, cameras, VR cameras, standalone microphone arrays). However, for such encoders, it is desirable to have other input types besides the signals captured by the microphone array, such as speaker signals, audio object signals, or Ambisonics signals. Summary of the Invention

[0008] According to a first aspect, an apparatus is provided, comprising components configured to: acquire a multi-channel audio signal; acquire directional parameter values ​​associated with at least two time-frequency portions of the multi-channel audio signal, the directional parameter values ​​associated with the at least two time-frequency portions including an elevation element and an azimuth element associated with the at least two time-frequency portions; and compand encode the acquired directional parameter values, the components configured to compand encode the acquired directional parameter values ​​further configured to: quantize the elevation element; determine a companding function based on the quantized elevation element and / or the multi-channel audio signal format; generate a companded azimuth element based on the companding function applied to the azimuth element; and quantize the companded azimuth element.

[0009] This component can also be configured to decompress the quantized companded azimuth elements based on the inverse of the companding function.

[0010] The component configured to determine the companding function based on the quantized elevation angle element and / or the multi-channel audio signal format can also be configured to determine the companding function based on the quantized elevation angle element and the multi-channel audio signal format.

[0011] The component configured to compand encode the acquired direction parameter values ​​can also be configured to generate codewords for each quantized elevation element and quantized companded azimuth element.

[0012] The component configured to compand encode the acquired direction parameter values ​​can also be configured to generate codewords for each quantized elevation element and decompressed quantized azimuth element.

[0013] The component can also be configured to determine the quantization error and average elevation angle coding for companding, wherein the component configured to determine the average elevation angle coding can be configured to: quantize the average elevation angle elements of the subbands within the frame; and quantize the azimuth elements based on a quantization raster with variable boundaries, and wherein the component is configured to: select either the companded coding output or the average elevation angle coding output based on the quantization error.

[0014] The component can also be configured to determine the quantization grid based on an allocated number of bits used to encode each subband within a frame, including subbands and time blocks, based on the value of the energy ratio, which is associated with the acquired direction parameter value. The component configured to quantize elevation elements is configured to quantize elevation elements based on the quantization grid, and the component configured to quantize companded azimuth elements is configured to quantize companded azimuth based on the quantization grid.

[0015] According to a second aspect, an apparatus is provided, comprising components configured to: acquire at least one coded bitstream, the at least one coded bitstream including an coded multichannel audio signal and companded coded direction parameter values, the companded coded direction parameter values ​​being associated with at least two time-frequency portions of the coded multichannel audio signal, and the coded direction parameter values ​​associated with the at least two time-frequency portions including coded elevation elements and companded azimuth elements associated with the at least two time-frequency portions; decode the coded elevation elements; determine a decompanoplastin function based on the coded elevation elements and / or the multichannel audio signal format; and generate decompanoplastin azimuth elements based on the decompanoplastin function applied to the companded azimuth elements.

[0016] The component configured to determine the decompression spreading function based on the encoded elevation angle element and / or the multi-channel audio signal format can also be configured to determine the decompression spreading function based on the encoded elevation angle element and the multi-channel audio signal format.

[0017] The component configured to decode the encoded elevation elements can also be configured to decode the codeword for each quantized elevation element.

[0018] The component can also be configured to: determine the quantization grid based on the allocated number of bits, the allocated number of bits being used to encode each sub-band within a frame, including sub-bands and time blocks, based on the value of the energy ratio, the value of the energy ratio being associated with the acquired direction parameter value, wherein the component configured to decode the codeword for each quantized elevation element can be configured to: decode the elevation element based on the quantization grid.

[0019] According to a third aspect, a method is provided, the method comprising: acquiring a multi-channel audio signal; acquiring directional parameter values ​​associated with at least two time-frequency components of the multi-channel audio signal, the directional parameter values ​​associated with the at least two time-frequency components including elevation and azimuth elements associated with the at least two time-frequency components; and companding the acquired directional parameter values, wherein companding the acquired directional parameter values ​​comprises: quantizing the elevation element; determining a companding function based on the quantized elevation element and / or the multi-channel audio signal format; generating a companded azimuth element based on the companding function applied to the azimuth element; and quantizing the companded azimuth element.

[0020] The method may also include decompressing the quantized companded azimuth elements based on the inverse of the companding function.

[0021] Determining the companding function based on the quantized elevation angle element and / or the multi-channel audio signal format may also include: determining the companding function based on the quantized elevation angle element and the multi-channel audio signal format.

[0022] Companding the acquired direction parameter values ​​may also include generating codewords for each quantized elevation element and quantized companded azimuth element.

[0023] Companding the acquired direction parameter values ​​may also include generating codewords for each quantized elevation element and the decompressed quantized azimuth element.

[0024] The method may further include determining the quantization error and average elevation angle coding of the companding coding, wherein the average elevation angle coding may include: quantizing the average elevation angle elements of the subbands within the frame; and quantizing the azimuth elements based on a quantization grid with variable boundaries, and the method may further include selecting the companding coding output or the average elevation angle coding output based on the quantization error.

[0025] The method may further include: determining a quantization grid based on an allocated number of bits, the allocated number of bits being used to encode each subband within a frame, including subbands and time blocks, based on the value of an energy ratio, the value of which is associated with an acquired direction parameter value, wherein quantizing elevation elements may include quantizing elevation elements based on a quantization grid, and quantizing companded azimuth elements may include quantizing companded azimuth based on a quantization grid.

[0026] According to a fourth aspect, a method is provided, the method comprising: acquiring at least one coded bitstream, the at least one coded bitstream including a coded multi-channel audio signal and companded coded direction parameter values, the companded coded direction parameter values ​​being associated with at least two time-frequency portions of the coded multi-channel audio signal, and the coded direction parameter values ​​associated with the at least two time-frequency portions including coded elevation elements and companded azimuth elements associated with the at least two time-frequency portions; decoding the coded elevation elements; determining a decompanopling function based on the coded elevation elements and / or the format of the coded multi-channel audio signal; and generating decompanopling azimuth elements based on the decompanopling function applied to the companded azimuth elements.

[0027] Determining the decompression spreading function based on the encoded elevation angle element and / or the multi-channel audio signal format may also include: determining the decompression spreading function based on the encoded elevation angle element and the multi-channel audio signal format.

[0028] Decoding the encoded elevation elements may also include decoding the codeword used for each quantized elevation element.

[0029] The method may further include: determining a quantization grid based on an allocated number of bits, the allocated number of bits being used to encode each subband within a frame, including subbands and time blocks, based on a value of an energy ratio associated with an acquired direction parameter value, wherein a component configured to decode the codeword for each quantized elevation element is configured to: decode the elevation element based on the quantization grid.

[0030] According to a fifth aspect, an apparatus is provided, comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured, together with the at least one processor, to cause the apparatus to at least: acquire a multi-channel audio signal; acquire direction parameter values ​​associated with at least two time-frequency portions of the multi-channel audio signal, the direction parameter values ​​associated with the at least two time-frequency portions including an elevation element and an azimuth element associated with the at least two time-frequency portions; and compand encode the acquired direction parameter values, wherein the apparatus causing to compand encode the acquired direction parameter values ​​is further caused to: quantize the elevation element; determine a companding function based on the quantized elevation element and / or the multi-channel audio signal format; generate a companded azimuth element based on the companding function applied to the azimuth element; and quantize the companded azimuth element.

[0031] The device can also be used to decompress the quantized companded azimuth elements based on the inverse of the companding function.

[0032] The apparatus that enables the determination of the companding function based on the quantized elevation element and / or the multi-channel audio signal format can also be configured to: determine the companding function based on the quantized elevation element and the multi-channel audio signal format.

[0033] The apparatus that compandsizes the acquired directional parameter values ​​can also be configured to generate codewords for each quantized elevation element and quantized companded azimuth element.

[0034] The apparatus that compandsizes the acquired directional parameter values ​​can also be configured to generate codewords for each quantized elevation element and each decompressed quantized companded azimuth element.

[0035] The apparatus can also be configured to determine the quantization error and average elevation angle coding of the companding coding, wherein the means of determining the average elevation angle coding can also be configured to: quantize the average elevation angle elements of the subbands within the frame; and quantize the azimuth elements based on a quantization grid with variable boundaries, and wherein the apparatus can also be configured to select either the companding coding output or the average elevation angle coding output based on the quantization error.

[0036] The apparatus can also be configured to: determine a quantization grid based on an allocated number of bits, the allocated number of bits being used to encode each subband within a frame, including subbands and time blocks, based on the value of an energy ratio associated with an acquired direction parameter value, wherein the means for quantizing elevation elements can be configured to quantize elevation elements based on a quantization grid, and the means for quantizing companded azimuth elements can be configured to quantize companded azimuth elements based on a quantization grid.

[0037] According to a sixth aspect, an apparatus is provided, comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured together with the at least one processor such that the apparatus at least: acquires at least one coded bitstream, the at least one coded bitstream comprising: an coded multichannel audio signal and companded coded direction parameter values, the companded coded direction parameter values ​​being associated with at least two time-frequency portions of the coded multichannel audio signal, and the coded direction parameter values ​​associated with the at least two time-frequency portions including coded elevation elements and companded azimuth elements associated with the at least two time-frequency portions; decodes the coded elevation elements; determines a decompanoplastin function based on the coded elevation elements and / or the format of the multichannel audio signal; and generates decompanoplastin azimuth elements based on the decompanoplastin function applied to the companded azimuth elements.

[0038] The apparatus for determining the decompression spreading function based on the encoded elevation angle element and / or the multi-channel audio signal format can also be configured to: determine the decompression spreading function based on the encoded elevation angle element and the multi-channel audio signal format.

[0039] The apparatus that enables the decoding of the encoded elevation elements can also be made to decode the codeword used for each quantized elevation element.

[0040] The apparatus can also be configured to: determine a quantization grid based on an allocated number of bits, the allocated number of bits being used to encode each subband within a frame, including subbands and time blocks, based on the value of an energy ratio associated with an acquired direction parameter value, wherein the apparatus configured to decode the codeword for each quantized elevation element can be configured to: decode the elevation element based on the quantization grid.

[0041] According to a seventh aspect, an apparatus is provided, comprising: means for acquiring a multi-channel audio signal; means for acquiring direction parameter values ​​associated with at least two time-frequency portions of the multi-channel audio signal, the direction parameter values ​​associated with the at least two time-frequency portions including an elevation element and an azimuth element associated with the at least two time-frequency portions; and means for companding the acquired direction parameter values, wherein the means for companding the acquired direction parameter values ​​includes: means for quantizing the elevation element; means for determining a companding function based on the quantized elevation element and / or the multi-channel audio signal format; means for generating a companded azimuth element based on a companding function applied to the azimuth element; and means for quantizing the companded azimuth element.

[0042] According to an eighth aspect, an apparatus is provided, comprising: means for acquiring at least one coded bitstream, the at least one coded bitstream including an coded multichannel audio signal and companded coded direction parameter values, the companded coded direction parameter values ​​being associated with at least two time-frequency portions of the coded multichannel audio signal, and the coded direction parameter values ​​associated with the at least two time-frequency portions including coded elevation elements and companded azimuth elements associated with the at least two time-frequency portions; means for decoding the coded elevation elements; means for determining a decompanoplastin function based on the coded elevation elements and / or the format of the multichannel audio signal; and means for generating a decompanoplastin azimuth element based on the decompanoplastin function applied to the companded azimuth element.

[0043] According to a ninth aspect, a computer program [or a computer-readable medium including program instructions] is provided, the instructions being configured to cause a device to perform at least the following operations: acquiring a multi-channel audio signal; acquiring direction parameter values ​​associated with at least two time-frequency portions of the multi-channel audio signal, the direction parameter values ​​associated with the at least two time-frequency portions including elevation and azimuth elements associated with the at least two time-frequency portions; and companding the acquired direction parameter values, wherein companding the acquired direction parameter values ​​includes: quantizing the elevation element; determining a companding function based on the quantized elevation element and / or the multi-channel audio signal format; generating a companded azimuth element based on the companding function applied to the azimuth element; and quantizing the companded azimuth element.

[0044] According to a tenth aspect, a computer program [or a computer-readable medium including program instructions] is provided, the instructions being configured to cause an apparatus to perform at least the following operations: acquiring at least one coded bitstream, the at least one coded bitstream including an coded multichannel audio signal and companded coded direction parameter values, the companded coded direction parameter values ​​being associated with at least two time-frequency portions of the coded multichannel audio signal, and the coded direction parameter values ​​associated with the at least two time-frequency portions including coded elevation elements and companded coded azimuth elements associated with the at least two time-frequency portions; decoding the coded elevation elements; determining a decompanoplastin function based on the coded elevation elements and / or the multichannel audio signal format; and generating a decompanoplastin azimuth element based on the decompanoplastin function applied to the companded azimuth elements.

[0045] According to an eleventh aspect, a non-transitory computer-readable medium is provided including program instructions for causing a device to perform at least the following operations: acquiring a multi-channel audio signal; acquiring direction parameter values ​​associated with at least two time-frequency portions of the multi-channel audio signal, the direction parameter values ​​associated with the at least two time-frequency portions including elevation and azimuth elements associated with the at least two time-frequency portions; and companding the acquired direction parameter values, wherein companding the acquired direction parameter values ​​includes: quantizing the elevation element; determining a companding function based on the quantized elevation element and / or the multi-channel audio signal format; generating a companded azimuth element based on the companding function applied to the azimuth element; and quantizing the companded azimuth element.

[0046] According to a twelfth aspect, a non-transitory computer-readable medium is provided comprising program instructions for causing an apparatus to perform at least the following operations: acquiring at least one coded bitstream, the at least one coded bitstream comprising: an coded multichannel audio signal and companded coded direction parameter values, the companded coded direction parameter values ​​being associated with at least two time-frequency portions of the coded multichannel audio signal, and the coded direction parameter values ​​associated with the at least two time-frequency portions including coded elevation elements and companded coded azimuth elements associated with the at least two time-frequency portions; decoding the coded elevation elements; determining a decompanoplastin function based on the coded elevation elements and / or the format of the multichannel audio signal; and generating a decompanoplastin azimuth element based on the decompanoplastin function applied to the companded azimuth elements.

[0047] According to a thirteenth aspect, an apparatus is provided, comprising: an acquisition circuit system configured to acquire a multi-channel audio signal; an acquisition circuit system configured to acquire directional parameter values ​​associated with at least two time-frequency portions of the multi-channel audio signal, the directional parameter values ​​associated with the at least two time-frequency portions including an elevation element and an azimuth element associated with the at least two time-frequency portions; and an encoding circuit system configured to compand encode the acquired directional parameter values, wherein the encoding circuit system configured to compand encode the acquired directional parameter values ​​is configured to: quantize the elevation element; determine a companding function based on the quantized elevation element and / or the multi-channel audio signal format; generate a companded azimuth element based on the companding function applied to the azimuth element; and quantize the companded azimuth element.

[0048] According to a fourteenth aspect, an apparatus is provided, comprising: an acquisition circuit system for acquiring at least one coded bitstream, the at least one coded bitstream including an coded multichannel audio signal and companded coded direction parameter values, the companded coded direction parameter values ​​being associated with at least two time-frequency portions of the coded multichannel audio signal, and the coded direction parameter values ​​associated with the at least two time-frequency portions including coded elevation elements and companded azimuth elements associated with the at least two time-frequency portions; a decoding circuit system configured to decode the coded elevation elements; determine a decompanoplastin function based on the coded elevation elements and / or the format of the multichannel audio signal; and a generation circuit system configured to generate decompanoplastin azimuth elements based on the decompanoplastin function applied to the companded azimuth elements.

[0049] According to a fifteenth aspect, a computer-readable medium including program instructions is provided to cause a device to perform at least the following operations: acquiring a multi-channel audio signal; acquiring direction parameter values ​​associated with at least two time-frequency portions of the multi-channel audio signal, the direction parameter values ​​associated with the at least two time-frequency portions including elevation and azimuth elements associated with the at least two time-frequency portions; and companding the acquired direction parameter values, wherein companding the acquired direction parameter values ​​includes: quantizing the elevation element; determining a companding function based on the quantized elevation element and / or the multi-channel audio signal format; generating a companded azimuth element based on the companding function applied to the azimuth element; and quantizing the companded azimuth element.

[0050] According to a sixteenth aspect, a computer-readable medium including program instructions is provided to cause an apparatus to perform at least the following operations: acquiring at least one coded bitstream, the at least one coded bitstream including an coded multichannel audio signal and companded coded direction parameter values, the companded coded direction parameter values ​​being associated with at least two time-frequency portions of the coded multichannel audio signal, and the coded direction parameter values ​​associated with the at least two time-frequency portions including coded elevation elements and companded azimuth elements associated with the at least two time-frequency portions; decoding the coded elevation elements; determining a decompanoplastin function based on the coded elevation elements and / or the multichannel audio signal format; and generating a decompanoplastin azimuth element based on the decompanoplastin function applied to the companded azimuth element.

[0051] An apparatus includes components for performing the actions described above.

[0052] An apparatus is configured to perform the actions described above.

[0053] A computer program includes program instructions for causing a computer to perform the methods described above.

[0054] A computer program product stored on a medium can cause a device to perform the methods described herein.

[0055] An electronic device may include the means as described herein.

[0056] A chipset may include devices as described herein.

[0057] The embodiments of this application are intended to solve problems associated with the prior art. Attached Figure Description

[0058] To better understand this application, reference will now be made to the accompanying drawings by way of example, in which:

[0059] Figure 1 A system suitable for implementing some embodiments is illustrated schematically;

[0060] Figure 2 An encoder according to some embodiments is illustrated schematically;

[0061] Figure 3 Examples of some embodiments are shown. Figure 2 The flowchart shown illustrates the operation of the encoder;

[0062] Figure 4 The illustration shows, for example, some embodiments. Figure 2 The direction encoder shown;

[0063] Figure 5 Examples of some embodiments are shown. Figure 4 The flowchart shown illustrates the operation of the orientation encoder;

[0064] Figure 6 and Figure 7 It shows that it is suitable for use in, for example Figure 4 The companderode function implemented in the directional encoder is shown.

[0065] Figure 8 The illustration shows, for example, some embodiments. Figure 2 The decoder shown;

[0066] Figure 9 Examples of some embodiments are shown. Figure 8 The flowchart of the decoder operation is shown; and

[0067] Figure 10 An example device suitable for implementing the illustrated apparatus is shown schematically. Detailed Implementation

[0068] The following describes in further detail suitable apparatuses and possible mechanisms for providing metadata parameters derived from combination and coding space analysis. In the following discussion, multi-channel systems are discussed in relation to multi-channel microphone implementations. However, as mentioned above, the input format can be any suitable input format, such as multi-channel speakers, Ambisonics (FOA / HOA), etc. It should be understood that in some embodiments, channel positions are based on the microphone's location, or a virtual location or orientation.

[0069] Furthermore, in the following examples, the output of the example system is a multi-channel speaker arrangement. In other embodiments, the output may be rendered to the user via components other than speakers. The multi-channel speaker signal may also be summarized as two or more playback audio signals.

[0070] As mentioned above, the directional metadata associated with the audio signal can include multiple parameters (such as multiple directions and the direction-to-total ratio, distance, etc. associated with each direction) for each time-frequency tile. The directional metadata can also include other parameters, or can be associated with other parameters considered non-directional (such as surround coherence, spread-to-total energy ratio, residual-to-total energy ratio), but when combined with directional parameters, it can be used to define the characteristics of the audio scene. For example, a reasonable design choice that produces high-quality output is one where the directional metadata includes two directions (and the direction-to-total ratio, distance value, etc. associated with each direction) for each time-frequency subframe. However, as mentioned above, bandwidth and / or storage limitations may require the codec not to send directional metadata parameter values ​​for each frequency band and time subframe.

[0071] Current proposals include those disclosed in GB patent application 1811071.8, which takes into account lossy compression of metadata, and vector quantization methods have been discussed for PCT / FI2019 / 050675 when the number of bits available for a given subband is very low. Even with a codebook of only up to 9 bits, vector quantization methods increase the codec's table ROM, where approximately 4kB of memory is used for 2, 3, 4, ... and 9-bit 4D codebooks.

[0072] The concept discussed in the embodiments herein is to provide a low-complexity codec with low ROM imprint that takes into account the characteristics of multi-channel directional metadata.

[0073] Although codecs such as those in UK patent application GB2000465.1 have considered lossy compression of metadata, the proposed flexible orientation codebook is uniformly distributed, meaning that for 3 bits, only the front, back, lateral, and center positions can be represented. However, it is useful to consider the representation of channel positions in multi-channel formats. Furthermore, the embodiments discussed herein offer improved performance compared to non-uniform scalar codebook implementations because it does not require storing the codebook for every possible number of bits (in other words, the embodiments require less codebook storage).

[0074] In the following embodiments, the codec employs a unified quantizer structure, but may optionally implement (e.g., based on the channel input format) an adjustable parameterized companding function.

[0075] about Figure 1Example apparatus and system for implementing embodiments of this application are shown. System 100 is shown having an "analysis" section 121 and a "synthesis" section 131. The "analysis" section 121 is the section from receiving multi-channel signals to encoding directional metadata and transmission signals, while the "synthesis" section 131 is the section from decoding the encoded directional metadata and transmission signals to the presentation of the regenerated signal (e.g., in the form of a multi-channel loudspeaker).

[0076] In the following description, the “analysis” section 121 is described as a series of parts; however, in some embodiments, this section may be implemented as a function of the same functional means or within a section. In other words, in some embodiments, the “analysis” section 121 is an encoder that includes at least one of a transmission signal generator or an analysis processor as described below.

[0077] The input to system 100 and the "analysis" section 121 is a multi-channel signal 102. The "analysis" section 121 may include a transmit signal generator 103, an analysis processor 105, and an encoder 107. In the following example, a microphone channel signal input is described; however, in other embodiments, any suitable input (or synthesized multi-channel) format can be implemented. In such embodiments, directional metadata associated with the audio signal may be provided to the encoder as a separate bitstream. The multi-channel signal is passed to the transmit signal generator 103 and the analysis processor 105.

[0078] In some embodiments, the transmit signal generator 103 is configured to receive multi-channel signals and generate a suitable audio signal format for encoding. The transmit signal generator 103 may, for example, generate stereo or single-channel audio signals. The transmit audio signal generated by the transmit signal generator can be any known format. For example, when the input is an audio signal from a mobile phone microphone array, the transmit signal generator 103 may be configured to select left and right microphone pairs and apply any suitable processing to the audio signal pairs, such as automatic gain control, microphone noise removal, wind noise removal, and equalization. In some embodiments, when the input is a first-order Ambisonic / higher-order Ambisomic (FOA / HOA) signal, the transmit signal generator may be configured to formulate directional beam signals directed to the left and right, such as two opposing cardioid signals. Furthermore, in some embodiments, when the input is a speaker surround mix and / or an object, the transmit signal generator 103 can be configured to generate a downmix signal that combines the left channel into a left downmix channel, combines the right channel into a right downmix channel, and adds the center channel to both transmit channels with appropriate gain.

[0079] In some embodiments, the transmit signal generator is bypassed (or, in other words, optional). For example, in some cases where analysis and synthesis occur at the same device in a single processing step, no transmit signal is generated without intermediate processing, and the input audio signal is transmitted unprocessed. The number of transmit channels generated can be any suitable number, rather than, for example, one or two channels.

[0080] The output of the signal generator 103 can be transmitted to the encoder 107.

[0081] In some embodiments, the analysis processor 105 is also configured to receive and analyze the multichannel signal to generate directional metadata 106 associated with the multichannel signal and therefore with the transmission signal 104.

[0082] The analysis processor 105 can be configured to generate directional metadata parameters, which, for each time-frequency analysis interval, may include at least one direction parameter 108 and at least one energy ratio parameter 110 (and in some embodiments, may also include other parameters, a non-exhaustive list of which includes the number of directions, circumferential coherence, diffusion-to-total energy ratio, residual-to-total energy ratio, extended coherence parameter, and distance parameter). The direction parameter can be represented in any suitable manner, for example, as spherical coordinates, where the spherical coordinates represent azimuth angles. and elevation angle .

[0083] In some embodiments, the number of directional metadata parameters may vary across time-frequency blocks. Thus, for example, in band X, all directional metadata parameters are acquired (generated) and transmitted, while in band Y, only one of the directional metadata parameters is acquired and transmitted; furthermore, in band Z, no parameters are acquired or transmitted. A practical example of this could be that for some time-frequency blocks corresponding to the highest frequency band, some directional metadata parameters are unnecessary for perceptual reasons. Directional metadata 106 can be passed to encoder 107.

[0084] In some embodiments, the analysis processor 105 is configured to apply a time-frequency transform to the input signal. Then, for example, in a time-frequency plot, when the input is a mobile phone microphone array, the analysis processor can be configured to estimate delay values ​​between microphone pairs that maximize the inter-microphone correlation. Based on these delay values, the analysis processor can then be configured to determine corresponding direction values ​​for directional metadata. Furthermore, the analysis processor can be configured to determine a direction-to-total ratio parameter based on the correlation values.

[0085] In some embodiments, for example, when the input is a FOA signal, the analysis processor 105 can be configured to determine an intensity vector. The analysis processor can then be configured to determine the direction parameter values ​​of the directional metadata based on the intensity vector. The spread-to-total ratio can then be determined, thereby determining the direction and total ratio parameter values ​​of the directional metadata. This analysis method is referred to in the literature as Directional Audio Coding (DirAC).

[0086] In some examples, such as when the input is a HOA signal, the analysis processor 105 can be configured to divide the HOA signal into multiple sectors, using the method described above in each sector. This sector-based method is referred to in the literature as Higher-Order DirAC (HO DirAC). In these examples, each time-frequency patch corresponding to multiple sectors has more than one simultaneous direction parameter value.

[0087] Furthermore, in some embodiments where the input is speaker surround mix and / or an audio object-based signal, the analysis processor can be configured to convert the signal into a FOA / HOA signal format and acquire the direction and direction-to-total ratio parameter values ​​as described above.

[0088] Encoder 107 may include an audio encoder core 109 configured to receive transmitted audio signals 104 and generate suitable encodings of these audio signals. In some embodiments, encoder 107 may be a computer (running suitable software stored in memory and at least one processor), or alternatively, a specific device utilizing, for example, an FPGA or ASIC. Audio encoding may be implemented using any suitable scheme.

[0089] Encoder 107 may further include a directional metadata encoder / quantizer 111, configured to receive directional metadata and output an encoded or compressed form of the information. In some embodiments, encoder 107 may further interleave the directional metadata, multiplex the directional metadata into a single data stream, or embed the directional metadata into an encoded downmixer signal before transmission or storage, such as... Figure 1 As shown by the dashed line in the diagram. Multiplexing can be implemented using any suitable scheme.

[0090] In some embodiments, the transmission signal generator 103 and / or analysis processor 105 may be located on a device separate from (or otherwise separate from) the encoder 107. For example, in such embodiments, directional metadata (and associated non-directional metadata) parameters associated with the audio signal may be provided to the encoder as a separate bitstream.

[0091] In some embodiments, the transmission signal generator 103 and / or analysis processor 105 may be part of the encoder 107, i.e., located inside the encoder and on the same device.

[0092] In the following description, the “composite” portion 131 is described as a series of portions; however, in some embodiments, the portion may be implemented as the same functional device or a function within the portion.

[0093] On the decoder side, the received or retrieved data (stream) can be received by decoder / demultiplexer 133. Decoder / demultiplexer 133 can demultiplex the encoded stream and pass the audio encoded stream to transport signal decoder 135, which is configured to decode the audio signal to obtain the transport audio signal. Similarly, decoder / demultiplexer 133 may include metadata decoder 137, which is configured to receive encoded directional metadata (e.g., a direction index representing a direction parameter value) and generate directional metadata.

[0094] In some embodiments, the decoder / demultiplexer 133 may be a computer (running suitable software stored in memory and at least one processor), or alternatively, a specific device utilizing, for example, an FPGA or ASIC.

[0095] The decoded metadata and transmitted audio signals can be passed to the synthesis processor 139.

[0096] The “synthesis” section 131 of system 100 also shows a synthesis processor 139, which is configured to receive transmitted audio signals and directional metadata, and recreate the synthesized spatial audio in the form of a multi-channel signal 110 based on the transmitted signals and directional metadata in any suitable format (these may be multi-channel speaker formats, or in some embodiments, depending on the use case, any suitable output format, such as binaural or binaural signals).

[0097] The synthesis processor 139 therefore creates the output audio signal based on any suitable known method, such as a multi-channel speaker signal or a binaural signal. This will not be explained in detail here. However, as a simplified example, the speaker output can be rendered according to any of the following methods. For example, the transmitted audio signal can be divided into a directional stream and an ambient stream based on the ratio of direction to total energy and the ratio of diffusion to total energy. The directional stream can then be rendered using amplitude translation based on the direction parameters(s). The ambient stream can be further rendered using decorrelation. The directional stream and the ambient stream can then be combined.

[0098] The output signal can be reproduced using a multi-channel speaker setup or headphones.

[0099] It should be noted that Figure 1 The processing blocks can reside in the same or different processing entities. For example, in some embodiments, microphone signals from a mobile device are processed using a spatial audio capture system (including an analysis processor and a transmit signal generator), and the resulting spatial metadata and transmit audio signals (e.g., in the form of a MASA stream) are forwarded to an encoder (e.g., an IVAS encoder) that includes the aforementioned encoder. In other embodiments, the input signal (e.g., a 5.1-channel audio signal) is forwarded directly to an encoder (e.g., an IVAS encoder) that includes the aforementioned analysis processor, transmit signal generator, and encoder.

[0100] In some embodiments, there may be two (or more) input audio signals, wherein the first audio signal is generated by... Figure 1 The illustrated apparatus processes the data (to generate data as input to the encoder), and the second audio signal is directly forwarded to the encoder (e.g., an IVAS encoder), which includes the aforementioned analysis processor, transmission signal generator, and encoder. The audio input signal can then be encoded independently in the encoder, or it can be combined in the parameter domain, for example, according to so-called MASA mixing.

[0101] In some embodiments, a compositing portion may exist comprising separate decoder and compositing processor entities or devices, or the compositing portion may comprise a separate entity comprising both a decoder and a compositing processor. In some embodiments, a decoder block may process more than one input data stream in parallel. In applications, the term compositing processor may be interpreted as an internal or external renderer.

[0102] Therefore, in general, firstly, the system (analysis section) is configured to receive multi-channel audio signals. Then, the system (analysis section) is configured to generate suitable transmission audio signals (e.g., by selecting some of the audio signal channels). Next, the system is configured to encode the transmission audio signals for storage / transmission. After this, the system can store / transmit the encoded transmission audio signals and metadata. The system can retrieve / receive the encoded transmission audio signals and metadata. Then, the system is configured to extract the transmission audio signals and metadata from the encoded transmission audio signals and metadata parameters, such as by demultiplexing and decoding the encoded transmission audio signals and metadata parameters.

[0103] The system (synthesis section) is configured to synthesize the output multi-channel audio signal based on the extracted transmitted audio signal and metadata.

[0104] about Figure 2Further detailed description of example analysis processor 105 and metadata encoder / quantizer 111 according to some embodiments (e.g. Figure 1 (As shown).

[0105] In some embodiments, the analysis processor 105 includes a time-frequency domain converter 201.

[0106] In some embodiments, the time-frequency domain converter 201 is configured to receive the multi-channel signal 102 and apply a suitable time-to-frequency domain transformation, such as the short-time Fourier transform (STFT), to convert the input time-domain signal into a suitable time-frequency signal. These time-frequency signals can be passed to the spatial analyzer 203 and the direction encoder 205.

[0107] Therefore, for example, the time-frequency signal 202 can be represented in the time-frequency domain as:

[0108] ,

[0109] Where b is the frequency bin index, n is the time-frequency block (frame) index, and i is the channel index. In another expression, n can be considered as the time index with a sampling rate lower than that of the original time-domain signal. These frequency bins can be grouped into subbands, which group one or more bins into subbands k=0, ..., K-1 with frequency band indices. Each subband k has the lowest interval... and the highest interval And the subband contains from arrive All intervals. The width of the subband can approximate any suitable distribution, such as the Equivalent Rectangular Bandwidth (ERB) scale or the Bark scale.

[0110] In some embodiments, the analysis processor 105 includes a spatial analyzer 203. The spatial analyzer 203 may be configured to receive time-frequency signals 202 and estimate direction parameters 108 based on these signals. The direction parameters may be determined based on any audio-based "direction" determination.

[0111] For example, in some embodiments, the spatial analyzer 203 is configured to estimate direction using two or more signal inputs. This represents the simplest configuration for estimating "direction," allowing for more complex processing on more signals.

[0112] Therefore, the spatial analyzer 203 can be configured to provide at least one azimuth and elevation angle for each frequency band and time-frequency block within a frame of the audio signal, denoted as azimuth angle. and elevation angle The direction parameter 108 can also be passed to the direction encoder 205.

[0113] The spatial analyzer 203 can also be configured to determine the energy ratio parameter 110. The energy ratio can be considered as the determination of the energy of an audio signal that can be considered to arrive from one direction. Direction is related to the total energy ratio. The energy ratio can be estimated, for example, using a stability metric for orientation estimation, or any relevant metric, or any other suitable method for obtaining the ratio parameter. The energy ratio can be passed to the energy ratio encoder 207.

[0114] Spatial analyzer 203 can also be configured to determine a number of coherence parameters 112, which may include the surround coherence analyzed in the time-frequency domain. ) and extended coherence ( ).

[0115] Therefore, in general, the analysis processor is configured to receive time-domain multichannel or other formats, such as microphone or dual-channel audio signals.

[0116] After this, the analysis processor can apply a time-to-frequency domain transformation (e.g., STFT) to generate a suitable time-frequency domain signal for analysis, and then apply directional analysis to determine the direction and energy ratio parameters.

[0117] The analyzer can then be configured to output the determined parameters.

[0118] Although the direction, energy ratio, and coherence parameter are represented here for each time index n, in some embodiments, these parameters can be combined across several time indices. The same applies to the frequency axis; as already expressed, the direction of several frequency intervals b can be expressed by a direction parameter in a frequency band k consisting of several frequency intervals b. This also applies to all spatial parameters discussed herein.

[0119] In some embodiments, the direction data can be represented using 16 bits, such that each azimuth parameter is approximately represented using 9 bits, and the elevation angle is represented using 7 bits. In such an embodiment, the energy ratio parameter can be represented using 8 bits. For each frame, there can be N subbands (where N can be between 1 and 24 and can be fixed at 5) and M time-frequency (TF) blocks (where the value of M can be M=4). Therefore, in this example, it is necessary to... Bits are used to store the uncompressed orientation and energy ratio metadata for each frame.

[0120] Similarly, Figure 2 As shown, an example metadata encoder / quantizer 111 is illustrated according to some embodiments.

[0121] The metadata encoder / quantizer 111 may include a direction encoder 205. The direction encoder 205 is configured to receive direction parameters (such as azimuth angle). and elevation angle 108 (and in some embodiments, the expected bit allocation) are received, thereby generating a suitable encoded output. In some embodiments, encoding is based on a quantization operation, where the quantization or codebook position is a spherical grid in which the arrangement of spheres is formed on a “surface” sphere in a ring, the ring being defined by a lookup table defined by a determined quantization resolution. In other words, the spherical grid uses the idea of ​​covering a sphere with smaller spheres, and regards the center of the smaller sphere as the point defining the grid of nearly equidistant orientations. Thus, the smaller sphere defines a cone or solid angle around the center point, which can be indexed according to any suitable indexing algorithm. Although spherical quantization is described herein, any suitable linear quantization raster can be used.

[0122] Then, the quantization values ​​can be further combined by determining whether the elevation angle values ​​of the corresponding directional parameters are sufficiently similar, using an embedded flexible boundary codebook.

[0123] The encoded direction parameter 206 can then be passed to the combiner 211.

[0124] The metadata encoder / quantizer 111 may include an energy ratio encoder 207. The energy ratio encoder 207 is configured to receive energy ratios and determine appropriate encoding for the energy ratios used in compressing subbands and time-frequency blocks. For example, in some embodiments, the energy ratio encoder 207 is configured to use 3 bits to encode each energy ratio parameter value.

[0125] Furthermore, in some embodiments, instead of transmitting or storing all energy ratios for all TF blocks, only a weighted average is transmitted or stored for each subband. This average can be determined by considering the total energy of each time block, thus favoring the values ​​of subbands with more energy.

[0126] In such an embodiment, the quantized energy ratio 208 is the same for all TF blocks of a given subband.

[0127] In some embodiments, the energy ratio encoder 207 is also configured to pass the quantized (encoded) energy ratio 208 to the combiner 211.

[0128] The metadata encoder / quantizer 111 may include a combiner 211. The combiner is configured to receive the direction parameter and energy ratio parameter of the encoded (or quantized / compressed) data and combine them to generate a suitable output (e.g., a metadata bitstream that can be transmitted or stored in combination with or separately from the transmitted signal).

[0129] about Figure 3 This illustrates, according to some embodiments, such as Figure 2 The example operation of the direction encoder / quantizer is shown.

[0130] The initial operation is to obtain metadata (such as azimuth angle, elevation angle, energy ratio, etc.), such as Figure 3 As shown in step 301.

[0131] Then, the direction values ​​(elevation, azimuth) can be compressed or encoded (e.g., by applying spherical quantization or any suitable compression), such as... Figure 3 As shown in step 303.

[0132] The energy ratio is compressed or encoded (e.g., by generating a weighted average per subband and then quantizing it into a 3-bit value), such as... Figure 3 As shown in step 305.

[0133] Then, the coded directionality value and energy ratio (and in some embodiments, other parameters such as coherence value) are combined to generate coded metadata, such as... Figure 3 As shown in step 307. In some embodiments, the encoded directivity values ​​(and power ratios) are directly multiplexed within the encoded transmitted audio signal data stream.

[0134] Directional encoder 205 about Figure 4 Further details are shown below.

[0135] In some embodiments, the directional encoder may include a quantization determiner / bit allocator 401. The quantization determiner / bit allocator 401 may be configured to receive the energy ratio 208 of the encoding / quantization for each sub-band. Furthermore, the quantization determiner / bit allocator 401 may be configured to receive allocated bits for a directional encoding value 400, which defines how many bits have been allocated for encoding the directional parameters of the time-frequency interval. For example, in the case where the audio metadata consists of azimuth, elevation, and energy ratio data for each sub-band, the directional data can be represented by 16 bits, such that the azimuth is approximately represented by 9 bits and the elevation by approximately 7 bits. The energy ratio can be represented by 8 bits. For each frame, there are N sub-bands and M = 4 time-frequency (TF) blocks, making it necessary to... Bits are used to store uncompressed metadata for each frame. The number of subbands can be any number between 1 and 24, depending on the codec's functional mode. For lower bit rates, the number of subbands is fixed to a lower value, such as N=5; otherwise, it can vary from frame to frame and depends on the number of similar time-frequency patches.

[0136] In some embodiments, the energy ratio can be encoded using 3 bits to encode each energy ratio value. Furthermore, instead of transmitting all energy ratio values ​​for all TF blocks, only a weighted average is transmitted for each subband. This average is calculated by taking into account the total energy of each time block, thus favoring the values ​​of subbands with more energy.

[0137] Furthermore, in some embodiments, the quantization determiner / bit allocator 401 may be configured to acquire an input format indicator 402. The input format indicator 402 may be acquired based on any suitable method. For example, in some embodiments, the input format indicator is determined by the device based on analysis of the input audio signal. In some other embodiments, the input format indicator is acquired by receiving a suitable indicator associated with the input audio signal (e.g., as metadata associated with the input audio signal).

[0138] The quantization determiner / bit allocator 401 can then determine encoding control information based on these values, such as the number of bits allocated for encoding each subband and the quantization resolution of the azimuth and elevation angles of all TF tiles (time blocks) in the current subband, and further control the quantization / encoding operation. The quantization resolution can be set, for example, by allowing a predetermined number of bits given by the energy ratio value and the allocated bits.

[0139] This control allows for the uniform quantization of a companded version of the azimuth value, and then companded back, to be performed on an indicator that has been acquired as multi-channel format data at a low bit rate (which may be 2-5 bits after allocating the number of bits per time frequency). Instead of directly using an azimuth uniform quantizer corresponding to the number of available bits, the indicator is used to obtain an indicator that has been acquired as multi-channel format data.

[0140] In some embodiments, the quantization determiner / bit allocator 401 is configured to determine whether the number of bits used to encode a subband above a TF patch is less than a determined threshold. For example, if fewer than 11 bits are used to encode the elevation elements of a subband above four TF patches, the quantization determiner / bit allocator 401 may be configured to control the encoding / quantization operation to check whether average elevation / flexible azimuth quantization is superior to elevation / azimuth companding quantization.

[0141] In some embodiments, the orientation encoder 205 includes a distance determiner 421. The distance determiner 421 may be controlled by a quantization determiner / bit allocator 401 such that distances d1 and d2 are determined when the maximum number of bits allocated for each TF patch of the current subband is less than or equal to a threshold. The angular distance is calculated as follows:

[0142]

[0143] in The average elevation angle is d1. The distance d1 is an estimate of the quantization distortion when using compander coding, and the distance d2 is an estimate of the quantization distortion when using average elevation / flexible azimuth coding.

[0144] In some embodiments, the estimation is based on the unquantized angles and actual values ​​in each codebook, without calculating quantized values.

[0145] In some embodiments, the variance of the elevation angle is considered because if its variance is greater than a certain value, more than one elevation angle value of the subband is encoded. This is further detailed in PCT / FI2019 / 050675.

[0146] Distance determiner 421 is configured to determine whether distance d2 is less than distance d1.

[0147] In some embodiments, the direction encoder 205 includes an average elevation / flexible azimuth encoder 420. When the maximum number of bits allocated to each TF block in the current sub-band is greater than the allowed number of bits, the average elevation / flexible azimuth encoder 420 can be controlled by a quantization determiner / bit allocator 401 to encode the elevation and azimuth values ​​of each TF block within the number of bits allocated to the current sub-band.

[0148] Furthermore, when the distance d2 is less than the distance d1, the average elevation / flexible azimuth encoder 420 can be controlled by the distance determiner 421 to encode the elevation and azimuth values ​​of each TF patch within the number of bits allocated to the current subband. In other words, encoding is performed when the estimated quantization distortion using encoding is less than the estimated quantization distortion using the companding method.

[0149] The average elevation encoder / flexible azimuth encoder 420 includes an average elevation quantizer 413. The average elevation quantizer 413 is configured to determine the average elevation value of a sub-band above a TF block, and then encode (and use) that average elevation value. The average elevation value is then quantized based on the determined quantization grid / configuration. For example, in some embodiments, the average elevation quantizer 413 is configured to encode the average elevation value with 1 or 2 bits (1 bit for a value of 0 degrees, 2 bits for + / - 36 degrees).

[0150] Furthermore, the average elevation encoder / flexible boundary encoder 420 includes a flexible azimuth quantizer configured to employ a flexible boundary embedded codebook for the azimuth value of each tile in the considered TF tiles. In some embodiments, the azimuth coding boundaries are (in degrees) 0, +30, -30, +110, -110, +135, -135. All azimuth values ​​are quantized within these values. Multi-bit estimation (which may include entropy coding, such as Golomb-Rice coding reduction) is then performed, and in cases where too many bits are used, the boundary direction is gradually shifted forward with less impairment (less distortion).

[0151] In some embodiments, the orientation encoder 205 includes a companding encoder 410. The companding encoder 410 includes an elevation quantizer 403 configured to quantize each elevation element in a TF block of a subband.

[0152] Therefore, in some embodiments, the elevation quantizer 415 is configured to determine the quantized elevation element value (or quantization information) based on the elevation element value and the quantization grid or other quantization configuration.

[0153] The quantified elevation angle information can be passed to the compressor 405, and in some embodiments to the inverse compressor 409.

[0154] The direction encoder 205 may also include a comparator 405. The comparator 405 may be configured to receive azimuth elements of the directional parameter 108, and is also configured to receive quantized elevation values ​​and control from the quantizer / bit allocator 401.

[0155] Compander 405 can then be configured to select a companding function based on the quantized elevation angle value. In some embodiments, the companding function can also be determined based on the input channel format, which can be provided as a control or indicator from the quantization determiner / bit allocator 401. Thus, for example, there may be one or more companding functions associated with a determined 5.1-channel input format, and one or more companding functions associated with a determined 7.1-channel input format.

[0156] Then, the companding function can be applied to the azimuth element of the directional parameter to generate a companded azimuth element that can be passed to the azimuth quantizer 407.

[0157] about Figure 6 and Figure 7 An example companding function is shown. Regarding Figure 6The diagram illustrates a first comparation function, which can be selected, for example, when the elevation angle is zero. The input azimuth (X-axis) value 601 can be mapped to a comparation azimuth (Y-axis) value 603 using function 605. Furthermore, Figure 6 A series of raw codewords (quantized values ​​shown as circles) 607 and companded codewords 609 (quantized values ​​shown as asterisks) are shown. The resulting codewords improve resolution on both the front and side sides, where the direct signal is more likely to originate from a multi-channel setup. Figure 6 In this example, there are 5 corresponding values, and they correspond to a 3-bit quantizer, because the other three codewords are used for the negative azimuth value. Although the example shown in this article is a 3-bit codeword example, the same companding function can be used for 4, 5, or more bits.

[0158] When the quantization elevation angle is greater than a given threshold, the forward direction is less likely to exist, and the compassion function is based on... Figure 7 The function within it changes. Regarding Figure 7 The diagram illustrates a second comparation function, which can be selected, for example, when the elevation angle is not zero. Function 705 can be used to map the input azimuth (X-axis) 701 value to the comparation azimuth (Y-axis) 703 value. Furthermore, Figure 7 A series of raw codewords (quantization values ​​shown as circles) 707 and compressed codewords 709 (quantization values ​​shown as asterisks) are shown. According to Figure 7 The companding function defined in the code has very few points quantized to zero or + / -180. The percentage of these points can be adjusted by the pre / post activation values ​​of the companding function, i.e., the first and last y-values ​​in the companding function definition (20 and 160, respectively).

[0159] The output of the compander 405 is then passed to the azimuth quantizer 407, where quantization (such as...) is applied. Figure 6 and Figure 7 (The code is shown in the image).

[0160] The companding encoder 410 may further include an azimuth quantizer 407 configured to receive the output of the compander 405 and quantize the azimuth values. These values ​​are then passed to the inverse compander 409. In some embodiments, the inverse compander 409 is implemented within the decoder 133, so that these values ​​are output from the compander encoder 410 as quantized azimuth elements.

[0161] In some cases, the compressor encoder 410 may also include an inverse compressor 409. The inverse compressor 409 may be configured to receive quantized compressive azimuth elements of the direction parameters, and is also configured to receive quantized compressive elevation values ​​and control from the quantization determiner / bit allocator 401.

[0162] The inverse compander 409 can then be configured to select an inverse companding function based on the quantized elevation angle value. In some embodiments, the inverse companding function can also be determined based on the input channel format, which can be provided as a control or indicator from the quantization determiner / bit allocator 401. The inverse companding function can then be applied to the quantized companded azimuth elements of the directional parameter to generate quantized azimuth elements.

[0163] The inverse companding function is the inverse of the companding function applied in compander 405. In some embodiments, the compander, quantizer, and inverse compander are the same functional element.

[0164] In some embodiments, when the quantization determiner / bit allocator 401 determines that there is more than a threshold bit allocation (e.g., 11 bits), it employs a companding encoder to encode the subband of the TF patch.

[0165] about Figure 5 This shows, as Figure 4 The diagram shows the operation flowchart of the direction encoder 205.

[0166] The initial operation involves acquiring directional metadata (such as azimuth and elevation values), encoded energy ratios, and bit allocations, such as... Figure 5 As shown in step 501.

[0167] Then, the quantization resolution is initially determined based on the energy ratio, such as... Figure 5 Step 503 is shown.

[0168] Encoding checks (where the number of available bits is checked according to a threshold), such as Figure 5 Step 505 is shown in the diagram.

[0169] When the number of available bits is greater than the threshold, perform the companding azimuth quantization operation shown in steps 512, 514, 516 and 518 as described below.

[0170] If the number of available bits is less than a threshold, a distance (error or similarity) check is performed to determine the loss between the quantization based on the compressive azimuth quantization operation shown in steps 512, 514, 516 and 518 as described below and the average elevation / flexible azimuth quantization operation as shown in steps 511 and 513 when compared with the direction parameter.

[0171] If as Figure 5 If the error distance of the elevation / compression azimuth quantization operation shown in step 509 is large, the direction parameter / value can be encoded based on the average elevation / flexible azimuth quantization operation, as shown in steps 511 and 513.

[0172] Therefore, the average elevation angle is quantized based on the quantization determined according to the quantization energy ratio, such as... Figure 5 Step 511 is shown.

[0173] Then, based on the flexible encoding operation described above, the azimuth elements are quantized, such as... Figure 5 Step 513 is shown.

[0174] If as Figure 5 If the error distance of the elevation / compression azimuth quantization operation shown in step 509 is small, then the direction parameter / value can be encoded based on the compression azimuth quantization operation, such as... Figure 5 Steps 512, 514, 516 and 518 are shown.

[0175] Quantize the elevation angle parameter, such as Figure 5 Step 512 is shown.

[0176] In some embodiments, the comprador function is determined based on the quantized elevation angle (and input function) and is applied to the azimuth value, such as... Figure 5 Step 514 is shown.

[0177] Then, the azimuth value of the compressive diffraction is quantized (based on a quantization grid determined according to the quantization energy ratio), such as... Figure 5 Step 516 is shown.

[0178] Then, in some embodiments, the quantized azimuth values ​​of the companded angle can be inversely companded, such as... Figure 5 Step 518 is shown. This operation, as described above, can be implemented within the decoder and is therefore optional relative to the encoder. For example, with respect to orientation encoding, the inverse companding operation can be optional (because the inverse companding of the orientation value can be implemented within the decoder).

[0179] However, in some embodiments where the quantized value of the azimuth (or direction value) is used to encode other parameters (e.g., the encoding of coherent values), an inverse compassion operation can be performed so that the inverse compassion value can be used to encode other parameters.

[0180] In other words, in some embodiments, the inverse companding operation can be implemented to assist in the encoding of other parameters, but not applied to the direction (or specifically, the azimuth value of the companding), because the inverse companding operation can be applied at the decoder.

[0181] Then, output the "quantized" azimuth and elevation values, such as Figure 5 Step 519 is shown.

[0182] Then, the encoded direction value can be output, such as Figure 5 Step 521 is shown.

[0183] Therefore, for systems with low bit allocation per subband or per group of TF blocks for the corresponding directional parameters, the elevation angle values ​​are checked. If they are not similar enough, the directional information in the considered subband is quantized individually for each TF block. Furthermore, if the input format is determined to be multi-channel, the following steps can be performed:

[0184] 1. The quantized elevation angle is limited to positive values ​​(including zero).

[0185] 2. If the quantized elevation angle is zero, then

[0186] a. Compare the azimuth using the comparea function F1 (e.g., ... Figure 6 (As shown)

[0187] b. Quantize the companded value uniformly using the available bits of the azimuth angle.

[0188] c. (Optional) Perform reverse compressive expansion on the quantized azimuth angle.

[0189] d. Identify the spherical index value associated with the elevation angle / quantified azimuth angle.

[0190] otherwise

[0191] a. Compare the azimuth using the comparea function F2 (e.g., ... Figure 7 (As shown)

[0192] b. Use the available bits of the azimuth angle to uniformly quantize the companded value.

[0193] c. (Optional) Perform reverse compressive expansion on the quantized azimuth angle.

[0194] d. Identify the spherical index value associated with the elevation angle / quantified azimuth angle.

[0195] 3. End

[0196] It can be mentioned that verification step 2 is applied to cases where the channel input format is different from the 5.1 or 7.1 channel input format, or more generally, it is not an input format as a single plane (because these input formats always return zero elevation angle).

[0197] Furthermore, in some embodiments, there may be functional differences between input formats such as 5.1-based and 7.1-based because these formats have different preferred azimuth values.

[0198] The companding function can be conveniently described using linear segments, resulting in low complexity and reduced ROM imprinting, because the same function can be used for both companding and decomanding (inverse companding). For example, companding / inverse companding operations can be implemented as shown in the following C code example:

[0199]

[0200]

[0201]

[0202]

[0203]

[0204]

[0205]

[0206] The proposed embodiment improves directional quantization resolution at low bit rates, and is particularly audible for sounds coming from "front". In this embodiment, it eliminates the need to store a non-uniform codebook for varying numbers of bits, requiring only 10 companding function values.

[0207] about Figure 8 Decoder 133 is shown in further detail.

[0208] In some embodiments, decoder 133 includes demultiplexer 801, which is configured to receive encoded audio signals (encoded transmission signals), encoded power ratios and encoded directional parameters (such as encoded azimuth and encoded elevation values), and demultiplex the data stream into separate encoded audio signals, encoded power ratios and encoded directional parameters.

[0209] In some embodiments, the decoder further includes an audio signal decoder 135 configured to receive encoded audio signals and decode those audio signals to generate a decoded audio signal 810 that can be passed to the synthesis processor 139.

[0210] Furthermore, in some embodiments, the decoder 133 includes an energy ratio decoder 803 configured to receive encoded energy ratios and decode these energy ratios to generate an energy ratio 804 that can be passed to the synthesis processor 139.

[0211] Additionally, decoder 133 includes direction decoder 805. Direction decoder 805 is configured to receive average elevation value and flexibly quantized azimuth value, and regenerate elevation and azimuth values ​​based on a known flexibly quantized method (when the direction value is encoded based on a known average elevation / flexibly azimuth quantization method).

[0212] Furthermore, the direction decoder receives the azimuth index corresponding to the uniform quantizer, obtains the value from the uniform quantizer, and then inversely compandspans it to obtain the actual codeword. Additionally, in some embodiments, the direction decoder 805 may also include an inverse compander 409. The inverse compander 409 may be configured to receive the quantized companded azimuth element of the direction parameter, and may also be configured to receive the quantized companded elevation value.

[0213] The inverse compander 409 can then be configured to select an inverse companding function based on the quantized elevation angle value. In some embodiments, the inverse companding function can also be determined based on a channel format, which can be provided as a control or indicator from the quantization determiner / bit allocator. The inverse companding function can then be applied to the quantized companded azimuth elements of the directional parameters to generate quantized azimuth elements.

[0214] The inverse companding function is the inverse of the companding function used in compander 405.

[0215] In some embodiments, when the elevation angle is encoded separately, the azimuth index is obtained separately, and when it is encoded jointly (e.g., when the quantization grid is a known spherical index), it is obtained jointly, and then the azimuth index is extracted and decoded.

[0216] about Figure 9 This shows, as Figure 8 The flowchart shown is an example of the decoder / synthesis processor operation.

[0217] Therefore, the encoded signal is demultiplexed, such as Figure 9 Step 901 is shown.

[0218] Decoding audio signals, such as Figure 9 Step 902 is shown.

[0219] Decoding of the energy-to-space parameter, such as Figure 9 Step 903 is shown.

[0220] Decoding is performed based on the energy comparison direction of the decoded data, such as... Figure 9 Step 905 is shown (where the inverse companding operation is applied when companding operation is used in the encoder).

[0221] Then, the audio signal can be rendered based on spatial parameters (direction and energy ratio) and the audio signal, such as... Figure 9 Step 907 is shown.

[0222] In some embodiments, companding can also be used when prior information about the direction of the audio source exists. Furthermore, in some embodiments, the companding operation, or the companding function selected to implement the companding operation, may depend on the use case or application.

[0223] about Figure 10 The illustration shows an example electronic device that can be used as an analysis or synthesis device. This device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc.

[0224] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. Processor 1407 may be configured to execute various program codes, such as the methods described herein.

[0225] In some embodiments, device 1400 includes memory 1411. In some embodiments, at least one processor 1407 is coupled to memory 1411. Memory 1411 can be any suitable storage component. In some embodiments, memory 1411 includes a program code portion for storing program code that can be implemented on processor 1407. Furthermore, in some embodiments, memory 1411 may also include a storage data portion for storing data, such as data that has been processed or will be processed according to the embodiments described herein. The implemented program code stored in the program code portion and the data stored in the storage data portion can be retrieved by processor 1407 via memory processor coupling when needed.

[0226] In some embodiments, device 1400 includes a user interface 1405. In some embodiments, user interface 1405 may be coupled to processor 1407. In some embodiments, processor 1407 may control the operation of user interface 1405 and receive input from user interface 1406. In some embodiments, user interface 1405 may enable a user to input commands to device 1400, for example, via a keypad. In some embodiments, user interface 1405 may enable a user to obtain information from device 1400. For example, user interface 1405 may include a display configured to display information from device 1400 to a user. In some embodiments, user interface 1405 may include a touchscreen or touch interface that enables information to be input to device 1400 and further displayed to a user of device 1400. In some embodiments, user interface 1405 may be a user interface for communicating with a location determiner as described herein.

[0227] In some embodiments, device 1400 includes an input / output port 1409. In some embodiments, input / output port 1409 includes a transceiver. In such embodiments, the transceiver may be coupled to processor 1407 and configured to enable communication with other devices or electronic devices, such as via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver component may be configured to communicate with other electronic devices or devices via wire or wired coupling.

[0228] The transceiver can communicate with other devices using any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol (e.g., IEEE 802.X), a suitable short-range radio frequency communication protocol (such as Bluetooth), or an Infrared Data Communication Path (IRDA).

[0229] Transceiver input / output port 1409 can be configured to receive signals, and in some embodiments, processor 1407 executing appropriate code is used to determine the parameters described herein.

[0230] Generally, various embodiments can be implemented using hardware or dedicated circuitry systems, software, logic, or any combination thereof. Some aspects of this disclosure can be implemented in hardware, while others can be implemented in firmware or software, which can be executed by a controller, microprocessor, or other computing device, although this disclosure is not limited thereto. While various aspects of this disclosure may be shown and described as block diagrams, flowcharts, or using some other illustrations, it should be clearly understood that, by way of non-limiting example, the blocks, apparatuses, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0231] The term "circuit system" as used in this application may refer to one or more or all of the following:

[0232] (a) Hardware circuit implementation only (such as implementation only in analog and / or digital circuit systems) and

[0233] (b) A combination of hardware circuitry and software, such as (if applicable):

[0234] (i) A combination of (multiple) analog and / or digital hardware circuits and software / firmware, and

[0235] (ii) Any part of a hardware processor(s) having software (including (multiple) digital signal processors, software, and (multiple) memories), which work together to enable a device (such as a mobile phone or server) to perform various functions, and

[0236] (c) (Multiple) hardware circuits and / or (multiple) processors, such as (multiple) microprocessors or a portion thereof, which require software (such as firmware) to operate, but may be absent when the software is not required to operate.

[0237] This definition of circuit system applies to all uses of the term in this application, including in any claim. As another example, as used in this application, the term circuit system also covers only the implementation of hardware circuitry or processor (or processors) or a portion thereof and its accompanying software and / or firmware.

[0238] For example, if applicable to a particular claim element, the term circuit system also covers baseband integrated circuits or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices or other computing or network devices.

[0239] Embodiments of this disclosure can be implemented by computer software executable by a mobile device's data processor, such as in a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products, including software routines, applets, and / or macros) can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. A computer program product may include one or more computer-executable components that, when the program runs, are configured to execute the embodiments. The one or more computer-executable components may be at least one piece of software code or a portion thereof.

[0240] Furthermore, it should be noted that any block of the logic flow shown in the figure can represent a program step, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. Software can be stored on physical media, such as memory chips or memory blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants, CDs. Physical media are non-transitory media.

[0241] The memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and, by way of non-limiting example, can include one or more of the following: general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), FPGAs, gate-level circuits, and processors based on multi-core processor architectures.

[0242] The embodiments of this disclosure can be practiced in a variety of components, such as integrated circuit modules. In general, integrated circuit design is a highly automated process. Complex and powerful software tools can be used to transform logic-level designs into semiconductor circuit designs ready to be etched and formed on semiconductor substrates.

[0243] The scope of protection sought by the various embodiments of this disclosure is defined by the independent claims. Embodiments and features described in this specification that are not within the scope of the independent claims (if any) are to be interpreted as examples that aid in understanding the various embodiments of this disclosure.

[0244] The foregoing description has provided a complete and informative description of exemplary embodiments of the present disclosure by way of non-limiting example. However, various modifications and adaptations will become apparent to those skilled in the art when read in conjunction with the accompanying drawings and appended claims, given the foregoing description. Nevertheless, all such and similar modifications to the teachings of this disclosure will still fall within the scope of the invention as defined in the appended claims. In fact, other embodiments exist, including combinations of one or more of the previously discussed embodiments with any other embodiments.

Claims

1. An apparatus for encoding, comprising components configured to: Acquire multi-channel audio signals; Obtain directional parameter values ​​associated with at least two time-frequency components of the multi-channel audio signal, wherein the directional parameter values ​​associated with at least two time-frequency components include elevation and azimuth elements associated with at least two time-frequency components; as well as The component configured to compand encode the acquired directional parameter values ​​is further configured to: Quantize the elevation angle elements; The companding function is determined based on the quantized elevation angle elements and / or the multi-channel audio signal format; Companded azimuth elements are generated based on the companding function applied to the azimuth elements; as well as The compressive azimuth elements are quantized.

2. The apparatus according to claim 1, wherein the component is further configured to: decompress the quantized companded azimuth element based on the inverse of the companding function.

3. The apparatus of claim 1, wherein the component configured to determine the companding function based on the quantized elevation element and / or the multi-channel audio signal format is configured to: determine the companding function based on the quantized elevation element and the multi-channel audio signal format.

4. The apparatus according to any one of claims 1 to 3, wherein the component configured to compand encode the acquired direction parameter values ​​is further configured to generate codewords for each quantized elevation element and quantized companded azimuth element.

5. The apparatus of claim 2, wherein the component configured to compand encode the acquired direction parameter values ​​is further configured to: generate codewords for each quantized elevation element and decompressed quantized companded azimuth element.

6. The apparatus of claim 3, wherein the component is further configured to: determine a quantization error for the companding encoding and an average elevation angle encoding, wherein the component configured to determine the average elevation angle encoding is configured to: Quantize the average elevation angle element used for sub-bands within the frame; and The azimuth element is quantized based on a quantization grid with variable boundaries, and the component is configured to select either companded encoding output or average elevation encoding output based on the quantization error.

7. The apparatus according to any one of claims 1 to 3, wherein the component is further configured to: determine a quantization grid based on an allocated number of bits, the allocated number of bits being used to encode each subband within a frame comprising subbands and time blocks based on a value of an energy ratio associated with the acquired direction parameter value, wherein the component configured to quantize the elevation element is configured to: quantize the elevation element based on the quantization grid, and the component configured to quantize the companded azimuth element is configured to: quantize the companded azimuth based on the quantization grid.

8. An apparatus for decoding, comprising a component configured to: Obtain at least one encoded bitstream, the at least one encoded bitstream comprising: The encoded multi-channel audio signal and the direction parameter value of the companded code, wherein the direction parameter value of the companded code is associated with at least two time-frequency components of the encoded multi-channel audio signal, and the coded direction parameter value associated with at least two time-frequency components includes an elevation element of the code associated with at least two time-frequency components and an azimuth element of the companded code; Decode the encoded elevation angle element; The decompression spreading function is determined based on the quantized elevation angle elements and / or multi-channel format; and Based on the decompressor function applied to the azimuth element of the compassion coding, a decompressor azimuth element is generated.

9. The apparatus of claim 8, wherein the component configured to determine the decompression spreading function based on the encoded elevation element and / or the multi-channel audio signal format is further configured to: determine the decompression spreading function based on the encoded elevation element and the multi-channel audio signal format.

10. The apparatus of claim 8, wherein the component configured to decode the encoded elevation elements is further configured to decode the codeword for each quantized elevation element.

11. The apparatus of any one of claims 8 to 10, wherein the component is further configured to: determine a quantization grid based on an allocated number of bits, the allocated number of bits being used to encode each subband within a frame comprising subbands and time blocks based on a value of an energy ratio associated with the acquired direction parameter value, wherein the component configured to decode codewords for each quantized elevation element is configured to: decode the elevation element based on the quantization grid.

12. A method for encoding, comprising: Acquire multi-channel audio signals; Obtain directional parameter values ​​associated with at least two time-frequency components of the multi-channel audio signal, wherein the directional parameter values ​​associated with at least two time-frequency components include elevation and azimuth elements associated with at least two time-frequency components; as well as The acquired directional parameter values ​​are companded and encoded, wherein the companding and encoding of the acquired directional parameter values ​​includes: Quantize the elevation angle elements; The companding function is determined based on the quantized elevation angle elements and / or the multi-channel audio signal format; Companded azimuth elements are generated based on the companding function applied to the azimuth elements; and The compressive azimuth elements are quantized.

13. The method of claim 12, further comprising: Based on the inverse of the companding function, the quantized companded azimuth elements are decomanded.

14. The method of claim 12, wherein determining the companding function based on the quantized elevation angle element and / or the multi-channel audio signal format further comprises: The companding function is determined based on the quantized elevation angle element and the multi-channel audio signal format.

15. The method according to any one of claims 12 to 14, wherein companding the acquired directional parameter values ​​further comprises: Generate codewords for each quantized elevation angle element and quantized compressive azimuth angle element.

16. The method of claim 13, wherein compascaling encoding the acquired directional parameter values ​​further comprises: Generate codewords for each quantized elevation angle element and the decompressed quantized compressed azimuth angle element.

17. The method of claim 14, further comprising: Determine the quantization error and average elevation angle coding used for the companding coding, wherein the average elevation angle coding includes: Quantize the average elevation angle element used for sub-bands within the frame; and The azimuth element is quantized based on a quantization raster with variable boundaries, and the method further includes: selecting companded encoding output or average elevation angle encoding output based on the quantization error.

18. The method according to any one of claims 12 to 14, further comprising: A quantization grid is determined based on an allocated number of bits, which are used to encode each subband within a frame, including subbands and time blocks, based on an energy ratio value associated with the acquired direction parameter value. Quantization of the elevation element includes quantizing the elevation element based on the quantization grid, and quantization of the companded azimuth element includes quantizing the companded azimuth based on the quantization grid.

19. A method for decoding, comprising: Acquire at least one coded bitstream, the at least one coded bitstream comprising: coded multi-channel audio signal and companded coded direction parameter values, wherein the companded coded direction parameter values ​​are associated with at least two time-frequency portions of the coded multi-channel audio signal, and the coded direction parameter values ​​associated with at least two time-frequency portions include coded elevation elements and companded azimuth elements associated with at least two time-frequency portions; Decode the encoded elevation angle element; The decompression spreading function is determined based on the encoded elevation angle elements and / or the encoded multi-channel audio signal format; and Based on the decompressor function applied to the azimuth element of the compassion coding, a decompressor azimuth element is generated.

20. The method of claim 19, wherein determining the decompression spreading function based on the encoded elevation angle element and / or the encoded multi-channel audio signal format comprises: The decompression spreading function is determined based on the encoded elevation angle element and the encoded multi-channel audio signal format.

21. The method of claim 19, wherein decoding the encoded elevation angle element further comprises: Decode the codewords used for each elevation angle element of the quantization.

22. The method according to any one of claims 19 to 21, further comprising: A quantization grid is determined based on an allocated number of bits, which are used to encode each subband within a frame, including subbands and time blocks, based on an energy ratio value associated with the acquired direction parameter value. Decoding the codeword for each quantized elevation element includes decoding the elevation element based on the quantization grid.