Apparatus and method for encoding or decoding directional audio coding parameters using different time / frequency resolutions

By encoding DirAC parameters with varying resolutions and applying advanced encoding techniques, the method reduces bitrates for spatial audio coding, addressing the inefficiencies of existing DirAC techniques and enabling efficient immersive audio transmission.

JP7799581B2Active Publication Date: 2026-01-15FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022133236
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-11-17
Filing Date
2022-08-24
Publication Date
2026-01-15
Estimated Expiration
2038-11-16

AI Technical Summary

Technical Problem

Existing directional audio coding (DirAC) techniques require high bitrates for transmitting spatial audio data, which is inefficient for immersive audio content transmission, especially when considering multiple sound sources and background noise.

Method used

The proposed solution involves encoding directional audio coding parameters with different resolutions for diffusion and direction parameters, applying grouping and averaging techniques, and using quantization and entropy coding to reduce bitrates while maintaining audio quality.

Benefits of technology

This approach achieves improved spatial audio coding with reduced bitrates by optimizing the resolution and encoding process, ensuring efficient transmission of immersive audio content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007799581000027
    Figure 0007799581000027
  • Figure 0007799581000028
    Figure 0007799581000028
  • Figure 0007799581000029
    Figure 0007799581000029
Patent Text Reader

Abstract

An apparatus is provided for encoding directional audio coding parameters, including diffusion parameters and direction parameters. [Solution] The apparatus comprises a parameter calculator (100) for calculating diffusion parameters using a first time or frequency resolution and directional parameters using a second time or frequency resolution, and a quantizer and encoder processor (200) for generating quantized and coded representations of the diffusion parameters and directional parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention is directed to audio signal processing, and in particular to an efficient coding scheme for directional audio coding parameters such as DirAC metadata.

[0002] The present invention aims to propose a low bitrate coding solution for coding spatial metadata from 3D audio scene analysis performed by Directional Audio Coding (DirAC), a perceptually motivated technique for spatial audio processing. [Background technology]

[0003] Transmitting an audio scene in three dimensions usually requires processing multiple channels, which generates a large amount of data to transmit. The directional audio coding (DirAC) technique [Non-Patent Document 1] is an efficient method for analyzing an audio scene and parametrically representing it. DirAC uses a perceptually motivated representation of the sound field based on the direction of arrival (DOA) and diffuseness measured for each frequency band. It is built on the assumption that, at a given time and for one critical band, the spatial resolution of the auditory system is limited to decoding one cue related to direction and another cue related to inter-aural coherence. The spatial sound is then reproduced in the frequency domain by crossfading two streams: an omnidirectional diffuse stream and a directional non-diffuse stream.

[0004] The present invention discloses a 3D audio coding method based on DirAC sound representation and reproduction to achieve immersive audio content transmission at low bit rates.

[0005] DirAC is a perceptually motivated spatial sound reproduction: at a given time, for one critical band, the spatial resolution of the auditory system is assumed to be limited to decoding one cue related to direction and another cue related to interaural coherence.

[0006] Based on these assumptions, DirAC represents spatial sound in a frequency band by crossfading two streams: an omnidirectional diffuse stream and a directional non-diffuse stream. DirAC processing is performed in two stages: analysis and synthesis, as depicted in Figures 10a and 10b.

[0007] In the DirAC analysis stage, the first-order coincident microphone in B format is taken as input and the diffuseness and direction of arrival of the sound are analyzed in the frequency domain.

[0008] In the DirAC synthesis stage, the sound is split into two streams: a non-diffuse stream and a diffuse stream. The non-diffuse stream is reproduced as a point source using amplitude panning, which can be performed using vector-based amplitude panning (VBAP) [Non-Patent Document 2]. The diffuse stream, which is responsible for the sensation of envelopment, is reproduced by transmitting mutually uncorrelated signals to the speakers.

[0009] DirAC parameters, hereafter also called spatial metadata or DirAC metadata, consist of a tuple of diffusivity and direction. Direction can be represented by two angles, azimuth and elevation, in spherical coordinates, while diffusivity is a scalar factor between 0 and 1.

[0010] 10a shows a filter bank 130 receiving a B-format input signal. An energy analysis 132 and an intensity analysis 134 are performed. Temporal averaging is performed on the energy results shown at 136 and on the intensity results shown at 138, and from the averaged data, spread values ​​for individual time / frequency bins are calculated as shown at 110. Direction values ​​for the time / frequency bins given by the time or frequency resolution of the filter bank 130 are calculated by block 120.

[0011] In the DirAC synthesis shown in Fig. 10b, an analysis filter bank 431 is again used. A virtual microphone processing block 421 is applied, where the virtual microphones correspond for example to the speaker positions of a 5.1 speaker setup. The diffuseness metadata is processed by a corresponding processing block 422 for diffuseness and a VBAP (Vector Based Amplitude Panning) gain table shown in block 423. A speaker averaging block 424 is configured to perform gain averaging, and a corresponding normalization block 425 is applied in order to have a corresponding defined loudness level in the individual final speaker signals. In block 426, microphone compensation is performed.

[0012] The resulting signal is used to generate a diffused stream 427, which includes a decorrelation step, and a non-diffused stream 428. Both streams are added in an adder 429 for the corresponding subband, and then summed with other subbands in block 431, i.e., a frequency-to-time conversion is performed. Therefore, block 431 can also be considered as a synthesis filter bank. Similar processing operations are performed for other channels from a particular speaker configuration, where for different channels the virtual microphone settings in block 421 will be different.

[0013] In the DirAC analysis stage, the first-order coincident microphones in B-format are considered as input and the diffuseness and direction of arrival of the sound are analyzed in the frequency domain.

[0014] In the DirAC synthesis stage, the sound is split into two streams: a non-diffuse stream and a diffuse stream. The non-diffuse stream is reproduced as a point source using amplitude panning, which can be performed using vector-based amplitude panning (VBAP) [Non-Patent Document 2]. The diffuse stream, which is responsible for the sensation of envelopment, is reproduced by transmitting mutually uncorrelated signals to the speakers.

[0015] DirAC parameters, hereafter also called spatial metadata or DirAC metadata, consist of a tuple of diffusivity and direction. Direction can be represented by two angles, azimuth and elevation, in spherical coordinates, while diffusivity is a scalar factor between 0 and 1.

[0016] If the STFT is considered as a time-frequency transform with a time resolution of 20 ms, as is usually recommended in some papers, and with 50% overlap between adjacent analysis windows, the DirAC analysis produces 288,000 values ​​per second for an input sampled at 48 kHz, which corresponds to an overall bitrate of approximately 2.3 Mbit / s when angles are quantized at 8 bits. The amount of data is not suitable for achieving low-bitrate spatial audio coding, and therefore an efficient coding scheme for DirAC metadata is needed.

[0017] Previous work on metadata reduction has focused primarily on videoconferencing scenarios, and DirAC's capabilities have been significantly reduced to allow for a minimal data rate for its parameters [Non-Patent Document 4]. Indeed, it has been proposed to limit directional analysis to the azimuth angle in the horizontal plane in order to reproduce only 2D audio scenes. Furthermore, diffuseness and azimuth angle are transmitted only up to 7 kHz, limiting communication to wideband audio. Finally, diffuseness is coarsely quantized at one or two bits, and sometimes only the diffuse stream is turned on or off in the synthesis stage, which is not comprehensive enough when considering multiple sound sources and two or more voices over background noise. In [Non-Patent Document 4], azimuth angle is quantized at three bits, and the source, in this case the loudspeaker, is assumed to have a very static position. Therefore, the parameters are transmitted only with an update frequency of 50 ms. Based on these many strong assumptions, the bit demand can be reduced to approximately 3 kbit / s. [Prior art documents] [Non-patent literature]

[0018] [Non-Patent Document 1] V. Pulkki, MV. Laitinen, J. Vilkamo, J. Ahonen, T. Lokki, and T. Pihlajamaeki,"Directional audio coding - perception-based reproduction of spatial sound", International Workshop on the Principles and Application on Spatial Hearing,Nov. 2009,Zao; Miyagi, Japan [Non-patent document 2] V. Pulkki, "Virtual source positioning using vector base amplitude panning", J. Audio Eng. Soc., 45(6):456-466, June 1997 [Non-patent document 3] J. Ahonen and V. Pulkki, "Diffuseness estimation using temporal variation of intensity vectors", in Workshop on Applications of Signal Processing to Audio and Acoustics WASPAA, Mohonk Mountain House, New Paltz, 2009. [Non-patent document 4] T. Hirvonen, J. Ahonen, and V. Pulkki, "Perceptual compression methods for metadata in Directional Audio Coding applied to audiovisual teleconference", AES 126th Convention, 2009, May 7-10, Munich, Germany. Summary of the Invention [Problem to be solved by the invention]

[0019] The object of the present invention is to provide an improved spatial audio coding concept. [Means for solving the problem]

[0020] This object is achieved by an apparatus for encoding directional audio coding parameters according to claim 1, a method for encoding directional audio coding parameters according to claim 17, a decoder for decoding an encoded audio signal according to claim 18, a method for decoding according to claim 33 or a computer program according to claim 34.

[0021] According to one aspect, the invention is based on the finding that if diffusion parameters on the one hand and directional parameters on the other hand are provided with different resolutions and the different parameters with different resolutions are quantized and coded to obtain coded directional audio coding parameters, an improved quality on the one hand and at the same time a reduced bit rate for coding the spatial audio coding parameters on the other hand is obtained.

[0022] In one embodiment, the time or frequency resolution of the diffuseness parameters is lower than the time or frequency resolution of the directional parameters. In a further embodiment, grouping is performed not only over frequency but also over time. The original diffuseness / directional audio coding parameters are calculated at a high resolution, e.g., for high-resolution time / frequency bins, and grouping, preferably grouping with averaging, is performed to calculate the resulting diffuseness parameters at a low time or low frequency resolution and the resulting directional parameters at a medium time or frequency resolution, i.e., a time or frequency resolution that lies between the time or frequency resolution of the diffuseness parameters and the high resolution at which the original raw parameters were calculated.

[0023] In some embodiments, the first and second time resolutions are different and the same as the first and second frequency resolutions, or vice versa, i.e., the first and second frequency resolutions are different but the first and second time-frequency are the same. In further embodiments, both the first and second time resolutions are different, and the first and second frequency resolutions are similarly different. Thus, the first time or frequency resolution may also be considered a first time-frequency resolution, and the second time or frequency resolution may also be considered a second time-frequency resolution.

[0024] In a further embodiment, the grouping of the diffusion parameters is performed by weighted summation, in which the weighting coefficients of the weighted summation are determined based on the power of the audio signal, such that time / frequency bins with higher power, or in general with a higher amplitude-related measure of the audio signal, have a greater influence on the result than diffusion parameters of time / frequency bins in which the signal to be analyzed has lower power or a lower energy-related measure.

[0025] In addition, it is preferable to perform a double-weighted averaging for the calculation of the grouped direction parameters. This double-weighted averaging is performed so that if the power of the original signal is very high in a time / frequency bin, the direction parameters from this time / frequency bin will have a greater influence on the final result. At the same time, the spread value of the corresponding bin is also taken into account so that, when the power is the same in both time / frequency bins, the direction parameters from the time / frequency bin associated with high spread will have a smaller influence on the final result compared to the direction parameters with low spread.

[0026] The parameter processing is preferably performed in frames, each frame being organized into a certain number of bands, each band comprising at least two original frequency bins for which the parameters were calculated. The bandwidth of the band, i.e., the number of original frequency bins, increases with the number of bands, with higher frequency bands being wider than lower frequency bands. In a preferred embodiment, the number of diffusion parameters per band and frame is equal to 1, and the number of directional parameters per frame and band is found to be greater than 2, e.g., 2 or 4. It has been found useful to have the same frequency resolution but different time resolution for the diffusion parameters and directional parameters, i.e., the number of bands for the diffusion parameters and directional parameters in a frame is equal to each other. These grouped parameters are then quantized and coded by a quantizer and encoder processor.

[0027] According to a second aspect of the present invention, the object of providing an improved processing concept for spatial audio coding parameters is achieved by a parameter quantizer for quantizing diffusion parameters and directional parameters, a subsequently connected parameter encoder for encoding the quantized diffusion parameters and the quantized directional parameters, and a corresponding output interface for generating an encoded parameter representation containing information about the encoded diffusion parameters and the encoded directional parameters. Thus, by quantization and subsequent entropy coding, a significant data rate reduction is obtained.

[0028] The diffusion parameters and directional parameters input to the encoder may be high-resolution diffusion / directional parameters, or may be grouped or ungrouped low-resolution directional audio coding parameters. One feature of a preferred parameter quantizer is that the quantization precision for quantizing the directional parameters is derived from the diffusion values ​​of the diffusion parameters associated with the same time / frequency region. Thus, in one feature of the second aspect, directional parameters associated with diffusion parameters having high diffusion properties are quantized with lower precision than directional parameters associated with time / frequency regions with diffusion parameters exhibiting low diffusion properties.

[0029] The diffusion parameters themselves may be entropy coded in raw coding mode, or coded in single-value coding mode if the diffusion parameters of the bands of a frame have the same value throughout the frame. In other embodiments, the diffusion values ​​may be coded in steps of only two consecutive values.

[0030] Another feature of the second embodiment is that the direction parameters are converted to an azimuth / elevation representation. In this feature, the elevation value is used to determine an alphabet for quantizing and encoding the azimuth value. Preferably, the azimuth alphabet has the largest amount of distinct values ​​when the elevation angle indicates zero angle, or generally, the equatorial angle, on the unit sphere. The smallest amount of values ​​in the azimuth alphabet is when the elevation angle indicates the north or south pole of the unit sphere. Thus, the alphabet value decreases with increasing absolute value of the elevation angle counted from the equator.

[0031] The elevation angle values ​​are quantized with a quantization precision determined from the corresponding diffusion values, with the quantization alphabet on the one hand and the quantization precision on the other hand determining the quantization and typically entropy coding of the corresponding azimuth angle values.

[0032] Thus, an efficient, parameter-adaptive process is performed that removes as many irrelevances as possible, while at the same time applying high resolution or high accuracy where it is worthwhile, but in other areas, such as the north or south poles of the unit sphere, the accuracy is not as high as at the equator of the unit sphere.

[0033] The decoder side operating according to the first aspect performs any kind of decoding and performs corresponding de-grouping using the coded or decoded diffusion parameters and the coded or decoded directional parameters. Thus, a parameter resolution conversion is performed to increase the resolution from the coded or decoded directional audio coding parameters to the resolution finally used by the audio renderer to perform rendering of the audio scene. In the course of this resolution conversion, different resolution conversions are performed on the one hand for the diffusion parameters and on the other hand for the directional parameters.

[0034] Diffusion parameters are typically coded at low resolution, so to obtain a high-resolution representation, one diffusion parameter needs to be doubled or copied several times, whereas the resolution of the directional parameters is already higher than that of the diffusion parameters in the coded audio signal, so the corresponding directional parameters need to be copied or doubled less compared to the diffusion parameters.

[0035] In one embodiment, the copied or doubled directional audio coding parameters are applied as is, or are processed, such as smoothed or low-pass filtered, to avoid artifacts caused by parameters that vary strongly over frequency and / or time. However, in a preferred embodiment, since the application of the resolution-converted parametric data is performed in the spectral domain, the corresponding frequency-to-time transformation of the rendered audio signal from the frequency domain to the time domain preferably performs inherent averaging by applying an overlap and add procedure, a function typically included within a synthesis filterbank.

[0036] At the decoder side according to the second aspect, certain procedures performed at the encoder side regarding entropy coding on the one hand and quantization on the other hand are undone: it is preferable to determine the dequantization precision at the decoder side from the typically quantized or dequantized diffusion parameters associated with the corresponding directional parameters.

[0037] Preferably, the alphabet for the elevation parameter is determined from the corresponding diffusion value and its associated dequantization precision. For the second aspect, it is also preferred to perform the determination of the dequantization alphabet for the azimuth parameter based on the quantized, or preferably dequantized, value of the elevation parameter.

[0038] According to a second aspect, a raw coding mode or an entropy coding mode is implemented at the encoder side, and the mode resulting in fewer bits is selected in the encoder and signaled to the decoder via some side information. Typically, the raw coding mode is always implemented for directional parameters associated with high spreading values, and the entropy coding mode is attempted for directional parameters associated with lower spreading values. In the entropy coding mode using raw coding, the azimuth and elevation values ​​are merged into a sphere index, which is then coded using a binary code or a punctured code, and at the decoder side, this entropy coding is undone accordingly.

[0039] In the entropy coding with modeling mode, average elevation and azimuth values ​​are calculated for a frame, and residual values ​​relative to these average values ​​are actually calculated. Thus, a kind of prediction is performed, and the predicted residual values, i.e., the distances for the elevation or azimuth angles, are entropy coded. For this purpose, it is preferable to implement an extended Golomb-Rice procedure, which relies on the preferably signed distances and average values ​​as well as Golomb-Rice parameters determined and coded on the encoder side. On the decoder side, as soon as entropy coding with modeling, i.e., this decoding mode, is signaled and determined by side information evaluation in the decoder, decoding using the extended Golomb-Rice procedure is performed using the coded averages, the coded, preferably signed distances, and the corresponding Golomb-Rice parameters for the elevation and azimuth angles.

[0040] Preferred embodiments of the present invention will now be discussed with reference to the accompanying drawings. [Brief explanation of the drawings]

[0041] [Figure 1a]FIG. 1 is a diagram showing a preferred embodiment of the encoder side of the first or second aspect. [Figure 1b] FIG. 10 is a diagram showing a preferred embodiment of the decoder side of the first or second aspect. [Figure 2a] 1 shows a preferred embodiment of an apparatus for encoding according to a first aspect; [Figure 2b] FIG. 2b shows a preferred implementation of the parameter calculator of FIG. 2a. [Figure 2c] FIG. 10 illustrates a further implementation for the calculation of the diffusion parameter. [Figure 2d] FIG. 2b shows a further preferred implementation of the parameter calculator 100 of FIG. 2a. [Figure 3a] 1a or 430 of FIG. 1b, showing the time / frequency representation obtained by the analysis filter bank 130 of FIG. 1a or 430 of FIG. 1b, with high time or frequency resolution. [Figure 3b] FIG. 1 illustrates an implementation of diffuse grouping at low time or frequency resolution, in particular at a specific low time resolution of a single diffuse parameter per frame. [Figure 3c] FIG. 1 shows a preferred example of a medium resolution for directional parameters with 5 bands on the one hand and 4 time domains on the other hand, resulting in 20 time / frequency domains. [Figure 3d] FIG. 10 shows an output bitstream with coded diffusion parameters and coded directional parameters. [Figure 4a] FIG. 10 illustrates an apparatus for encoding directional audio coding parameters according to a second aspect. [Figure 4b] FIG. 1 illustrates a preferred implementation of a parameter quantizer and a parameter encoder for the calculation of coded diffusion parameters. [Figure 4c] FIG. 4b shows a preferred implementation of the encoder of FIG. 4a with regard to the cooperation of the different elements. [Figure 4d] FIG. 10 illustrates the quasi-uniform coverage of the unit sphere applied for quantization purposes in the preferred embodiment. [Figure 5a] 4b shows an overview of the operation of the parameter encoder of FIG. 4a operating in different coding modes; [Figure 5b] FIG. 5b illustrates pre-processing of direction indices for both modes of FIG. 5a. [Figure 5c] FIG. 2 illustrates a first coding mode in a preferred embodiment. [Figure 5d] FIG. 1 illustrates a preferred embodiment of a second coding mode. [Figure 5e] FIG. 1 illustrates a preferred implementation of entropy coding of signed distances and corresponding averages using the GR coding procedure. [Figure 5f] FIG. 1 illustrates a preferred embodiment for the determination of optimal Golomb-Rice parameters. [Figure 5g] FIG. 5B illustrates an implementation of the extended Golomb-Rice procedure for encoding the reordered signed distances, as shown in block 279 of FIG. 5e. [Figure 6a] FIG. 4b illustrates an implementation of the parameter quantizer of FIG. 4a. [Figure 6b] FIG. 10 illustrates a preferred implementation of functionality related to a parameter inverse quantizer that is also used in certain aspects of the encoder-side implementation. [Figure 6c] FIG. 1 shows an overview of the implementation of the raw orientation encoding procedure. [Figure 6d] FIG. 10 illustrates an implementation of calculation of mean direction for azimuth and elevation angles and quantization and dequantization. [Figure 6e] FIG. 1 shows a projection of mean elevation and azimuth data. [Figure 6f] FIG. 1 illustrates distance calculations for elevation and azimuth angles. [Figure 6g] FIG. 10 shows an overview of the coding of the mean direction in the entropy coding mode with modeling. [Figure 7a] 1 shows a decoder for decoding an encoded audio signal according to a first embodiment; [Figure 7b] FIG. 7b illustrates a preferred implementation of the parameter resolution converter of FIG. 7a and subsequent audio rendering. [Figure 8a] 4 shows a decoder for decoding an encoded audio signal according to a second embodiment; FIG. [Figure 8b] FIG. 1 illustrates a schematic bitstream representation of coded spreading parameters in one embodiment. [Figure 8c] FIG. 1 illustrates an implementation of a bitstream when raw coding mode is selected. [Figure 8d] FIG. 10 shows a schematic bitstream when another coding mode, namely entropy coding mode with modeling, is selected. [Figure 8e] FIG. 1 illustrates a preferred implementation of a parameter decoder and parameter dequantizer, where the dequantization precision is determined based on the diffuseness in the time / frequency domain. [Figure 8f] FIG. 1 illustrates a preferred implementation of a parameter decoder and parameter dequantizer in which the elevation alphabet is determined from the dequantization precision and the azimuth alphabet is determined based on the dequantization precision and the time / frequency domain elevation data. [Figure 8g] FIG. 8b shows an overview of the parameter decoder of FIG. 8a showing two different decoding modes. [Figure 9a] FIG. 10 illustrates the decoding operation when raw encoding mode is active. [Figure 9b] FIG. 10 illustrates decoding of the mean direction when the entropy decoding mode with modeling is active. [Figure 9c] FIG. 10 illustrates elevation and azimuth angle reconstruction and subsequent inverse quantization when the decoding with modeling mode is active. [Figure 10a] FIG. 1 illustrates the well-known DirAC analyzer. [Figure 10b] FIG. 1 shows the well-known DirAC synthesizer. DETAILED DESCRIPTION OF THE INVENTION

[0042] The present invention generalizes DirAC metadata compression to any kind of scenario and is applied in a spatial coding system as shown in Figures 1a and 1b, where a DirAC-based spatial audio encoder and decoder are shown.

[0043] The encoder typically analyzes the spatial audio scene in B format. Alternatively, the DirAC analysis can be tailored to analyze different audio formats, such as audio objects or multi-channel signals, or any combination of spatial audio formats. The DirAC analysis extracts a parametric representation from the input audio scene. The direction of arrival (DOA) and diffuseness measured per time-frequency unit form the parameters. The DirAC analysis is followed by a spatial metadata encoder, which quantizes and encodes the DirAC parameters to obtain a low-bitrate parametric representation. The latter module is the subject of this invention.

[0044] The downmix signal, derived from various sources or audio input signals along with the parameters, is coded for transmission by a conventional audio core coder. In a preferred embodiment, an EVS audio coder is used to code the downmix signal, but the present invention is not limited to this core coder and can be applied to any audio core coder. The downmix signal is composed of various channels, called transport channels, which can be, for example, a B-format signal, a stereo pair, or four coefficient signals constituting a monophonic downmix depending on the target bit rate. The coded spatial parameters and the coded audio bitstream are multiplexed before being transmitted over a communication channel.

[0045] At the decoder, the transport channels are decoded by the core decoder, but the DirAC metadata is first decoded before being passed along with the decoded transport channels to the DirAC synthesis. The DirAC synthesis uses the decoded metadata to control the reproduction of the direct sound stream and its mix with the diffuse sound stream. The reproduced sound field can be reproduced in any loudspeaker layout or generated in Ambisonics format (HOA / FOA) with any order.

[0046] An audio encoder for encoding an audio signal, such as a B-format input signal, is shown in Fig. 1a. The audio encoder comprises a DirAC analyzer 100, which may include an analysis filterbank 130, a diffuseness estimator 110, and a direction estimator 120. The diffuseness and direction data are output to a spatial metadata encoder 200, which finally outputs encoded metadata on line 250. The B-format signal may also be forwarded to a beamformer / signal selector 140, which generates a mono or stereo transport audio signal from the input signal, which is then encoded in an audio encoder 150, i.e., preferably an EVS (Enhanced Voice Service) encoder. The encoded audio signal is output at 170. The encoded coding parameters, shown at 250, are input to a spatial metadata decoder 300. The encoded audio signal 170 is input to an audio decoder 340, which in a preferred embodiment is implemented as an EVS decoder in accordance with an encoder-side embodiment.

[0047] The decoded transport signal together with the decoded directional audio coding parameters is input to a DirAC synthesizer 400. In the embodiment shown in Fig. 1b, the DirAC synthesizer comprises an output synthesizer 420, an analysis filterbank 430 and a synthesis filterbank 440. At the output of the synthesis filterbank 440, a decoded multi-channel signal 450 is obtained, which may be transmitted to a loudspeaker or, alternatively, may be an audio signal in any other format, such as First Order Ambisonics (FOA) or Higher Order Ambisonics (HOA) format. Naturally, any other parametric data, such as MPS (MPEG Surround) data or SAOS (Spatial Audio Object Coding) data, may be generated together with the transport channel, which may be a mono or stereo channel.

[0048] In general, the output synthesizer operates by calculating, for each time-frequency bin determined by the analysis filterbank 430, a direct audio signal on the one hand and a diffuse audio signal on the other hand. The direct audio signal is calculated based on a direction parameter and a relationship between the direct and diffuse audio signals in the final audio signal for this time / frequency bin, this relationship being determined based on the diffuseness parameter such that time / frequency bins with a high diffuseness parameter result in an output signal with a large amount of diffuse signal and a small amount of direct signal, and time / frequency bins with a low diffuseness parameter result in an output signal with a large amount of direct signal and a small amount of diffuse signal.

[0049] 2a shows an apparatus for encoding directional audio coding parameters including diffusion parameters and directional parameters according to a first embodiment. The apparatus comprises a parameter calculator 100 for calculating diffusion parameters using a first time or frequency resolution and directional parameters using a second time or frequency resolution. The apparatus comprises a quantizer and encoder processor 200 for generating quantized and coded representations of the diffusion parameters and directional parameters shown at 250. The parameter calculator 100 may comprise elements 110, 120, 130 of FIG. 1a, where the various parameters have already been calculated at the first or second time or frequency resolution.

[0050] Alternatively, a preferred implementation is shown in Fig. 2b. Here, the parameter calculators, in particular blocks 110 and 120 in Fig. 1a, are configured as shown in item 130 of Fig. 2b, i.e., they calculate parameters using a third or fourth, typically higher, time or frequency resolution. A grouping operation is performed. To calculate the diffusion parameters, grouping and averaging is performed as shown in block 141 to obtain diffusion parameter representations at a first time or frequency resolution, and for the calculation of the directional parameters, grouping (and averaging) is performed in block 142 to obtain directional parameter representations at a second time or frequency resolution.

[0051] The diffusion parameters and directional parameters are calculated such that the second time or frequency resolution is different from the first time or frequency resolution and the first time resolution is lower than the second time resolution, or the second frequency resolution is higher than the first frequency resolution, or again alternatively such that the first time resolution is lower than the second time resolution and the first frequency resolution is equal to the second frequency resolution.

[0052] Typically, the diffusion parameters and directional parameters are calculated for a set of frequency bands, where bands with lower center frequencies are narrower than bands with higher center frequencies. As already discussed with respect to Figure 2b, parameter calculator 100 is configured to obtain initial diffusion parameters having a third time or frequency resolution, and parameter calculator 100 is also configured to obtain initial directional parameters having a fourth time or frequency resolution, where typically the third and fourth time or frequency resolutions are equal to each other.

[0053] The parameter calculator is then configured to group and average the initial spreading parameters so that a third time or frequency resolution is higher than the first time or frequency resolution, i.e., a reduced resolution is performed. The parameter calculator is also configured to group and average the initial direction parameters so that a fourth time or frequency resolution is higher than the second time or frequency resolution, i.e., a reduced resolution is performed. Preferably, the third time or frequency resolution is a constant time resolution so that each initial spreading parameter is associated with a time slot or frequency bin having the same size. The fourth time or frequency resolution is also a constant frequency resolution so that each initial direction parameter is associated with a time slot or frequency bin having the same size.

[0054] The parameter calculator is configured to average over a first plurality of spreading parameters associated with a first plurality of time slots, the parameter calculator 100 is also configured to average over a second plurality of spreading parameters associated with a second plurality of frequency bins, the parameter calculator is also configured to average over a third plurality of directional parameters associated with a third plurality of time slots, or the parameter calculator is also configured to average over a fourth plurality of directional parameters associated with a fourth plurality of frequency bins.

[0055] As discussed with respect to Figures 2c and 2d, the parameter calculator 100 is configured to perform a weighted average calculation, in which diffusion parameters or directional parameters derived from input signal portions with higher amplitude-related measurements are weighted using a higher weighting factor compared to diffusion parameters or directional parameters derived from input signal portions with lower amplitude-related measurements. The parameter calculator 100 is configured to calculate bin-wise and amplitude-related measurements at a third or fourth time or frequency resolution 143, as shown in item 143 of Figure 2c. In block 144, a weighting factor for each bin is calculated, and in block 145, grouping and averaging are performed using a weighted combination, such as weighted summation, where the diffusion parameters for each bin are input to block 145. At the output of block 145, diffusion parameters with a first time or frequency resolution are obtained, which may then be normalized in block 146, although this procedure is merely optional.

[0056] FIG. 2d illustrates the calculation of directional parameters with a second resolution. In block 146, amplitude-related measurements are calculated for each bin at the third and fourth resolutions, similar to item 143 in FIG. 2c. In block 147, weighting factors are calculated for each bin, not only based on the amplitude-related measurements obtained from block 147 but also using the corresponding diffusivity parameters for each bin, as shown in FIG. 2d. Thus, for the same amplitude-related measurement, a higher coefficient is typically calculated for lower diffusivity. In block 148, grouping and averaging are performed using a weighted combination, such as summation, and the results may be normalized, as shown in optional block 146. Thus, at the output of block 146, directional parameters are obtained as unit vectors corresponding to a two- or three-dimensional region, such as Cartesian vectors, which can be easily converted to polar form with azimuth and elevation values.

[0057] FIG. 3a shows a time / frequency raster as obtained by the filter bank analysis 430 of FIGS. 1a and 1b or applied by the filter bank synthesis 440 of FIG. 1b. In one embodiment, the entire frequency range is separated into 60 frequency bands, and a frame has 16 time slots. This high time or frequency resolution is preferably third or fourth-highest. Thus, starting from 60 frequency bands and 16 time slots, 960 time / frequency tiles or bins are obtained per frame.

[0058] 3b shows the resolution reduction performed by the parameter calculator, specifically by block 141 of FIG. 2b, to obtain a first time or frequency resolution representation of the spreading values. In this embodiment, the entire frequency bandwidth is separated into five grouping bands and only a single time slot. Therefore, for one frame, ultimately, only five spreading parameters per frame are obtained, which are then further quantized and coded.

[0059] Figure 3c shows the corresponding procedure performed by block 142 in Figure 2b. The high-resolution directional parameters from Figure 3a, in which one directional parameter is calculated per bin, are grouped and averaged into a medium-resolution representation in Figure 3c, which now has five frequency bands per frame, but in contrast to Figure 3a, now has four time slots. Thus, ultimately, one frame receives 20 directional parameters, i.e., 20 grouped bins per frame for the directional parameters and only five grouped bins per frame for the diffusion parameters in Figure 3b. In a preferred embodiment, the ends of the frequency bands are limited to their upper ends.

[0060] 3b and 3c, it should be noted that the spreading parameter for the first band, i.e., spreading parameter 1, corresponds to or is associated with the four directional parameters for the first band. As will be outlined later, the quantization precision of all directional parameters in the first band is determined by the spreading parameter for the first band, or, illustratively, the quantization precision of the directional parameters for the fifth band, i.e., the corresponding four directional parameters covering the fifth band and the four time slots within the fifth band, is determined by the single spreading parameter for the fifth band.

[0061] Therefore, in this embodiment where only a single diffusion parameter is configured per band, all directional parameters within a band have the same quantization / dequantization precision. As outlined below, the alphabet for quantizing and encoding the azimuth parameters depends on the value of the original / quantized / dequantized elevation parameter. Therefore, while each directional parameter in each band has the same quantization / dequantization parameter, each azimuth parameter in each grouped bin or time / frequency domain in Figure 3c can have a different alphabet for quantization and encoding.

[0062] The resulting bitstream generated by the quantizer and encoder processor 200, shown at 250 in FIG. 2a, is shown in more detail in FIG. 3d. The bitstream may include a resolution indicator 260 indicating the first and second resolutions. However, if the first and second resolutions are fixedly set by the encoder and decoder, this resolution indicator is not necessary. Items 261 and 262 indicate the coded spreading parameters for the corresponding bands. Because FIG. 3d shows only five bands, only five spreading parameters are included in the coded data stream. Items 263 and 264 indicate coded directional parameters. For the first band, there are four coded directional parameters, where the first index of the directional parameters indicates the band and the second index indicates the time slot. The directional parameters for the fifth band and the fourth time slot, i.e., the top right frequency bin in FIG. 3c, are indicated as DIR54.

[0063] Subsequently, further preferred implementations are discussed in detail.

[0064] Time-Frequency Decomposition In DirAC, both analysis and synthesis are performed in the frequency domain. Time-frequency analysis and synthesis can be performed using various block transforms, such as the short-term Fourier transform (STFT), or filter banks, such as complex-modulated quadrature mirror filter banks (QMF). In a preferred embodiment, we aim to share framing between the DirAC processing and the core encoder. Since the core encoder is preferably based on the 3GPP EVS codec, a 20 ms framing is desired. Furthermore, important criteria such as time and frequency resolution and robustness to aliasing are related to the highly active time-frequency processing in DirAC. Since the system is designed for communications, algorithmic delay is another important feature.

[0065] For all these reasons, a complex modulated low-delay filterbank (CLDFB) is the preferred choice. The CLDFB has a time resolution of 1.25 ms, dividing a 20 ms frame into 16 time slots. The frequency resolution is 400 Hz, which means that the input signal is decomposed into (fs / 2) / 400 frequency bands. The filterbank operation is described in general form by the following equation:

[0066]

number

[0067] where X CR and X CI are the real and imaginary subband values, respectively, t is the subband time index, 0≦t≦15, and k is 0≦k≦L C -1 is the subband index. Analysis prototype w c is S HP is an asymmetric low-pass filter with an adaptive length that depends on w c The length of

number

[0068] For example, CLDFB decomposes a signal sampled at 48 kHz into 60 x 16 = 960 time-frequency tiles per frame. The delay after analysis and synthesis can be adjusted by selecting different prototype filters. A delay of 5 ms (analysis and synthesis) was found to be a good compromise between the delivered quality and the incurred delay. For each time-frequency tile, the diffuseness and direction are calculated.

[0069] Estimation of DirAC parameters In each frequency band, the sound's diffuseness and the sound's direction of arrival are estimated. i (n), x i (n), y i (n), z i From the time-frequency analysis of (n), the pressure and velocity vectors are

number

number

number

number

number

[0070] The diffuseness of a sound field is defined as the ratio between sound intensity and energy density and has a value between 0 and 1.

[0071] The direction of arrival (DOA) is

number

[0072] The direction of arrival is determined by energy analysis of the B-format input and can be defined as the opposite direction of the intensity vector. The direction is defined in Cartesian coordinates, but can easily be converted to spherical coordinates defined by a single radius, azimuth, and elevation angle.

[0073] Overall, if the parameter values ​​were converted directly to bits, three values ​​would need to be coded per time-frequency tile: azimuth, elevation, and diffusivity. The metadata would consist of 2880 values ​​per frame, or 144000 values ​​per second, in the CLDFB example. This large amount of data needs to be significantly reduced to achieve low bitrate coding.

[0074] DirAC metadata grouping and averaging To reduce the number of parameters, the parameters calculated in each time-frequency tile are first grouped and averaged over several time slots along the frequency parameter band. The grouping is separated between diffuseness and direction, which is an important aspect of the present invention. In fact, the separation takes advantage of the fact that diffuseness preserves longer-term sound field characteristics than direction, which is a more reactive spatial cue.

[0075] The parameter bands constitute a non-uniform, non-overlapping decomposition of the frequency bands, roughly according to integer multiples of the Equivalent Rectangular Bandwidth (ERB) scale. By default, a 9x ERB scale is adopted for a total of five parameter bands for a 16 kHz audio bandwidth. The diffuseness is

number

[0076] The direction vector in Cartesian coordinates is

number

[0077] The parameter α allows for compressing or expanding the power-based weights in the weighted summation performed to average the parameters. In the preferred mode, α=1.

[0078] In general, this value can be a non-negative real number, as exponents smaller than 1 can also be useful. For example, 0.5 (the square root) still gives more weight to higher amplitude related signals, but is conservative compared to exponents of 1 or greater.

[0079] After grouping and averaging, the resulting direction vector dv[g,b] is generally no longer a unit vector, so normalization is necessary.

[0080]

number

[0081]

[0046] Subsequently, preferred embodiments of the second aspect of the present invention will be discussed. Figure 4a shows an apparatus for encoding directional audio coding parameters including diffusion parameters and directional parameters according to a further second aspect. The apparatus comprises a parameter quantizer 210 that receives at its input the grouped parameters discussed with respect to the first aspect or ungrouped or differently grouped parameters.

[0082] Thus, the parameter quantizer 210 and the subsequently connected parameter encoder 220 for encoding the quantized diffusion parameters and the quantized directional parameters, together with an output interface for generating an encoded parametric representation including information about the encoded diffusion parameters and the encoded directional parameters, are included, for example, in block 200 of Fig. 1a. The quantizer and encoder processor 200 of Fig. 2a may be implemented, for example, as discussed below with respect to the parameter quantizer 210 and the parameter encoder 220, although the quantizer and encoder processor 200 may also be implemented in any manner different from the first aspect.

[0083] Preferably, the parameter quantizer 210 of FIG. 4a is configured to quantize the diffusion parameters using a non-uniform quantizer to generate diffusion indices, as shown at 231 in FIG. 4b. The parameter encoder 220 of FIG. 4a is configured to entropy code the diffusion values ​​obtained for a frame, as shown at 232, i.e., preferably using three different modes, although a single mode or only two different modes may be used. One mode is a raw mode, in which individual diffusion values ​​are coded, for example, using a binary code or a punctured binary code. Alternatively, differential coding may be performed, in which each difference and the original absolute value are coded using the raw mode. However, the situation may be such that the same frame has the same diffusion characteristic across all frequency bands, and only one code value may be used. Again, the situation may be such that there are only continuous values ​​for diffusion characteristic, i.e., continuous diffusion indices within a frame, and then a third coding mode may be applied, as shown at 232.

[0084] Figure 4c shows an implementation of the parameter quantizer 210 of Figure 4a. The parameter quantizer 210 of Figure 4a is configured to convert the direction parameters into polar form, as shown at 233. At block 234, the quantization precision of the bin is determined. This bin may be the original high-resolution bin, or alternatively, preferably, may be a grouped bin of lower resolution.

[0085] As discussed above with respect to Figures 3b and 3c, each band has the same diffusion value but four different directional values. The same quantization precision is determined for the entire band, i.e., for all directional parameters within the band. In block 235, the elevation angle parameter output by block 233 is quantized using the quantization precision. The quantization alphabet for quantizing the elevation angle parameter is preferably also obtained from the bin quantization precision determined in block 234.

[0086] For purposes of processing the azimuth angle values, an azimuth angle alphabet is determined from the elevation angle information for the corresponding (grouped) time / frequency bin (236). The elevation angle information may be a quantized elevation angle value, an original elevation angle value, or a quantized and dequantized elevation angle value, where the latter value, i.e., a quantized and dequantized elevation angle value, is preferred to have the same situation at the encoder and decoder sides. In block 237, the azimuth angle parameter is quantized using the alphabet of this time / frequency bin. The quantization precision may be the same as discussed above with respect to FIG. 3b, but may nevertheless have a different azimuth angle alphabet for each individual grouped time / frequency bin associated with the direction parameter.

[0087] DirAC Metadata Coding For each frame, the DirAC spatial parameters are calculated on a grid consisting of nbands bands across frequencies, and for each frequency band b, the num_slots time slots are grouped into a number of equal-sized nblocks(b) time groups. A spreading parameter is transmitted for each frequency band, and a directional parameter is transmitted for each time group in each frequency band.

[0088] For example, if nbands=5 and nblocks(b)=4, and num_slots=16, this results in 5 diffusion parameters and 20 directional parameters per frame, which are further quantized and entropy coded.

[0089] Quantization of diffusion parameters Each diffusion parameter diff(b) is quantized to one of the diff_alph discrete levels using a non-uniform quantizer that generates a diffusion index diff_idx(b). For example, the quantizer may be derived from the ICC quantization tables used in the MPS standard, where the thresholds and reconstruction levels are calculated by the generate_diffuseness_quantizer function.

[0090] Preferably, only non-negative values ​​from the ICC quantization table are used, such as icc=[1.0, 0.937, 0.84118, 0.60092, 0.36764, 0.0], which includes only 6 levels out of the original 8. Since an ICC of 0.0 corresponds to a diffuseness of 1.0, and an ICC of 1.0 corresponds to a diffuseness of 0.0, a set of y coordinates is created as y=1.0-icc, and the corresponding set of x coordinates is created as x=[0.0, 0.2, 0.4, 0.6, 0.8, 1.0]. A shape-preserving piecewise cubic interpolation method known as the Piecewise Cubic Hermite Interpolating Polynomial (PCHIP) is used to derive a curve that passes through the set of points defined by x and y. The number of steps in the diffusion quantizer is diff_alph, which is 8 in the proposed implementation, regardless of the total number of levels in the ICC quantization table, which is also 8.

[0091] A new set of equally spaced coordinates of diff_alph are generated, x-interpolated from 0.0 to 1.0 (or close to but less than 1.0, if pure diffuseness of 1.0 is avoided due to sound rendering considerations), and the corresponding y-values ​​on the curve are used as reconstruction values, with those reconstruction values ​​spaced non-linearly. Intermediate points between successive x-interpolated values ​​are also generated, and the corresponding y-values ​​on the curve are used as thresholds to determine which values ​​map to particular diffusion indices and therefore reconstruction values. For the proposed implementation, the generated (rounded to 5 digits) reconstruction values ​​and thresholds calculated by the generate_diffuseness_quantizer function are: reconstructions=[0.0,0.03955,0.08960,0.15894,0.30835,0.47388,0.63232,0.85010] thresholds=[0.0,0.01904,0.06299,0.11938,0.22119,0.39917,0.54761,0.73461,2.0] is.

[0092] To make the search easier, a placeholder for a large out-of-range threshold value (2.0) is added at the end of the thresholds. For example, for a specific band b, if diff(b) = 0.33, then thresholds[4] <= diff(b) < thresholds[5], and thus diff_idx(b) = 4, and the corresponding reconstruction value is reconstructions[4] = 0.30835.

[0093] The above procedure is just one possible choice of a non-linear quantizer for the diffusion value.

[0094] Entropy coding of diffusion parameters The EncodeQuasiUniform(value, alphabet_sz) function is used to encode value using quasi-uniform probabilities with a punctured code. For value ∈ {0,..., alphabet_sz - 1}, some of the smallest ones are

Number

Number

[0095] Depending on those values, the quantized diffusion index can be entropy-coded using one of three available methods: raw coding, one value only, and two consecutive values only. The first bit (diff_use_raw_coding) indicates whether the raw coding method is used. In the case of raw coding, each diffusion index value is encoded using the EncodeQuasiUniform function.

[0096] If all index values ​​are equal, the one-value-only method is used. The second bit (diff_have_unique_value) is used to indicate this method, and the unique value is then encoded using the EncodeQuasiUniform function. If all index values ​​consist of only two consecutive values, the two-consecutive-value-only method indicated by the second bit above is used. The smaller of the two consecutive values ​​is encoded using the EncodeQuasiUniform function, taking into account that the alphabet size is reduced to diff_alpha-1. Then, for each value, the difference between that value and the minimum value is encoded using one bit.

[0097] The preferred EncodeQuasiUniform(value,alphabet_sz) function implements a so-called punctured code, which in pseudo-code is: bits = floor(log2(alphabet_sz)) thresh = 2 ^ (bits + 1) - alphabet_sz if (value < thresh) write_bits(value, bits) else write_bits(value + thresh, bits + 1) is defined as:

[0098] If alphabet_sz is a power of 2, then alphabet_sz=2^bits and thresh=2^bits, so the else branch is not used and binary coding results. Otherwise, the first minimum value of thresh is encoded using a binary code with bits bits, and the rest starting with value=thresh is encoded using a binary code with bits+1 bits. The first binary code encoded using bits+1 bits has the value value+thresh=thresh+thresh=thresh*2, so the decoder can understand whether it needs to read one more additional bit by reading only the first bit bits and comparing that value with thresh. The decoding function DecodeQuasiUniform(alphabet_sz) is written in pseudocode as follows: bits = floor(log2(alphabet_sz)) thresh = 2 ^ (bits + 1) - alphabet_sz value = read_bits(bits) if (value >= thresh) value = (value * 2 + read_bits(1)) - thresh return value It can be defined as:

[0099] Converting direction parameters to polar coordinates dv[0] 2 +dv[1] 2 +dv[2] 2 Each 3D direction vector dv, normalized so that = 1, is converted to a polar coordinate representation with elevation angle el ∈ [-90,90] and azimuth angle az ∈ [0,360] using the function DirectionVector2AzimuthElevation. The reverse conversion from polar coordinates to normalized direction vectors is achieved using the function AzimuthElevation2DirectionVector.

[0100] Quantization of direction parameters The directions, represented as elevation and azimuth angle pairs, are further quantized: for each quantized diffusion index level, the required angular precision is selected from the angle_spacing configuration vector as deg_req = angle_spacing(diff_idx(b)) and used to generate a set of quasi-uniformly distributed quantized points on the unit sphere.

[0101] The angle interval value deg_req is preferably not calculated from the diffusivity diff(b), but from the diffusion index diff_idx(b). Thus, there are one possible deg_req value of diff_alph for each possible diffusion index. On the decoder side, the original diffusivity diff(b) is not available, only the diffusion index diff_idx(b) is available, which can be used to select the same angle interval value as in the encoder. In the proposed implementation, the angle interval table is angle_spacing_table=[5.0,5.0,7.5,10.0,18.0,30.0,45.0,90.0] is.

[0102] The quasi-uniformly distributed points on the unit sphere are generated in a way that satisfies several important desirable properties: the points should be symmetrically distributed with respect to the X, Y, and Z axes; the quantization of a given direction to the nearest point and the mapping to an integer index should be a constant-time operation; and finally, computing the corresponding point on the sphere from the inverse quantization to the integer index and direction should be a constant-time or logarithmic-time operation with respect to the total number of points on the sphere.

[0103] There are two types of axial symmetry for points on a horizontal plane: two points exist where the orthogonal axes intersect the unit sphere in the current plane, and no points exist. For the example of any horizontal plane, there are three possible cases: if the number of points is a multiple of four, such as 8, there is symmetry about the X (left-right) axis, two points exist on the Y axis at 90 degrees and 270 degrees, and there is symmetry about the Y (front-back) axis, two points exist on the X axis at 0 degrees and 180 degrees. If the number of points is only a multiple of two, such as 6, there is symmetry about the X axis, but no points exist on the Y axis at 90 degrees and 270 degrees, and there is symmetry about the Y axis, but two points exist on the X axis at 0 degrees and 180 degrees. Finally, if the number of points is any integer, such as 5, there is symmetry about the X axis, but no points exist on the Y axis at 90 degrees and 270 degrees, and there is no symmetry about the Y axis.

[0104] In a preferred embodiment, having points at 0, 90, 180, and 270 degrees on all horizontal planes (corresponding to all quantized elevation angles) is considered useful from a psychoacoustic perspective, and means that the number of points on each horizontal plane is always a multiple of 4. However, depending on the particular application, the condition on the number of points on each horizontal plane can be relaxed to only multiples of 2, or to any integer.

[0105] Additionally, in the preferred embodiment, for each elevation angle, there is always an "origin" azimuth point in the privileged direction of 0 degrees (forward). This property can be relaxed by distributing the azimuth points relative to the 0 degree direction instead of in it, and choosing a pre-calculated quantization offset angle for each elevation angle separately. This can be easily implemented by adding the offset before quantization and subtracting it after dequantization.

[0106] The required angular accuracy is deg_req and should be a divisor of 90 degrees. If not, the required angular accuracy should be determined before actual use.

number

number

[0107] At the equator, the azimuth angle az is uniformly quantized with a step size of deg_req to produce an az_idx quantization index out of 4·n_points. For other elevation angles, the horizontal angular interval from the center of the unit sphere corresponding to the chord length between two consecutive points can be approximated by the arc length of a horizontal circle located at the q_el elevation angle. Thus, the number of points corresponding to 90 degrees on this horizontal circle is reduced in proportion to its radius compared to the equatorial circle, so that the arc length between two consecutive points remains approximately the same everywhere. At the poles, the total number of points is reduced to 1.

[0108] There exists a quantization index of az_alph=max(4·round(radius_len·n_points),1), where radius_len=cos(q_el), corresponding to the elevation angle of q_el. The corresponding quantization index is az_idx=round((az÷360)·az_alph), where the resulting value of az_alph is replaced by 0. This index corresponds to the dequantized azimuth angle of q_az=az_idx·(360÷az_alph). As a note, except for the poles where az_alph=1, the minimum values ​​near the poles are az_alph=4 for deg_req=90 and deg_req=45, and az_alph=8 for all the rest.

[0109] If the requirement on the number of points on each horizontal plane is relaxed to only multiples of 2, then there are 2·n_points corresponding to 180 degrees in the equatorial plane, so the azimuthal alphabet becomes az_alph=max(2·round(radius_len·(2·n_points)),1). If the requirement on the number of points is relaxed to any integer, there are 4·n_points corresponding to 360 degrees in the equatorial plane, so the azimuthal alphabet becomes az_alph=max(round(radius_len·(4·n_points)),1). In either case, the number of points is always a multiple of 4, since radius_len=1 and n_points is an integer in the equatorial plane.

[0110] The quantization and dequantization processes described above are accomplished using the QuantizeAzimuthElevation and DequantizeAzimuthElevation functions, respectively.

[0111] Preferably, the round(x) function rounds x to the nearest integer and is usually implemented in fixed point as round(x) = floor(x + 0.5). Rounding for ties that are exactly halfway between integers, such as 1.5, can be done in several ways. The above definition rounds ties to +infinity (1.5 rounds to 2 and 2.5 rounds to 3). Floating point implementations usually have a native round-to-integer function that rounds ties to an even integer (1.5 rounds to 2 and 2.5 rounds to 2).

[0112] Figure 4d, denoted as "Quasi-uniform coverage of the unit sphere," shows an example of quasi-uniform coverage of the unit sphere using 15-degree angular precision, showing quantized directions. The 3D view is from above, and only the upper hemisphere is drawn for better visualization; the connecting dotted spiral lines are simply for easier visual identification of points from the same horizontal circle or plane.

[0113] Next, a preferred implementation of the parameter encoder 220 of FIG. 4a for encoding quantized direction parameters, i.e., a quantized elevation index and a quantized azimuth index, is shown. As shown in FIG. 5a, the encoder is configured to classify (240) each frame with respect to the diffusion values ​​within the frame. Block 240 receives diffusion values, which in the embodiment of FIG. 3b are only five diffusion values ​​for the frame. If the frame consists of only low diffusion values, a low diffusion coding mode 241 is applied. If the five diffusion values ​​in the frame are only high diffusion values, a high diffusion coding mode 242 is applied. If it is determined that the diffusion values ​​in the frame are both above and below the diffusion threshold ec_max, a mixed diffusion coding mode 243 is applied. In both the low diffusion coding mode 241 and the high diffusion coding mode 242, as well as for the low diffusion band of the mixed diffusion frame, raw coding is attempted on the one hand and entropy coding is attempted on the other hand, i.e., as shown in 244a, 244b, and 244c. However, for high diffuseness bands in mixed diffuseness frames, raw coding mode is always used, as shown at 244d.

[0114] When different coding modes, i.e., raw coding mode and entropy coding mode (with modeling), are used, the result is that the encoder controller selects the mode that results in fewer bits for encoding the quantized indexes, as shown at 245a, 245b, and 245c.

[0115] Alternatively, only raw coding mode can be used for all frames and bands, or only entropy coding mode with modeling can be used for all bands, or any other coding mode that codes indices, such as Huffman coding mode or arithmetic coding mode with or without context adaptation, can be used.

[0116] Depending on the outcome of the selected procedures in blocks 245a, 245b, and 245c, the side information is set for the entire frame, as shown in blocks 246a, 246b, or for only the corresponding band, i.e., the low-diffusivity band, in block 246c. Alternatively, for item 246c, the side information can also be set for the entire frame. In this case, even if the side information is set for the entire frame, the determination of the high-diffusivity band can be made solely in the decoder, so that the decoder determines that a mixed-diffusivity frame is present, and so that the side information for the frame indicates an entropy coding mode with modeling, but the directional parameters for bands with high-diffusivity values ​​within this mixed-diffusivity frame are coded in raw coding mode.

[0117] In the preferred embodiment, diff_alph=8. The ec_max threshold was then chosen to be 5 by minimizing the average compressed size in a large test corpus. This threshold ec_max is used in the following modes depending on the range of values ​​of the diffusion index of the current frame: For low to medium diffusivity frames where diff_idx(b)<=ec_max for all bands b, all directions are coded using both raw and entropy coding, and the best one is selected and indicated by 1 bit as side information (identified above as dir_use_raw_coding); For mixed diffusivity frames where diff_idx(b)<=ec_max for some bands b, the directions corresponding to those bands are coded as in the first case, but for other high diffusivity bands b where diff_idx(b)>ec_max, the directions corresponding to these other bands are always coded as raw (to avoid mixing the entropy coding statistics of directions with low to medium diffusivity with directions with high diffusivity that are also very coarsely quantized); For high-diffusivity frames where diff_idx(b)>ec_max for all bands b, the ec_max threshold is pre-set to ec_max=diff_alph for the current frame (this setting can also be pre-set on the decoder side as well, since the diffusion index is coded before the direction), so this case becomes the first case.

[0118] 5b shows a preferred but optional preprocessing of the direction indices for both modes. For both modes, the quantized direction indices, i.e., the quantized azimuth indices and the quantized elevation indices, are processed in block 247 into a conversion of the elevation / azimuth indices that results in signed values, where a zero index corresponds to an elevation or azimuth angle of zero. To obtain a more compact representation of the permuted unsigned azimuth / elevation indices, a subsequent conversion 248 to unsigned values ​​is performed that includes interleaving of positive / negative values.

[0119] FIG. 5c shows a preferred implementation of the first coding mode 260, i.e., raw coding mode without modeling. The preprocessed azimuth / elevation indices are input to block 261 to merge both indices into a single spherical index. Based on the quantization precision, i.e., deg_req, derived from the associated diffusion index, encoding is performed (262) using a coding function such as EncodeQuasiUniform or a (punctured) binary code. Thus, coded spherical indices for either a band or an entire frame are obtained. Coded spherical indices for the entire frame are obtained only for low-diffusivity frames for which raw coding is selected, or only for high-diffusivity frames for which raw coding is again selected, or coded spherical indices for only the high-diffusivity band of the frame are obtained for mixed-diffusivity frames, as shown in 243 in FIG. 5a, for which a second coding mode, such as entropy coding with modeling, is selected for the other bands with low or medium diffusivity.

[0120] FIG. 5d illustrates this second encoding mode, which may be, for example, an entropy coding mode with modeling. The preprocessed indices classified into, for example, a mixed diffusive frame, as shown in FIG. 5a at 240, are input to block 266, which collects corresponding quantized data, such as elevation index, elevation alphabet, azimuth index, and azimuth alphabet, which are collected into individual vectors for the frame. In block 267, averages are explicitly calculated for the elevation and azimuth angles based on information derived from the inverse quantization and the corresponding vector transformation, as discussed below. These averages are quantized to the highest angular precision used in the frame, as shown in block 268. Predicted elevation and azimuth indices are generated from the averages, as shown in block 269, and the signed distances for the elevation and azimuth angles associated with the predicted elevation and azimuth indices from the original indices are calculated and, optionally, reduced to another smaller interval of values.

[0121] As shown in Figure 5e, the data generated by the modeling operation using the projection operation to derive the predicted values ​​shown in Figure 5d is entropy coded. This coding operation, shown in Figure 5e, ultimately generates coded bits from the corresponding data. In block 271, the azimuth and elevation mean values ​​are converted to signed values, and a specific permutation 272 is performed to have a more compact representation. These mean values ​​are then coded using a binary code or a punctured binary code (273) to generate an elevation mean bit 274 and an azimuth mean bit. In block 275, the Golomb-Rice parameters are determined, as shown in Figure 5f. These parameters are then also coded using a (punctured) binary code (276) to have a Golomb-Rice parameter for the elevation angle and another Golomb-Rice parameter for the azimuth angle shown at 277. In block 278, the (reduced) signed distances calculated by block 270 are reordered and coded using the extended Golomb-Rice method shown at 279 to have coded elevation and azimuth distances shown at 280.

[0122] 5f shows a preferred implementation for determining the Golomb-Rice parameters in block 275, which is performed for determining both the elevation Golomb-Rice parameters or the azimuth Golomb-Rice parameters. In block 281, an interval is determined for the corresponding Golomb-Rice parameter. In block 282, for each candidate value, the total number of bits for all reduced signed distances is calculated, and in block 283, the candidate value resulting in the smallest number of bits is selected as the Golomb-Rice parameter for either the azimuth or elevation processing.

[0123] Next, Figure 5g will be discussed to further illustrate the procedure in block 279 of Figure 5e, i.e., the extended Golomb-Rice method. Based on the selected Golomb-Rice parameter p, the range index for either elevation or azimuth is separated into a most significant portion MSP and a least significant portion LSP, as shown on the right side of block 284. In block 285, if the MSP is the maximum possible value, the trailing zero bits of the MSP portion are removed, and in block 286 the result is encoded using a (punctured) binary code.

[0124] The LSP portion is also coded using a (punctured) binary code as shown at 287. Thus, on lines 288 and 289, the coded bits for the most significant portion MSP and the coded bits for the least significant portion LSP are obtained, which together represent the corresponding coded reduced signed distance for either elevation or azimuth.

[0125] Figure 8d shows an example of coded directions. Mode bits 806 indicate, for example, an entropy coding mode with modeling. Item 808a indicates azimuth mean bits, and item 808b indicates elevation mean bits, as discussed above with respect to item 274 of Figure 5e. Golomb-Rice azimuth parameters 808c and Golomb-Rice elevation parameters 808d are also included in coded form in the bitstream of Figure 8d, corresponding to those discussed above with respect to item 277. Coded elevation distances and coded azimuth distances 808e and 808f are included in the bitstream as obtained at 288 and 289, or as discussed above with respect to item 280 in Figures 5e and 5g. Item 808g indicates additional payload bits for additional elevation / azimuth distances. The averages over elevation and azimuth angles, and the Golomb-Rice parameters for elevation and azimuth angles are only needed once per frame, but if the frame is very long or the signal statistics vary significantly within a frame, they can be calculated twice per frame, etc., if necessary.

[0126] Figure 8c shows the bitstream when the mode bits indicate raw coding as defined by block 260 of Figure 5c. Mode bits 806 indicate the raw coding mode, and item 808 indicates the payload bits for the sphere index, i.e., the result of block 262 of Figure 5c.

[0127] Entropy coding of directional parameters When coding the quantized directions, the elevation index el_idx is always coded first, followed by the azimuth index az_idx. If the current configuration only considers the horizontal equatorial plane, nothing is coded for the elevation angle, which is assumed to be zero everywhere.

[0128] Before coding, signed values ​​are mapped to unsigned values ​​using a generic reordering transformation, which interleaves positive and negative numbers into unsigned numbers as u_val=2·|s_val|-(s_val<0), implemented by the ReorderGeneric function. The expression (condition) evaluates to 1 if the condition is true, and to 0 if the condition is false.

[0129] Since some smaller unsigned values ​​are coded more efficiently using one bit less using the EncodeQuasiUniform function, both the already unsigned elevation and azimuth indices are converted to signed, and only afterwards is the ReorderGeneric function applied, so that a signed index value of zero corresponds to an elevation or azimuth angle of zero. By converting to signed first, the zero value is located in the middle of the signed interval of possible values, and after applying the ReorderGeneric function, the resulting unsigned reordered elevation index value is

number

[0130] In raw coding without modeling, the two unsigned permuted indices are merged into one unsigned sphere index: sphere_idx = sphere_offsets(deg_req, el_idx_r) + az_idx_r, where the sphere_offsets function calculates the sum of all azimuth alphabets az_alph corresponding to unsigned permuted elevation indices less than el_idx_r. For example, for deg_req = 90, if el_idx_r = 0 (elevation angle 0 degrees) has az_alph = 4, el_idx_r = 1 (elevation angle -90 degrees) has az_alph = 1, and el_idx_r = 2 (elevation angle 90 degrees) has az_alph = 1, then sphere_offsets(90,2) takes the value 4 + 1. If the current configuration only considers the horizontal equatorial plane, el_idx_r is always 0 and the unsigned sphere index simplifies to sphere_idx = az_idx_r. In general, the total number of points on the sphere, or the sphere point count, is sphere_alph = sphere_offsets(deg_req,el_alph+1).

[0131] The unsigned spherical index shpere_idx is coded using the EncodeQuasiUniform function. For entropy coding with modeling, the quantized directions are grouped into two categories. The first category includes quantized directions for entropy-coded diffusion indices diff_idx(b)≦ec_max, and the second category includes quantized directions for raw-coded diffusion indices diff_idx(b)>ec_max, where ec_max is a threshold optimally selected depending on diff_alph. This approach implicitly excludes frequency bands with high diffusivity from entropy coding if frequency bands with low to medium diffusivity are also present in the frame to avoid mixing residual statistics. For mixed-diffusivity frames, raw coding is always used for frequency bands with high diffusivity. However, if all frequency bands have high diffusivity and diff_idx(b)>ec_max, the threshold is preset to ec_max=diff_alph to enable entropy coding for all frequency bands.

[0132] For the first category of entropy coded quantized directions, the corresponding elevation index el_idx, elevation alphabet el_alph, azimuth index az_idx, and azimuth alphabet az_alph are collected into separate vectors for further processing.

[0133] The average direction vector is derived by converting each entropy-coded quantized direction back to a direction vector, calculating either the mean, median, or mode of the direction vectors including renormalization, and converting the average direction vector to an average elevation angle el_avg and azimuth angle az_avg. These two values ​​are quantized using the highest angular precision deg_req used by the entropy-coded quantized directions, indicated by deg_req_avg. This angular precision is typically the required angular precision corresponding to the smallest spreading index min(diff_idx(b)), for b ∈ {0,...,nbands-1} and diff_idx(b) ≤ ec_max.

[0134] el_avg is quantized normally, using the corresponding n_points_avg value derived from deg_req_avg to produce el_avg_idx and el_avg_alph, while az_avg is quantized using the precision at the equator to produce az_avg_idx and az_avg_alph=4·n_points_avg.

[0135] For each direction to be entropy coded, the dequantized average elevation angle q_el_avg and azimuth angle q_az_avg are projected using the precision of that direction to obtain a predicted elevation angle and azimuth angle index. For an elevation angle index el_idx, its precision, which may be derived from el_alph, is used to calculate a projected average elevation angle index el_avg_idx_p. For a corresponding azimuth angle index az_idx, its precision in the horizontal circle located at q_el elevation angle, which may be derived from az_alph, is used to calculate a projected average azimuth angle index az_avg_idx_p.

[0136] The projections to obtain the predicted elevation and azimuth indices can be calculated in several equivalent ways. For elevation,

number

number

number

number

[0137] The signed distance el_idx_dist is calculated as the difference between each elevation index el_idx and its corresponding el_avg_idx_p. In addition, since the difference produces values ​​in the interval {-el_alph+1,...,el_alph-1}, they are scaled like a modular operation by adding el_alph for values ​​that are too small and subtracting el_alph for values ​​that are too large.

number

[0138] Similarly, the signed distance az_idx_dist is calculated as the difference between each azimuth index az_idx and its corresponding az_avg_idx_p. The difference operation produces values ​​in the interval {-az_alph+1,...,az_alph-1}, which are reduced to the interval {-az_alph÷2,...,az_alph÷2-1} by adding az_alph for values ​​that are too small and subtracting az_alph for values ​​that are too large. If az_alph=1, the azimuth index is always az_idx=0 and nothing needs to be coded.

[0139] Depending on their values, the quantized elevation and azimuth indices can be coded using one of two available methods: raw coding or entropy coding. The first bit (dir_use_raw_coding) indicates whether the raw coding method is used. In the case of raw coding, the merged sphere_index single unsigned sphere index is coded directly using the EncodeQuasiUniform function.

[0140] The entropy coding consists of several parts: all quantized elevation and azimuth indices corresponding to the diffusion index diff_idx(b)>ec_max are coded in the same way as raw coding; then, in all other cases, the elevation part is entropy coded first, followed by the azimuth part.

[0141] The elevation part consists of three components: the average elevation index, the Golomb-Rice parameters, and the reduced signed elevation distance. The average elevation index el_avg_idx is converted to signed so that the zero value is in the center of the signed interval of possible values, the ReorderGeneric function is applied, and the result is coded using the EncodeQuasiUniform function. The Golomb-Rice parameters, with an alphabet size depending on the maximum alphabet size of the elevation index, are coded using the EncodeQuasiUniform function. Finally, for each reduced signed elevation distance el_idx_dist, the ReorderGeneric function is applied to generate el_idx_dist_r, and the result is coded using the extended Golomb-Rice method with the parameters shown above.

[0142] For example, if the highest angle precision used, deg_req_min, is 5 degrees, then the maximum elevation alphabet size, el_alph, should be:

number

number

[0143] The azimuth part also consists of three components: the average azimuth index, the Golomb-Rice parameter, and the reduced signed azimuth distance. The average azimuth index az_avg_idx is converted to a signed value so that the zero value is in the center of the signed interval of possible values, the ReorderGeneric function is applied, and the result is coded using the EncodeQuasiUniform function. The Golomb-Rice parameter, with an alphabet size according to the maximum alphabet size of the azimuth index, is coded using the EncodeQuasiUniform function. Finally, for each reduced signed azimuth distance az_idx_dist, the ReorderGeneric function is applied to generate az_idx_dist_r, and the result is coded using the extended Golomb-Rice method with the parameters shown above.

[0144] For example, if the highest angular precision used, deg_req_min, is 5 degrees, then the maximum azimuth alphabet size, az_alph, is:

number

[0145] An important property to take into account for efficient entropy coding is that each permuted reduced elevation distance el_idx_dist_r may have a different alphabet size, which is exactly el_alph of the original elevation index value el_idx and depends on the corresponding spreading index diff_idx(b), and each permuted reduced azimuth distance az_idx_dist_r may have a different alphabet size, which is exactly az_alph of the original azimuth index value az_idx and depends on both the corresponding q_el of its horizontal circle and the spreading index diff_idx(b).

[0146] The existing Golomb-Rice entropy coding method with integer parameter p ≥ 0 is used to code an unsigned integer u. First, u is divided into a least significant part u_lsp = u mod 2 with p bits. p and the top part

number

[0147] Because arbitrarily large integers can be coded, some coding efficiency may be lost if the actual values ​​to be coded have a known, relatively small alphabet size. Another drawback is the possibility of decoding out-of-range or invalid values, or reading too many 1 bits, in the event of a transmission error or a deliberately crafted invalid bit stream.

[0148] The extended Golomb-Rice method combines three improvements to the existing Golomb-Rice method for coding vectors of values, each with a known and potentially different alphabet size u_alpha. First, the alphabet size of the most significant part is:

number

[0149] For u_msp_alph=3, a threshold of 3 is particularly preferable since the restricted Golomb-Rice codewords for the most significant part are 0, 10, 11, and therefore the total code lengths are 1+p, 2+p, and 2+p, where p is the number of bits for the least significant part; a punctured code is always optimal for lengths up to two, so it is used instead, replacing both the most significant and least significant parts.

[0150] Furthermore, it should be noted that if the function EncodeQuasiUniform is exactly a punctured code and the alphabet size is a power of 2, it implicitly becomes a binary code. In general, punctured codes are optimal and uniquely determined given the alphabet size, and generate codes of only one or two lengths. For three or more consecutive code lengths, the number of possible codes is no longer quasi-uniform, and various choices exist for the number of possible codes of each length.

[0151] The invention is not limited to the exact description above. Alternatively, the invention can be easily extended in the form of an inter-frame predictive coding scheme, where, instead of calculating a single average directional vector for the entire current frame for each parameter band, the average directional vector is calculated by using previous directional vectors from the current frame and optionally from previous frames throughout time, and quantizing and coding them as side information. This solution has the advantage of being more efficient in coding, but is less robust to possible packet losses.

[0152] Figures 6a to 6g illustrate further steps performed in the encoder as discussed above. Figure 6a shows a general overview of a parameter quantizer 210, consisting of a quantized elevation function 210a, a quantized azimuth function 210b, and a dequantized elevation function 210c. The preferred embodiment of Figure 6a shows a parameter quantizer with azimuth function 210c that depends on the quantized and then dequantized elevation value q_el.

[0153] Figure 6c shows a corresponding inverse quantizer for inverse quantizing the elevation angle, as discussed above with respect to Figure 6a for the encoder. However, the embodiment of Figure 6b is also useful for the inverse quantizer shown in item 840 of Figure 8a. Based on the inverse quantization precision deg_req, the elevation angle index, on the one hand, and the azimuth angle index, on the other hand, are inverse quantized to finally obtain the inverse quantized elevation angle value q_el and the inverse quantized azimuth angle value q_az. Figure 6c shows the first encoding mode, i.e., the raw coding mode discussed with respect to items 260 to 262 in Figure 5c. Figure 6c further shows the preprocessing discussed in Figure 5b, showing the conversion of the elevation angle data to signed values ​​at 247a and the corresponding conversion of the azimuth angle data to signed values ​​at 247b. Reordering is performed for the elevation angle, as shown at 248a, and for the azimuth angle, as shown at 248b. A sphere point counting procedure is performed in block 248c to calculate a sphere alphabet based on the quantization or dequantization precision. In block 261, merging of both indices into a single sphere index is performed, and encoding in block 262 is performed using a binary or punctured binary code, where in addition to this sphere index, the sphere alphabet for the corresponding dequantization precision is also derived as shown in Figure 5c.

[0154] Figure 6d shows the procedure performed for the entropy coding mode with modeling. In item 267a, dequantization of azimuth and elevation data is performed based on the corresponding index and dequantization precision. The dequantized values ​​are input to block 267b to calculate a direction vector from the dequantized values. In block 267c, averaging is performed on vectors with associated spreading indices below the corresponding threshold to obtain an averaged vector. In block 267d, the direction-averaged direction vector is again converted to an elevation average and an azimuth average, and these values ​​are then quantized using the highest precision as determined by block 268e. This quantization is shown in 268a and 268b, and results in a corresponding quantized index and a quantization alphabet, where the alphabet is determined by the quantization precision for the average value. In blocks 268c and 268d, dequantization is again performed to obtain dequantized average values ​​for the elevation and azimuth angles.

[0155] In Figure 6e, the projected elevation angle average is calculated in block 269a, and the projected azimuth angle average is calculated in block 269b, i.e., Figure 6e shows a preferred implementation of block 269 of Figure 5d. As shown in Figure 6e, blocks 269a, 269b preferably receive quantized and dequantized average values ​​for elevation and azimuth angles. Alternatively, projection can be performed directly on the output of block 267d, but the quantization and dequantization procedure is preferred for higher accuracy and better compatibility with the state at the encoder and decoder sides.

[0156] Figure 6f shows a procedure corresponding to block 270 of Figure 5d in a preferred embodiment. In blocks 278a, 278b, the corresponding difference, or "distance" as it is called in block 270 of Figure 5d, is calculated between the original index and the projected index. Corresponding interval reduction is performed in block 270c for the elevation data and in 270d for the azimuth data. Following the reordering in blocks 270e, 270f, the data to be subjected to extended Golomb-Rice coding is obtained, as discussed above with respect to Figures 5e through 5g.

[0157] Figure 6g shows further details regarding the procedure performed to generate coded bits for the elevation and azimuth averages. Blocks 271a and 271b show the conversion of the elevation and azimuth average data to signed data, followed by the ReorderGeneric function shown for both types of data in blocks 272a and 272b. Items 273a and 273b show the encoding of this data using a (punctured) binary code, such as the encoding quasi-uniform function discussed above.

[0158] 7a shows a decoder according to a first embodiment for decoding an encoded audio signal including encoded directional audio coding parameters, where the encoded directional audio coding parameters include encoded diffusion parameters and encoded directional parameters. The apparatus includes a parameter processor 300 for decoding the encoded directional audio coding parameters to obtain decoded diffusion parameters having a first time or frequency resolution and decoded directional parameters having a second time or frequency resolution. The parameter processor 300 is connected to a parameter resolution converter 710 for converting the decoded diffusion parameters or decoded directional parameters into transformed diffusion parameters or transformed directional parameters. Alternatively, as indicated by the hedged line, the parameter resolution converter 710 can already perform parameter resolution processing using the encoded parametric data, and the transformed encoded parameters are sent from the parameter resolution converter 710 to the parameter processor 300. In this latter case, the parameter processor 300 then directly supplies the processed, i.e., decoded, parameters to the audio renderer 420. However, it is preferred to perform parameter resolution conversion using the decoded diffusion parameters and the decoded directional parameters.

[0159] The decoded directional and diffusion parameters typically have a third or fourth time or frequency resolution when fed to the audio renderer 420, where the third or fourth resolution is higher than the resolution inherent in those parameters when output by the parameter processor 300.

[0160] Since the time or frequency resolutions inherent to the decoded diffusion parameters and the decoded directional parameters are different from each other, and typically the decoded diffusion parameters have a lower time or frequency resolution compared to the decoded directional parameters, the parameter resolution converter 710 is configured to perform different parameter resolution conversions using the decoded diffusion parameters and the decoded directional parameters. As discussed above with respect to Figures 3a to 3c, the highest resolution used by the audio renderer 420 is that shown in Figure 3b, the medium resolution shown in Figure 3c is that inherent to the decoded directional parameters, and the low resolution inherent to the decoded diffusion parameters is that shown in Figure 3b.

[0161] Figures 3a through 3c are merely examples showing three very specific time or frequency resolutions. Any other time or frequency resolutions, which follow the same pattern of high, medium, and low time or frequency resolutions, can also be applied by the present invention. As shown in the examples of Figures 3b and 3c, both resolutions have the same frequency resolution but different time resolutions, or vice versa, where one time or frequency resolution is lower than another. In this example, the frequency resolution is the same in Figures 3b and 3c, but the time resolution is higher in Figure 3c, as Figure 3c shows medium resolution while Figure 3b shows low resolution.

[0162] The results of the audio renderer 420, operating at the third or fourth higher time or frequency resolution, are then transferred to the spectrum-to-time converter 440, which then generates a time-domain multi-channel audio signal 450, as already discussed above with respect to FIG. 1b. The spectrum-to-time converter 440 converts the data from the spectral domain generated by the audio renderer 420 to the time domain on line 450. The spectral domain in which the audio renderer 420 operates includes a first number of time slots and a second number of frequency bands for a frame. The frame includes a number of time / frequency bins equal to the product of the first number and the second number, where the first and second numbers define a third time or frequency resolution, i.e., a higher time or frequency resolution.

[0163] The resolution converter 710 is configured to generate a number of at least four spreading parameters from the spreading parameters associated with the first time or frequency resolution, where two of the spreading parameters are for time / frequency bins that are adjacent in time and two other of the at least four spreading parameters are for time / frequency bins that are adjacent to each other in frequency.

[0164] Since the time or frequency resolution for the diffusion parameters is lower than that for the directional parameters, the parameter resolution converter is configured to generate, for the decoded diffusion parameters, a number of transformed diffusion parameters, and, for the decoded directional parameters, a second number of transformed directional parameters, where the second number is greater than the first number.

[0165] 7b shows a preferred procedure performed by the parameter resolution converter. In block 721, the parameter resolution converter 710 obtains the diffusion / directional parameters for the frame. In block 722, a multiplication or copy operation of the diffusion parameters to at least four high-resolution time / frequency bins is performed. In block 723, optional processing such as smoothing or low-pass filtering is performed on the copied parameters in the high-resolution representation. In block 724, the high-resolution parameters are applied to the corresponding audio data in the corresponding high-resolution time / frequency bins.

[0166] 8a shows a preferred implementation of a decoder for decoding an encoded audio signal including encoded directional audio coding parameters, including encoded diffusion parameters and encoded directional parameters, according to a first embodiment. The encoded audio signal is input to an input interface 800. The input interface 800 receives the encoded audio signal and separates the encoded diffusion parameters and the encoded directional parameters from the encoded audio signal, typically frame by frame. This data is input to a parameter decoder 820, which generates quantized diffusion parameters and quantized directional parameters from the encoded parameters, where the quantized directional parameters are, for example, azimuth and elevation indices. This data is input to a parameter inverse quantizer 840, which determines dequantized diffusion parameters and dequantized directional parameters from the quantized diffusion parameters and quantized directional parameters. This data can then be used to convert one audio format to another, or to render the audio signal into a multi-channel signal or in any other representation, such as an Ambisonics, MPS, or SAOC representation.

[0167] The dequantized parameters output by block 840 may be input to an optional parameter resolution converter, as discussed above with respect to FIG. 7a, in block 710. The transformed or untransformed parameters may be input to the audio renderers 420, 440 shown in FIG. 8a. If the encoded audio signal additionally includes an encoded transport signal, the input interface 800 is configured to separate the encoded transport signal from the encoded audio signal and provide this data to the audio transport signal decoder 340 already discussed above with respect to FIG. 8b. The result is input to the time-to-spectral converter 430, which provides the audio renderer 420. If the audio renderer 420 is implemented as shown in FIG. 1b, the transformation to the time domain is performed using the synthesis filter bank 440 of FIG. 1b.

[0168] Figure 8b shows a portion of an encoded audio signal, typically organized in a bitstream that references an encoded spreading parameter. The spreading parameter has associated with it preferably two mode bits 802 to indicate the three different modes shown in Figure 8b and discussed above. The encoded data relating to the spreading parameter includes payload data 804.

[0169] The bitstream portions relating to directional parameters are shown in Figures 8c and 8d discussed above, where Figure 8c shows the situation where raw coding mode is selected and Figure 8d shows the situation where entropy decoding mode with modeling is selected / indicated by mode bits or mode flags 806.

[0170] The parameter decoder 820 of Figure 8a is configured to decode the spread payload data for the time / frequency domain, which in a preferred embodiment is a time / frequency domain with low resolution, as shown in block 850. In block 851, a dequantization precision for the time / frequency domain is determined. Based on this dequantization precision, block 852 of Figure 8e illustrates decoding and / or dequantization of the directional parameters using the same dequantization precision for the time / frequency domain with which the spread parameters are associated. The output of Figure 8e is a set of decoded directional parameters for the time / frequency domain, such as one band in Figure 3c, i.e., in the illustrated example, four directional parameters for one band in the frame.

[0171] Figure 8f illustrates further features of the decoder, and in particular the parameter decoder 820 and parameter dequantizer 840 of Figure 8a. Regardless of whether the dequantization precision is determined based on the spreading parameters or explicitly signaled or determined elsewhere, block 852a illustrates determining an elevation alphabet from the signaled dequantization precision for the time / frequency domain. In block 852b, the elevation data is decoded and optionally dequantized using the elevation alphabet for the time / frequency domain to obtain dequantized elevation parameters at the output of block 852b. In block 852c, the azimuth alphabet for the time / frequency domain is determined not only from the dequantization precision from block 851, but also from the quantized or dequantized elevation data to reflect the situation discussed above with respect to the quasi-uniform coverage of the unit sphere of Figure 4d. In block 852d, decoding and optionally dequantization of the azimuth data using the azimuth alphabet is performed for the time / frequency domain.

[0172] The invention according to its second aspect preferably combines those two features, although the two features, i.e. one feature of FIG. 8a or the other feature of FIG. 8f, may also be applied separately from each other.

[0173] Figure 8g shows an overview of parameter decoding depending on whether raw decoding mode or decoding mode with modeling is selected as indicated by mode bit 806 discussed in Figures 8c and 8d. If raw decoding is to be applied, the sphere index for the band is decoded as shown in 862, and the quantized azimuth / elevation parameters for the band are calculated from the decoded sphere index as shown in block 864.

[0174] If decoding with modeling is indicated by mode bit 806, then averages for the azimuth / elevation data in the band / frame are decoded as indicated by block 866. In block 868, distances for the azimuth / elevation information in the band are decoded, and in block 870, the corresponding quantized elevation and azimuth parameters are calculated, typically using summation operations.

[0175] Regardless of whether raw decoding mode or decoding with modeling mode is applied, the decoded azimuth / elevation indices are dequantized (872), as shown at 840 in FIG. 8a, and the results may be converted to Cartesian coordinates for the bands in block 874. Alternatively, if the azimuth and elevation data can be used directly in the audio renderer, such conversion in block 874 is not necessary. If the conversion to Cartesian coordinates is performed anyway, any potentially used parameter resolution conversion may be applied before or after the conversion.

[0176] 9a to 9c are also referred to for further preferred implementations of the decoder. FIG. 9a shows the decoding operation shown in block 862. Depending on the inverse quantization precision determined by block 851 in FIG. 8e or 8f, functional sphere point counting in block 248c is performed to determine the actual sphere alphabet also applied during encoding. The bits related to the sphere index are decoded in block 862, and decomposition into two indices is performed as shown in 864a and given in more detail in FIG. 9a. To finally obtain the elevation index, azimuth index, and corresponding alphabet for subsequent inverse quantization in block 872 of FIG. 8g, permutation functions 864b, 864c, and corresponding transformation functions in blocks 864d and 864e are performed.

[0177] Figure 9b shows the corresponding procedure for another decoding mode, i.e., decoding mode with modeling. In block 866a, the average inverse quantization precision is calculated in line with that discussed above for the encoder side. The alphabet is calculated in block 866b, and the corresponding bits 808a, 808b of Figure 8d are decoded in blocks 866c and 866d. To undo or mimic the corresponding operations performed on the encoder side, reordering functions 866e, 866f are performed in subsequent transform operations 866g, 866h.

[0178] Figure 9c additionally illustrates the complete inverse quantization operation 840 in a preferred embodiment. Block 852a determines the elevation alphabet as discussed above with respect to Figure 8f, and the corresponding calculation of the azimuth alphabet is similarly performed in block 852c. Projection calculation operations 820a, 820e are also performed for the elevation and azimuth angles. Reordering procedures 820b and 820f for the elevation and azimuth angles are similarly performed, as are corresponding summation operations 820c, 820g. Corresponding interval reductions in block 820d for the elevation angle and block 820h for the azimuth angle are similarly performed, and inverse quantization of the elevation angle is performed in blocks 840a and 840b. Figure 9c illustrates that this procedure implies a specific order, i.e., the elevation data is first processed based on the dequantized elevation data, and the decoding and inverse quantization of the azimuth data are performed in a preferred embodiment of the present invention.

[0179] This is followed by a summary of the benefits and advantages of the preferred embodiments. Efficient coding of spatial metadata generated by DirAC without compromising the generality of the model, which is the key to successful integration of DirAC into low bitrate coding schemes. Grouping and averaging of direction and diffusion parameters with different time (or optionally frequency) resolution: Diffusion preserves longer term sound field properties than direction, which is a more reactive spatial cue, so diffusion is averaged over a longer time period than direction. Quasi-uniform dynamic coverage of a 3D sphere, fully symmetric about the X, Y, and Z coordinate axes, and any desired angular resolution is possible. · Quantization and dequantization operations are of constant complexity (no search for the closest code vector is required). · The encoding and decoding of quantized point indices has constant or at most logarithmic complexity with respect to the total number of quantized points on the sphere. · The worst-case entropy coding size of the entire DirAC spatial metadata for one frame is always limited to only 2 bits more than that of the raw coding. · An extended Golomb-Rice coding method that is optimal for coding vectors of symbols with potentially different alphabet sizes. Use the mean direction for efficient entropy coding of the direction, and map the quantized mean direction from the highest resolution to each azimuth and elevation angle resolution. For mixed diffusion frames, always use raw coding for directions with high diffusivity above a predefined threshold. Use the angular resolution for each direction as a function of its corresponding diffusivity.

[0180] A first aspect of the present invention is directed to processing diffusion parameters and directional parameters using first and second time or frequency resolutions, followed by quantization and encoding of each value. This first aspect additionally refers to grouping of parameters using different time / frequency resolutions. Further aspects relate to performing amplitude measurement-related weighting within the grouping, and yet additional aspects relate to weighting for averaging and grouping of directional parameters using corresponding diffusion parameters as the basis for the corresponding weights. The above aspects are explained and elaborated in the first set of claims.

[0181] A second aspect of the invention, detailed subsequently in the enclosed set of examples, is directed to performing quantization and coding, which may be performed without the features outlined in the first aspect, or may be used in conjunction with the corresponding features detailed in the first aspect.

[0182] Thus, all of the various aspects recited in the claims and example set and in the various dependent claims of the claims and examples may be used independently of one another or together, and it is particularly preferred for the most preferred embodiments that all aspects of the claim set be used together with all aspects of the example set.

[0183] The set of examples includes the following examples: 1. An apparatus for encoding directional audio coding parameters including diffusion parameters and direction parameters, comprising: a parameter quantizer (210) for quantizing the diffusion parameters and the directional parameters; a parameter encoder (220) for encoding the quantized diffusion parameters and the quantized directional parameters; an output interface (230) for generating an encoded parametric representation including information about the encoded diffusion parameters and the encoded directional parameters; An apparatus comprising: 2. The apparatus of Example 1, wherein the parameter quantizer (210) is configured to quantize the spreading parameters using a non-uniform quantizer to generate the spreading index. 3. The apparatus of Example 2, wherein the parameter quantizer (210) is configured to derive the non-uniform quantizer using an inter-channel coherence quantization table to obtain thresholds and reconstruction levels for the non-uniform quantizer. 4. The parameter encoder (220) encoding the quantized spreading parameters in raw coding mode using a binary code when the coding alphabet has a size that is a power of two, or encoding the quantized spreading parameters in raw coding mode using a punctured code when the coding alphabet is different from a power of two, or encoding the quantized spreading parameters in a one-value-only mode using the first specific instruction and a codeword for one value from the raw coding mode; or Encoding the quantized spreading parameter in a two consecutive values ​​only mode using a second specific instruction, a code for the smaller of two consecutive values, and a bit for the difference between one or each actual value and the smaller of the two consecutive values. 4. The apparatus of any one of Examples 1 to 3, configured to: 5. The parameter encoder (220) is configured to determine, for all spreading values ​​associated with the time portion or the frequency portion, whether the coding mode is a raw coding mode, a mode of only one value, or a mode of only two consecutive values; Raw mode is signaled using one of two bits, one value only mode is signaled using another of the two bits having a first value, and two consecutive values ​​only mode is signaled using another of the two bits having a second value. The apparatus of Example 4. 6. The parameter quantizer (210) For each direction parameter, it accepts a Cartesian vector with two or three components, Convert a Cartesian vector into a representation with azimuth and elevation values 10. The apparatus of any one of the preceding examples, configured to: 7. The apparatus of any one of the preceding examples, wherein the parameter quantizer (210) is configured to determine a quantization precision for quantization of the directional parameters, the quantization precision being dependent on a diffusion parameter associated with the directional parameters such that directional parameters associated with lower diffusion parameters are quantized more precisely than directional parameters associated with higher diffusion parameters. 8. The parameter quantizer (210) so that the quantized points are quasi-uniformly distributed on the unit sphere, or The quantized points are distributed symmetrically across the x-, y-, or z-axis, or By mapping to integer indices, quantization in a given direction to the nearest quantization point or to one of several nearest quantization points is a constant time operation, or so that the computation of the corresponding point on the sphere from the inverse quantization to integer index and direction is a constant or logarithmic time behavior with respect to the total number of points on the sphere. 8. The apparatus of example 7 configured to determine a quantization precision. 9. The apparatus of Examples 6, 7, or 8, wherein the parameter quantizer (210) is configured to quantize elevation angles having positive and negative values ​​into a set of unsigned quantization indexes, a first group of quantization indexes indicating negative elevation angles and a second group of quantization indexes indicating positive elevation angles. 10. The apparatus of any one of the above examples, wherein the parameter quantizer (210) is configured to quantize the azimuth angles using a number of possible quantization indexes, the number of quantization indexes decreasing from lower elevation angles to higher elevation angles such that a first number of possible quantization indexes for a first elevation angle having a first magnitude is higher than a second number of possible quantization indexes for a second elevation angle having a second magnitude, the second magnitude being larger in absolute value than the first magnitude. 11. The parameter quantizer (210) determining the required accuracy from the spread value associated with the azimuth angle; quantizing the elevation angle associated with the azimuth angle with the required precision; Quantize the azimuth angle using the quantized elevation angle The device of Example 10, configured as follows: 12. The apparatus of any one of the preceding examples, wherein the quantized direction parameters have a quantized elevation angle and a quantized azimuth angle, and wherein the parameter encoder (220) is configured to first encode the quantized elevation angle and then encode the quantized azimuth angle. 13. The quantized orientation parameters include unsigned indices for azimuth and elevation angle pairs; a parameter encoder (220) configured to convert unsigned indices to signed indices such that the index representing a zero angle is located in the center of a signed interval of possible values; a parameter encoder (220) configured to perform a reordering transformation on the signed indices to interleave positive and negative numbers into unsigned numbers; Any one of the above example devices. 14. The quantized orientation parameters include sorted or unsorted unsigned azimuth and elevation indices; A parameter encoder (220) merges the paired indices into a sphere index; Perform raw coding of sphere indices 10. The apparatus of any one of the preceding examples, configured to: 15. The parameter encoder (220) is configured to derive a sphere index from the sphere offset and the current sorted or unsorted azimuth index; The sphere offset is derived from the sum of the azimuth alphabets corresponding to the sorted or unsorted elevation indexes that are less than the current sorted or unsorted elevation index; The device of Example 14. 16. A parameter encoder (220) performs entropy coding on the quantized direction parameters associated with diffusion values ​​below a threshold; Perform raw coding on the quantized direction parameters associated with diffusion values ​​greater than the threshold. 10. The apparatus of any one of the preceding examples, configured to: 17. The parameter encoder (220) is configured to dynamically determine the threshold using the quantization alphabet and the quantization of the spreading parameters, or a parameter encoder (220) configured to determine a threshold based on a quantization alphabet of the diffusion parameters; The apparatus of Example 16. 18. The parameter quantizer (210) is configured to determine, as quantized direction parameters, an elevation index, an elevation alphabet associated with the elevation index, an azimuth index, and an azimuth alphabet associated with the azimuth index; A parameter encoder (220) deriving a mean direction vector from the quantized direction vectors for the time or frequency portion of the input signal; quantizing the mean direction vector using the best angular precision of the vector for the time or frequency portion; Encode the quantized mean direction vector or an output interface (230) configured to include the coded mean direction vector in the coded parameter representation as additional side information; Any one of the above example devices. 19. Parameter Encoder (220) Using the mean direction vector, calculate a predicted elevation index and a predicted azimuth index; Calculate the signed distance between the elevation index and the predicted elevation index, and between the azimuth index and the predicted azimuth index The device of Example 18, configured as follows: 20. The apparatus of Example 19, wherein the parameter encoder (220) is configured to convert the signed distance to a reduced distance by adding a value for the small value and subtracting a value for the large value. 21. The parameter encoder (220) is configured to determine whether the quantized direction parameters are coded in a raw coding mode or an entropy coding mode; an output interface (230) configured to introduce corresponding instructions into the encoded parameter representation; Any one of the above example devices. 22. The apparatus of any one of the preceding examples, wherein the parameter encoder (220) is configured to perform entropy coding using the Golomb-Rice method or a modification thereof. 23. The parameter encoder (220) converting the components of the mean direction vector to a signed representation such that the corresponding zero value is located in the center of a signed interval of possible values; Performs a permutation transformation of signed values ​​to interleave positive and negative numbers into unsigned numbers, Encoding the result using an encoding function to obtain an encoded component of the mean direction vector; Encode the Golomb-Rice parameters using an alphabet size that corresponds to the maximum alphabet size for the corresponding component of the direction vector. 23. The apparatus of any one of Examples 18 to 22, configured to: 24. The parameter encoder (220) is configured to perform a signed distance or reduced signed distance reordering transform to interleave positive and negative numbers into unsigned numbers; a parameter encoder (220) configured to encode the reordered signed distances or the reordered reduced signed distances using the Golomb-Rice method or a modification thereof; The device of any one of Examples 19 to 23. 25. The parameter encoder (220) determining the most significant and least significant portions of the value to be coded; Computing the alphabet for the most significant part; Calculating the alphabet for the least significant part; encoding the most significant part in unary using the alphabet for the most significant part and encoding the least significant part in binary using the alphabet for the least significant part; configured to apply the Golomb-Rice method or a modification thereof using The device of Example 24. 26. The parameter encoder (220) is configured to use the determination of the most significant and least significant parts of the value to be coded, calculate the alphabet for the most significant part, and apply the Golomb-Rice method or a modification thereof; If the most significant portion of the alphabet is less than or equal to a predefined value, such as 3, then the EncodeQuasiUniform method is used to encode the entire value, and exemplary EncodeQuasiUniform methods, such as punctured codes, generate codes with only one length or codes with only two lengths, or The parameter encoder (220) encodes the least significant portion in raw coding mode using a binary code if the coding alphabet has a size that is a power of two, or encodes the least significant portion in raw coding mode using a punctured code if the coding alphabet is not a power of two. It is configured as follows: Any one of the above example devices. 27. The apparatus of any one of the preceding examples, further comprising a parameter calculator for calculating diffusion parameters using a first time or frequency resolution and for calculating directional parameters using a second time or frequency resolution, as defined in any one of claims 1 to 15 below. 28. A method for encoding directional audio coding parameters including diffusion parameters and direction parameters, comprising: quantizing the diffusion parameter and the directional parameter; encoding the quantized diffusion parameters and the quantized directional parameters; generating an encoded parametric representation comprising information about the encoded diffusion parameters and the encoded directional parameters; A method comprising: 29. A decoder for decoding an encoded audio signal comprising encoded directional audio coding parameters comprising encoded diffusion parameters and encoded direction parameters, the decoder comprising: an input interface (800) for receiving an encoded audio signal and separating the encoded diffusion parameters and the encoded directional parameters from the encoded audio signal; a parameter decoder (820) for decoding the coded diffusion parameters and coded directional parameters to obtain quantized diffusion parameters and quantized directional parameters; a parameter dequantizer (840) for determining dequantized diffusion parameters and dequantized directional parameters from the quantized diffusion parameters and the quantized directional parameters; A decoder comprising: 30. The decoder of Example 29, wherein the input interface (800) is configured to determine from a coding mode indication (806) included in the encoded audio signal whether the parameter decoder (820) will use a first decoding mode, which is a raw decoding mode, or a second decoding mode, which is a decoding mode with modeling and is different from the first decoding mode, to decode the encoded directional parameters. 31. A parameter decoder (820) configured to decode the coded spreading parameters (804) for a frame of the coded audio signal to obtain quantized spreading parameters for the frame; an inverse quantizer (840) configured to use the quantized or inverse quantized diffusion parameters to determine an inverse quantization precision for inverse quantization of at least one directional parameter for the frame; a parameter inverse quantizer (840) configured to inverse quantize the quantized direction parameters using the inverse quantization precision; Decoder for example 29 or 30. 32. The parameter decoder (820) is configured to determine a decoding alphabet for decoding coded directional parameters for a frame from the inverse quantization precision; a parameter decoder (820) configured to decode the coded directional parameters using the decoding alphabet to obtain quantized directional parameters; Decoder for example 29, 30, or 31. 33. The decoder of any one of Examples 29 to 32, wherein the parameter decoder (820) is configured to derive a quantized sphere index from the encoded orientation parameters and to decompose the quantized sphere index into a quantized elevation index and a quantized azimuth index. 34. The parameter decoder (820) Determine the elevation alphabet from the inverse quantization precision, or Determine the azimuth alphabet from quantized or dequantized elevation parameters 34. The decoder of any one of examples 29 to 33, configured to 35. The parameter decoder (820) configured to decode quantized elevation parameters from the coded direction parameters and to decode quantized azimuth parameters from the coded direction parameters; a parameter dequantizer (840) configured to determine an azimuth alphabet from the quantized elevation parameters or the dequantized elevation parameters, the size of the azimuth alphabet being larger for elevation data indicative of an elevation angle at a first absolute elevation angle compared to elevation data indicative of an elevation angle at a second absolute elevation angle, the second absolute elevation angle being larger than the first absolute elevation angle; the parameter decoder (820) is configured to use the azimuth angle alphabet to generate quantized azimuth angle parameters, or the parameter dequantizer is configured to use the azimuth angle alphabet to dequantize the quantized azimuth angle parameters; The decoder of any one of Examples 29 to 34. 36. The input interface (800) is configured to determine a decoding mode with modeling from a decoding mode indication (806) in the encoded audio signal; a parameter decoder (820) configured to obtain a mean elevation index or a mean azimuth index; The decoder of any one of Examples 29 to 35. 37. The parameter decoder (820) determines (851) an inverse quantization precision for the frame from the quantized diffusion index for the frame; determining an elevation mean alphabet or an azimuth mean alphabet from the frame-wise dequantization precision (852a); Calculating an average elevation index using the bits (808b) in the encoded audio signal and the elevation average alphabet, or calculating an average azimuth index using the bits (808a) in the encoded audio signal and the azimuth average alphabet. The decoder of example 36, configured as follows. 38. A parameter decoder (820) configured to decode certain bits (808c) in the encoded audio signal to obtain decoded elevation Golomb-Rice parameters and to decode further bits (808c) in the encoded audio signal to obtain decoded elevation distance; a parameter decoder (820) configured to decode certain bits (808a) in the encoded audio signal to obtain decoded azimuth Golomb-Rice parameters and to decode further bits (808f) in the encoded audio signal to obtain decoded azimuth distances; a parameter decoder (820) configured to calculate a quantized elevation parameter from the elevation Golomb-Rice parameter, the decoded elevation distance, and the elevation average index, or to calculate a quantized azimuth parameter from the azimuth Golomb-Rice parameter, the decoded azimuth distance, and the azimuth average index; Decoder for example 36 or 37. 39. The parameter decoder (820) is configured to decode (850) the spreading parameters for the time and frequency parts from the encoded audio signal to obtain quantized spreading parameters; a parameter inverse quantizer (840) configured to determine (851) an inverse quantization precision from the quantized or inverse quantized diffusion parameters; a parameter decoder (820) configured to derive (852a) an elevation alphabet from the inverse quantization precision and use the elevation alphabet to obtain quantized elevation parameters for the time and frequency portions of the frame; an inverse quantizer configured to inverse quantize the quantized elevation angle parameter using the elevation angle alphabet to obtain a dequantized elevation angle parameter for the time and frequency portion of the frame; The decoder of any one of Examples 29 to 38. 40. A parameter decoder (820) configured to decode the encoded orientation parameters to obtain quantized elevation parameters; a parameter dequantizer (840) configured to determine (852c) an azimuth alphabet from the quantized elevation parameters or the dequantized elevation parameters; the parameter decoder (820) is configured to calculate (852d) quantized azimuth angle parameters using the azimuth angle alphabet, or the parameter dequantizer (840) is configured to dequantize the quantized azimuth angle parameters using the azimuth angle alphabet; The decoder of any one of Examples 29 to 39. 41. The parameter inverse quantizer (840) determining the elevation alphabet using the inverse quantization precision (852a); Determine the azimuth alphabet using the dequantization precision and the quantized or dequantized elevation parameters generated using the elevation alphabet (852c). It is configured as follows: the parameter decoder (820) is configured to use the elevation alphabet to decode the coded direction parameters to obtain quantized elevation parameters and to use the azimuth alphabet to decode the coded direction parameters to obtain quantized azimuth parameters, or the parameter dequantizer (840) is configured to dequantize the quantized elevation parameters using the elevation alphabet and to dequantize the quantized azimuth parameters using the azimuth alphabet; The decoder of any one of Examples 29 to 40. 42. The parameter decoder (820) calculating a predicted elevation index or a predicted azimuth index using the mean elevation index or the mean azimuth index; performing a Golomb-Rice decoding operation or a modification thereof to obtain distances related to azimuth or elevation parameters; Add the distance related to the azimuth or elevation parameter to the mean elevation or azimuth index to obtain the quantized elevation or quantized azimuth index. The decoder of example 33, configured as follows: 43. A parameter resolution converter (710) for converting the time / frequency resolution of the dequantized diffusion parameters, or the time or frequency resolution of the dequantized azimuth or elevation parameters, or a parametric representation derived from the dequantized azimuth or elevation parameters to a target time or frequency resolution; an audio renderer (420) for applying the diffusion parameters and directional parameters to the audio signal at a target time or frequency resolution to obtain a decoded multi-channel audio signal; 43. The decoder of any one of Examples 29 to 42, further comprising: 44. The decoder of Example 43, comprising a spectrum-to-time converter (440) for converting the multi-channel audio signal from a spectral domain representation to a time domain representation having a time resolution higher than the time resolution of the target time or frequency resolution. 45. The encoded audio signal includes an encoded transport signal, and the input interface (800) is configured to extract the encoded transport signal; the decoder comprises a transport signal audio decoder (340) for decoding the encoded transport signal; the decoder further comprising a time-to-spectral converter (430) for converting the decoded transport signal into a spectral representation; the decoder comprises an audio renderer (420, 440) for rendering a multi-channel audio signal using the dequantized diffusion parameters and the dequantized directional parameters; the decoder further comprises a spectral-to-time converter (440) for converting the rendered audio signal into a time-domain representation; The decoder of any one of Examples 29 to 44. 46. ​​A method for decoding an encoded audio signal comprising encoded directional audio coding parameters comprising encoded diffusion parameters and encoded direction parameters, the method comprising: receiving an encoded audio signal and separating (800) encoded diffusion parameters and encoded directional parameters from the encoded audio signal; - decoding (820) the coded diffusion parameters and the coded directional parameters to obtain quantized diffusion parameters and quantized directional parameters; determining (840) dequantized diffusion parameters and dequantized directional parameters from the quantized diffusion parameters and the quantized directional parameters; A method comprising: 47. A computer program for performing the method of example 28 or 46 when running on a computer or processor.

[0184] The inventive encoded audio signal including the parameter representation may be stored on a digital or non-transitory storage medium or transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0185] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or apparatus corresponds to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.

[0186] Depending on particular implementation requirements, embodiments of the present invention may be implemented in hardware or software. Implementations may be performed using digital storage media, such as floppy disks, DVDs, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memories, on which electronically readable control signals are stored that cooperate (or can cooperate) with a programmable computer system to perform the respective methods.

[0187] Some embodiments according to the invention comprise a data carrier having electronically readable control signals that can cooperate with a programmable computer system to cause one of the methods described herein to be performed.

[0188] In general, embodiments of the present invention may be implemented as a computer program product having program code that operates to perform one of the methods when the computer program product is run on a computer, and that may for example be stored on a machine-readable carrier.

[0189] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier or a non-transitory storage medium.

[0190] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0191] A further embodiment of the inventive method is, therefore, a data carrier (or digital storage medium, or computer readable medium) comprising a computer program recorded thereon for performing one of the methods described herein.

[0192] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or the sequence of signals may be adapted to be transferred via a data communication connection, for example via the Internet.

[0193] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0194] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0195] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0196] The above-described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore the intention to be limited only by the scope of the appended claims, and not by the specific details presented by the description and illustration of the embodiments herein. [Explanation of symbols]

[0197] 100 DirAC analyzer, parameter calculator 110 Diffusivity Estimator 120 Direction Estimator 130 Filter Bank 140 Beamformer / Signal Selector 150 Audio Encoder 170 encoded audio signals 200 Spatial Metadata Encoder, Quantizer and Encoder Processor 210 Parameter Quantizer 210a Quantized Elevation Function 210b Quantized azimuth angle function 210c Inverse quantized elevation and azimuth functions 220 Parameter Encoder 230 Output Interface 241 Low Diffusion Coding Mode 242 High Diffusion Coding Mode 243 Mixed Diffusive Coding Mode 248 conversion 260 Resolution Indication, First Coding Mode 274 elevation average bits 300 Spatial Metadata Decoder, Parameter Processor 340 Audio Decoder, Audio Transport Signal Decoder 400 DirAC synthesizer, synthesis filter bank 420 output synthesizer, audio renderer 430 Analysis filter bank, filter bank analysis, time-spectrum converter, time-frequency converter 440 synthesis filter bank, filter bank synthesis, spectrum / time converter, audio renderer, 450 multi-channel signals, time-domain multi-channel audio signals 710 Parameter resolution converter, resolution converter 800 input interfaces 802 mode bits 804 Payload Data 806 mode bits, mode flags 808 items 808a Items, Bits 808b Items, Bits 808c Golomb-Rice azimuth parameters, bits 808d Golomb-Rice elevation parameters 808e encoded elevation distance 808f coded azimuth distance 808g Item 820 Parameter Decoder 820a Projection calculation operation 820b Sorting procedure for elevation 820c Addition operation 820d Elevation Block 820e Projection calculation operation Sorting procedure for 820f azimuth angles 820g Addition operation 820h Azimuth angle block 840 parameter inverse quantizer, complete inverse quantization operation

Claims

1. 1. An apparatus for encoding directional audio coding parameters including diffusion parameters and direction parameters, comprising: a parameter calculator (100) for calculating the diffusion parameters using a first time or frequency resolution and for calculating the directional parameters using a second time or frequency resolution; a quantizer and encoder processor (200) for generating quantized and coded representations of said diffusion parameters and said directional parameters; Equipped with The parameter calculator (100) obtaining initial diffusion parameters having a third time or frequency resolution; and obtaining initial direction parameters having a fourth time or frequency resolution; grouping and averaging the initial diffusion parameters to obtain the diffusion parameters such that the third time or frequency resolution is higher than the first time or frequency resolution, and grouping and averaging the initial direction parameters to obtain the direction parameters such that the fourth time or frequency resolution is higher than the second time or frequency resolution. The apparatus is configured to:

2. 2. The apparatus of claim 1, wherein the parameter calculator is configured to calculate the diffusion parameter and the directional parameter for a set of frequency bands, wherein bands with lower center frequencies are narrower than bands with higher center frequencies.

3. The apparatus of claim 1 , wherein the third time or frequency resolution and the fourth time or frequency resolution are equal to each other.

4. the third time or frequency resolution is a constant time or frequency resolution such that each initial spreading parameter is associated with a time slot or frequency bin having the same size; or the fourth time or frequency resolution is a constant time or frequency resolution such that each initial direction parameter is associated with a time slot or frequency bin having the same size; the parameter calculator (100) is configured to average over a first plurality of spreading parameters associated with a first plurality of time slots; or the parameter calculator (100) is configured to average over a second plurality of spreading parameters associated with a second plurality of frequency bins; or the parameter calculator (100) is configured to average over a third plurality of directional parameters associated with a third plurality of time slots; or 4. The apparatus of claim 1 or 3, wherein the parameter calculator (100) is configured to average over a fourth plurality of directional parameters associated with a fourth plurality of frequency bins.

5. 5. The apparatus of claim 1, wherein the parameter calculator is configured to perform averaging using a weighted average, in which diffusion or directional parameters derived from input signal portions having higher amplitude-related measures are weighted using a higher weighting factor compared to diffusion or directional parameters derived from input signal portions having lower amplitude-related measures.

6. 6. The apparatus of claim 5, wherein the amplitude-related measure is power or energy in a time portion or a frequency portion, or power or energy in the time portion or the frequency portion raised to a non-negative real number equal to or different from 1.

7. 2. The apparatus of claim 1, wherein the parameter calculator is configured to perform the averaging such that the diffusion parameter or the directional parameter is normalized to an amplitude-related measure derived from a time portion of the input signal corresponding to the first time or frequency resolution or the second time or frequency resolution.

8. 2. The apparatus of claim 1, wherein the parameter calculator is configured to group and average the initial directional parameters using a weighted average, wherein a first directional parameter associated with a first time portion having a first diffusion parameter indicative of lower diffusion is weighted more heavily than a second directional parameter associated with a second time portion having a second diffusion parameter indicative of higher diffusion.

9. the parameter calculator (100) is configured to calculate the initial direction parameters such that each of the initial direction parameters comprises a Cartesian vector having a component for each of two or three directions; 2. The apparatus of claim 1, wherein the parameter calculator (100) is configured to perform the averaging separately for each component of the Cartesian vector, or the components are normalized such that the sum of the squared components of the Cartesian vector with respect to a direction parameter is equal to 1.

10. further comprising a time-frequency decomposer for decomposing an input signal having a plurality of input channels into a time-frequency representation for each input channel; or 10. The apparatus of claim 9, further comprising a time-frequency decomposer for decomposing the input signal having multiple input channels into a time-frequency representation for each input channel having the third time or frequency resolution or the fourth time or frequency resolution.

11. 11. The apparatus of claim 10, wherein the time-frequency decomposer comprises a modulated filter bank resulting in complex values ​​for each subband signal, each subband signal having multiple time slots per frame and frequency band.

12. 12. The apparatus of claim 1, wherein the apparatus is configured to associate an indication of the first time or frequency resolution or the second time or frequency resolution with the quantized and coded representation for transmission to a decoder or storage.

13. 13. The apparatus of claim 1, wherein the quantizer and encoder processor (200) for generating quantized and coded representations of the diffusion parameters and the directional parameters comprises a parameter quantizer for quantizing the diffusion parameters and the directional parameters, and a parameter encoder for encoding the quantized diffusion parameters and the quantized directional parameters.

14. 1. A method for encoding directional audio coding parameters including diffusion parameters and directional parameters, comprising: calculating the diffusion parameters using a first time or frequency resolution and the directional parameters using a second time or frequency resolution; generating a quantized and coded representation of said diffusion parameters and said directional parameters; Including, The calculating step obtaining initial diffusion parameters with a third time or frequency resolution and obtaining initial direction parameters with a fourth time or frequency resolution; grouping and averaging the initial diffusion parameters to obtain the diffusion parameters such that the third time or frequency resolution is higher than the first time or frequency resolution, and grouping and averaging the initial direction parameters to obtain the direction parameters such that the fourth time or frequency resolution is higher than the second time or frequency resolution. A method comprising:

15. 1. A decoder for decoding an encoded audio signal comprising directional audio coding parameters comprising encoded diffusion parameters and encoded direction parameters, the decoder comprising: a parameter processor (300) for decoding the encoded directional audio coding parameters to obtain decoded diffusion parameters having a first time or frequency resolution and decoded directional parameters having a second time or frequency resolution, the second time or frequency resolution being different from the first time or frequency resolution; a parameter resolution converter (710) for converting the coded or decoded spreading parameters or the coded or decoded directional parameters into transformed spreading parameters or transformed directional parameters having a third time or frequency resolution, wherein the third time or frequency resolution is higher than the first time or frequency resolution, or higher than the second time or frequency resolution, or higher than both the first time or frequency resolution and the second time or frequency resolution; A decoder comprising:

16. an audio renderer (420) operating in the spectral domain; 16. The decoder of claim 15, wherein the spectral region includes, for a frame, the first number of time slots and the second number of frequency bands, such that the frame includes a number of time / frequency bins equal to a product of the first number and a second number, the first number and the second number defining the third time or frequency resolution.

17. an audio renderer (420) operating in the spectral domain; 17. A decoder as claimed in claim 15 or 16, wherein the spectral region includes, for one frame, the first number of time slots and the second number of frequency bands, such that the frame includes a number of time / frequency bins equal to the product of the first number and the second number, and wherein the first number and the second number define a fourth time or frequency resolution, the fourth time or frequency resolution being equal to the third time or frequency resolution.

18. the first time or frequency resolution is lower than the second time or frequency resolution; 18. The decoder of claim 15, wherein the parameter resolution converter (710) is configured to generate a first plurality of transformed diffusion parameters from decoded diffusion parameters and a second plurality of transformed directional parameters from decoded directional parameters, the second plurality being larger than the first plurality.

19. the encoded audio signal comprises an encoded audio transport signal; The decoder: an audio decoder (340) for decoding the encoded transport audio signal to obtain a decoded audio signal; a time / frequency converter (430) for converting the decoded audio signal into a frequency representation having the third time or frequency resolution; 19. A decoder according to any one of claims 15 to 18, comprising:

20. an audio renderer (420) for applying the transformed diffuseness parameters and the transformed directional parameters to the frequency representation of the decoded audio signal at the third time or frequency resolution to obtain a synthetic spectral representation; a spectrum-to-time converter (440) for converting the synthesized spectral representation at the third time or frequency resolution to obtain a synthesized time-domain spatial audio signal; 20. The decoder of claim 19, further comprising:

21. 21. The decoder of claim 15, wherein the parameter resolution converter (710) is configured to copy decoded directional parameters, or copy decoded diffusion parameters, or smooth or low-pass filter a set of copied directional parameters or a set of copied diffusion parameters.

22. 22. A decoder according to any one of claims 15 to 21, wherein the second time or frequency resolution is different from the first time or frequency resolution.

23. 23. The decoder of claim 15, wherein the first temporal resolution is lower than the second temporal resolution, or the second frequency resolution is higher than the first frequency resolution, or the first temporal resolution is lower than the second temporal resolution and the first frequency resolution is equal to the second frequency resolution.

24. 24. The decoder of claim 15, wherein the parameter resolution converter (710) is configured to copy the decoded diffusion parameters and the decoded directional parameters into a corresponding number of frequency-adjacent transformed parameters for a set of bands, wherein bands with lower center frequencies receive fewer copied parameters than bands with higher center frequencies.

25. the parameter processor (300) is configured to decode coded spreading parameters for a frame of the coded audio signal to obtain quantized spreading parameters for the frame; the parameter processor (300) is configured to use the quantized or dequantized diffusion parameters to determine an inverse quantization precision for inverse quantization of at least one directional parameter for the frame; 25. The decoder of any one of claims 15 to 24, wherein the parameter processor (300) is configured to dequantize quantized directional parameters using the dequantization precision.

26. the parameter processor (300) is configured to determine a decoding alphabet for decoding coded directional parameters for a frame from an inverse quantization precision to be used by the parameter processor (300) for inverse quantization; 16. The decoder of claim 15, wherein the parameter processor (300) is configured to use the determined decoding alphabet to decode the coded directional parameters and determine dequantized directional parameters.

27. the parameter processor (300) is configured to determine an elevation alphabet for processing the encoded elevation parameters from an inverse quantization precision to be used by the parameter processor (300) to inverse quantize the direction parameters, and to determine an azimuth alphabet from an elevation index obtained using the elevation alphabet; 27. A decoder according to any one of claims 15 to 26, wherein the parameter processor (300) is configured to dequantize encoded azimuth parameters using the azimuth alphabet.

28. 1. A method for decoding an encoded audio signal comprising directional audio coding parameters comprising an encoded diffusion parameter and an encoded directional parameter, the method comprising: a step (300) of decoding the encoded directional audio coding parameters to obtain decoded diffusion parameters having a first time or frequency resolution and decoded directional parameters having a second time or frequency resolution, the second time or frequency resolution being different from the first time or frequency resolution; a step (710) of transforming the coded or decoded spreading parameters or the coded or decoded directional parameters into transformed spreading parameters or transformed directional parameters having a third time or frequency resolution, wherein the third time or frequency resolution is higher than the first time or frequency resolution, or higher than the second time or frequency resolution, or higher than both the first time or frequency resolution and the second time or frequency resolution; A method comprising:

29. A computer program which, when run on a computer or processor, performs the method of claim 14.

30. 29. A computer program which, when run on a computer or processor, performs the method of claim 28.

Citation Information

Patent Citations

  • Signal analyzing method and signal composing method for complex index modulation filter bank, and program therefor and recording medium therefor

    JP2005148274A

  • JPP7175979B