Device and method for encoding multiple audio objects, or device and method for decoding using two or more related audio objects

The parametric encoding of audio objects using directional information and covariance synthesis addresses the challenge of high bit consumption and signal degradation, enabling efficient and high-quality audio reproduction at low bitrates.

JP2025170289APending Publication Date: 2025-11-18FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025134370
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-07
Filing Date
2025-08-12
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing audio encoding technologies face challenges in efficiently encoding multiple audio objects at low bitrates while maintaining high quality, particularly when the number of objects increases, leading to significant bit consumption and signal degradation.

Method used

A parametric encoding approach is adopted where at least two associated audio objects are defined per frequency bin, utilizing directional information and covariance synthesis to downmix objects into transport channels, with enhanced covariance combining and upmixing to improve audio quality and efficiency.

Benefits of technology

This method achieves high-quality audio reproduction with improved scalability and reduced bit rate by using directional cues and covariance synthesis, effectively handling multiple objects without introducing artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025170289000001_ABST
    Figure 2025170289000001_ABST
Patent Text Reader

Abstract

To provide an improved concept for encoding multiple audio objects or decoding encoded audio signals.SOLUTION: There is provided a device for encoding multiple audio objects that comprises: an object parameter calculator (100), which is the object parameter calculator (100) configured to calculate parameter data for at least two related audio objects for one or more frequency bins associated with a time frame, in which the number of the at least two related audio objects is less than the total number of the multiple audio objects; and an output interface (200) for outputting an encoded audio signal containing information about the parameter data of the at least two related audio objects for the one or more frequency bins.SELECTED DRAWING: Figure 1a
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to encoding audio signals, for example audio objects, and decoding encoded audio signals, such as encoded audio objects. [Background technology]

[0002] Introduction This document describes a parametric approach for low-bitrate encoding and decoding of object-based audio content using Directional Audio Coding (DirAC). The presented embodiment operates as part of the 3GPP® Immersive Voice and Audio Services (IVAS) codec, in which it provides an advantageous low-bitrate alternative to the Independent Stream with Metadata (ISM) mode, a discrete encoding approach.

[0003] prior art Discrete coding of objects The simplest way to code object-based audio content is to code each object individually and transmit it along with the corresponding metadata. The main drawback of this approach is that as the number of objects increases, the bit consumption required to encode each object increases significantly. A simple solution to this problem is to use a "parametric approach," in which several relevant parameters are calculated from the input signal, quantized, and transmitted along with an appropriate downmix signal that combines multiple object waveforms.

[0004] Spatial Audio Object Coding (SAOC) Spatial Audio Object Coding [SAOC_STD, SAOC_AES] is a parametric approach in which the encoder computes a downmix signal based on a downmix matrix D and a set of parameters, and sends both to the decoder. The parameters describe the psychoacoustically relevant properties and relationships of all the individual objects. At the decoder, the downmix is ​​rendered to a specific speaker layout using a rendering matrix R.

[0005] The main parameter of SAOC is the object covariance matrix E of size NxN, where N is the number of objects. This parameter is forwarded to the decoder as object level differences (OLD) and optional inter-object covariance (IOC).

[0006] Individual elements e of matrix E i,j is given by the following equation:

[0007]

number

[0008] The object level difference (OLD) is defined as follows:

[0009]

number

[0010] During the ceremony,

[0011]

number

[0012] and the absolute matter energy (NRG) is written as:

[0013]

number

[0014] and

[0015]

number

[0016] where i and j are object x i and x j where n denotes a time index, k denotes a frequency index, l denotes a series of time indexes, m denotes a series of frequency indexes, and ε is an additional constant to avoid division by zero, e.g., ε = 10.

[0017] The similarity of input objects (IOC) is given, for example, by cross-correlation.

[0018]

number

[0019] The downmix matrix D of size N_dmx rows and N columns has elements d i,j where i refers to the channel index of the downmix signal and j refers to the object index. For stereo downmix (N_dmx = 2), d i,j is calculated from the parameters DMG and DCLD as follows:

[0020]

number

[0021] During the ceremony, DMG i and DCLD i is given by the following formula:

[0022]

number

[0023] For mono downmix (N_dmx = 1), d i,j is calculated from only the DMG parameters as follows:

[0024]

number

[0025] During the ceremony,

[0026]

number

[0027] is.

[0028] Spatial Audio Object Coding-3D (SAOC-3D) Spatial Audio Object Coding for 3D Audio Reproduction (SAOC-3D) [MPEGH_AES, MPEGH_IEEE, MPEGH_STD, SAOC_3D_PAT] is an extension of the MPEG SAOC technology described above, compressing and rendering both the channel signal and the object signal in a very bitrate-efficient manner.

[0029] The main differences from SAOC are: While the original SAOC only supports up to two downmix channels, SAOC-3D can map multi-object input to any number of downmix channels (and associated side information). Rendering to multi-channel output is done directly as opposed to traditional SAOCs which used MPEG Surround as the multi-channel output processor. · Some tools, such as the remaining coding tools, have been removed.

[0030] Despite these differences, SAOC-3D is identical to SAOC from a parameter perspective: an SAOC-3D decoder receives the multichannel downmix X, the covariance matrix E, the rendering matrix R, and the downmix matrix D, just like an SAOC decoder.

[0031] The rendering matrix R is defined by the input channel and input object, received from the format converter (channel) and object renderer (object), respectively.

[0032] The downmix matrix D has elements d i,j where i refers to the channel index of the downmix signal and j refers to the object index, and is calculated from the downmix gain (DMG).

[0033]

number

[0034] During the ceremony,

[0035]

number

[0036] is.

[0037] The output covariance matrix C of size N_out * N_out is defined as: C=RER*

[0038] Related schemes There are several other schemes that are essentially similar to the SAOC described above, but with slight variations. Binaural Cue Coding of Objects (BCC), described in [BCC2001] and elsewhere, is a precursor to the SAOC technique. Joint Object Coding (JOC) and Advanced Joint Object Coding (A-JOC) perform a similar function to SAOC, delivering coarsely separated objects at the decoder side without rendering to a specific output speaker layout [JOC_AES, AC4_AES]. This technique sends the elements of the upmix matrix from the downmix to the separated objects as parameters (instead of OLD).

[0039] Directional Audio Coding (DirAC) Another parametric approach is directional audio coding. DirAC [Pulkki2009] is a perceptually motivated spatial sound reproduction. It assumes that the spatial resolution of the human auditory system is limited to decoding one cue of direction and another cue of interaural coherence for one critical band at a time.

[0040] Based on these assumptions, DirAC represents spatial sound in a frequency band by crossfading two streams: an omnidirectional diffuse stream and a directional non-diffuse stream. DirAC processing is performed in two phases: analysis and synthesis, as shown in Figures 12a and 12b.

[0041] In the DirAC analysis stage, the B-format primary matching microphone is considered as the input and the sound diffusion and direction of arrival are analyzed in the frequency domain.

[0042] In the DirAC synthesis stage, the sound is split into two streams: a non-diffuse stream and a diffuse stream. The non-diffuse stream is reproduced as a point source using amplitude panning, which can be done using Vector-Based Amplitude Panning (VBAP) [Pulkki1997]. The diffuse stream, which is responsible for the sense of envelopment, is created by delivering mutually decorrelated signals to the loudspeakers.

[0043] The analysis stage of FIG. 12a includes a bandpass filter 1000, an energy estimator 1001, an intensity estimator 1002, time averaging elements 999a and 999b, a diffuseness calculator 1003, and a direction calculator 1004. The calculated spatial parameters are a diffuseness value between 0 and 1 for each time / frequency tile and a direction of arrival parameter for each time / frequency tile generated by block 1004. In FIG. 12a, the direction parameters include azimuth and elevation angles and indicate the direction of sound arrival relative to a reference or listening position, particularly relative to the position of the microphone where the four component signals input to bandpass filter 1000 are collected. These component signals are first-order Ambisonics components, which in the diagram of FIG. 12a include an omnidirectional component W, a directional component X, another directional component Y, and another directional component Z.

[0044] The DirAC synthesis stage shown in FIG. 12b includes a bandpass filter 1005 that generates time / frequency representations of the microphone signals W, X, Y, and Z in B format. The signals corresponding to the individual time / frequency tiles are input to a virtual microphone stage 1006, which generates a virtual microphone signal for each channel. In particular, to generate the virtual microphone signal, for example, for the center channel, the virtual microphone is directed toward the center channel, and the resulting signal is the corresponding component signal of the center channel. The signal is then processed via a direct signal branch 1015 and a diffuse signal branch 1014. Both branches include corresponding gain adjusters or amplifiers controlled by the diffuseness values ​​derived from the original diffuseness parameters in blocks 1007 and 1008 and further processed in blocks 1009 and 1010 to obtain specific microphone compensation.

[0045] The component signals of the direct signal branch 1015 are also gain adjusted using gain parameters derived from directional parameters consisting of azimuth and elevation angles. In particular, these angles are input into a VBAP (vector base amplitude panning) gain table 1011. The results are input into a loudspeaker gain averaging stage 1012 for each channel and a further normalizer 1013, and the resulting gain parameters are forwarded to the amplifier or gain adjuster of the direct signal branch 1015. The diffused signal produced at the output of the decorrelator 1016 and the direct signal or non-diffused stream are combined in a combiner 1017, after which the other subbands are summed in another combiner 1018, which can be, for example, a synthesis filter bank. Thus, a loudspeaker signal for a particular loudspeaker is generated, and the same procedure is performed for the other channels of the other loudspeakers 1019 in the particular loudspeaker setup.

[0046] A high-quality version of DirAC synthesis is shown in Figure 12b. Here, the synthesizer receives all B-format signals, from which virtual microphone signals are calculated for each loudspeaker direction. The directivity pattern used is typically dipole. The virtual microphone signals are then modified in a nonlinear manner depending on the metadata, as described for branches 1016 and 1015. A low-bitrate version of DirAC is not shown in Figure 12b. However, in this low-bitrate version, only one channel of audio is transmitted. The processing difference is that all virtual microphone signals are replaced with this single channel of received audio. The virtual microphone signals are split into two streams: a diffuse stream and a non-diffuse stream, which are processed separately. Vector-based amplitude panning (VBAP) is used to reproduce non-diffuse sounds as point sources. In panning, a monophonic sound signal is applied to a subset of loudspeakers after being multiplied by a loudspeaker-specific gain factor. The gain factor is calculated using information about the loudspeaker setup and the specified pan direction. In the low-bitrate version, the input signal is simply panned in the direction implied by the metadata. In the high-quality version, each virtual microphone signal is multiplied by a corresponding gain factor, which achieves the same effect as panning but is less likely to introduce non-linear artifacts.

[0047] The goal of diffuse sound synthesis is to create the perception of sound surrounding the listener. In low-bitrate versions, the diffuse stream is reproduced by decorrelating the input signal and playing it from all loudspeakers. In high-quality versions, the virtual microphone signal of the diffuse stream already has a certain degree of incoherence and needs to be gently decorrelated.

[0048] DirAC parameters, also known as spatial metadata, consist of a tuple of diffusivity and direction, expressed in spherical coordinates as two angles: azimuth and elevation. If both the analysis and synthesis stages are performed at the decoder, the time-frequency resolution of the DirAC parameters can be chosen to be the same as the filterbank used for DirAC analysis and synthesis, i.e., different parameter sets can be chosen for each time slot and frequency bin of the filterbank representation of the audio signal.

[0049] To make the DirAC paradigm usable in spatial audio coding and teleconferencing scenarios, some work has been done to reduce the size of the metadata [Hirvonen2009].

[0050] [WO2019068638] introduced a universal spatial audio coding system based on DirAC. In contrast to conventional DirAC, which is designed for B-format (first-order Ambisonics format) input, this system can accept first-order or higher Ambisonics, multi-channel, or object-based audio input, and also allows for mixed-type input signals. All signal types can be efficiently coded and transmitted individually or in combination. While the former combines different representations in the renderer (decoder side), the latter uses an encoder-side combination of different audio representations in the DirAC domain.

[0051] Compatible with DirAC framework This embodiment builds on the unified framework for any input type presented in [WO2019068638] and aims to overcome the problem of not being able to efficiently apply DirAC parameters (direction and diffuseness) to object input (similar to what [WO2020249815] does for multi-channel content). It turns out that a single directional cue per time / frequency unit is insufficient to reproduce high-quality object content, even though diffuseness parameters are not actually necessary. Therefore, this embodiment proposes using multiple directional cues per time / frequency unit, thus introducing an adapted parameter set that replaces conventional DirAC parameters in the case of object input.

[0052] Low bitrate flexible system In contrast to DirAC, which uses a scene-based representation from the listener's perspective, SAOC and SAOC-3D are designed for channel- and object-based content, with parameters describing the relationships between channels / objects. To use a scene-based representation for object input and be compatible with DirAC renderers while ensuring efficient representation and high-quality reproduction, an adapted parameter set is required to allow signaling of multiple directional cues.

[0053] An important goal of this embodiment was to find a way to efficiently code object inputs at low bit rates and with good scalability to increasing numbers of objects. Coding each object signal individually would not provide such scalability. Each additional object significantly increases the overall bit rate. If the number of objects increases beyond the allowed bit rate, this directly leads to significant degradation of the output signal. This degradation is yet another argument in favor of this embodiment. [Prior art documents] [Patent documents]

[0054] [Patent Document 1] WO2019068638 [Patent Document 2] WO2020249815 Summary of the Invention [Problem to be solved by the invention]

[0055] It is an object of the present invention to provide an improved concept for encoding multiple audio objects or for decoding an encoded audio signal.

[0056] This object is achieved by an encoding device according to claim 1, a decoder according to claim 18, an encoding method according to claim 28, a decoding method according to claim 29, a computer program according to claim 30 or an encoded audio signal according to claim 31. [Means for solving the problem]

[0057] In one aspect, the invention is based on the discovery that for one or more frequency bins of a plurality of frequency bins, at least two associated audio objects are defined, and parameter data relating to these at least two associated objects is included on the encoder side and used on the decoder side to obtain a high quality and efficient audio encoding / decoding concept.

[0058] According to a further aspect of the invention, the invention is based on the discovery that a specific downmix adapted to the directional information associated with each object is performed such that each object with associated directional information is valid for the entire object, i.e. is used to downmix this object into a number of transport channels for all frequency bins within a time window. The use of directional information is for example equivalent to generating the transport channels as virtual microphone signals with specific adjustable characteristics.

[0059] On the decoder side, in certain embodiments, specific synthesis is performed that relies on covariance synthesis, which is particularly suited for high-quality covariance synthesis that does not suffer from artifacts introduced by the decorrelator. In other embodiments, advanced covariance synthesis is used that relies on specific improvements relative to standard covariance synthesis to improve speech quality and / or reduce the amount of computation required to calculate the mixing matrices used in covariance synthesis.

[0060] However, even in more classical synthesis, where audio rendering is performed by explicitly determining individual contributions within a time / frequency bin based on transmitted selection information, audio quality is superior to prior art object coding or channel downmix approaches. In such situations, each time / frequency bin has object identification information, and when performing audio rendering, i.e., when considering the directional contribution of each object, this object identification is used to look up the direction associated with this object information to determine the gain values ​​of the individual output channels for each time / frequency bin. Thus, if there is only one object associated with a time / frequency bin, then only the gain value of this single object per time / frequency bin is determined based on a "codebook" of object IDs and associated object directional information.

[0061] However, if there are multiple associated objects in a time / frequency bin, then a gain value for each associated object is calculated and the corresponding time / frequency bins of the transport channels are distributed to the corresponding output channels governed by the output format provided by the user, such as a particular channel format being a stereo format, a 5.1 format, etc. Regardless of whether the gain values ​​are used for the purposes of covariance combining, i.e., applying a mixing matrix for mixing the transport channels into output channels, or whether the gain values ​​are used to explicitly determine the individual contribution of each object in a time / frequency bin by multiplying the gain values ​​by the corresponding time / frequency bins of one or more transport channels and then summing the contributions of each output channel for the corresponding time / frequency bins, possibly enhanced by the addition of diffuse signal components, the output audio quality is nevertheless improved due to the flexibility afforded by determining one or more associated objects per frequency bin.

[0062] This determination is possible very efficiently because only one or more object IDs for a time / frequency bin need to be encoded and transmitted to the decoder together with per-object orientation information, which is also very efficient. This is due to the fact that for a frame there is only a single orientation information for all frequency bins.

[0063] Thus, regardless of whether the synthesis is preferably performed using enhanced covariance synthesis or using a combination of explicit transport channel contributions for each object, a highly efficient and high quality object downmix is ​​obtained, which is preferably enhanced by using a specific object direction dependent downmix that relies on downmix weights that reflect the generation of the transport channels as virtual microphone signals.

[0064] The aspect relating to two or more associated objects per time / frequency bin can preferably be combined with the aspect of performing a specific direction-dependent downmix of the objects into the transport channel, although both aspects can also be applied independently of each other. Furthermore, in certain embodiments, covariance combining with two or more associated objects per time / frequency bin is performed, but advanced covariance combining and advanced upmixing from the transport channel to the output channel can also be performed by transmitting only one object ID per time / frequency bin.

[0065] Furthermore, regardless of whether there are single or multiple relevant objects per time / frequency bin, upmixing can also be performed by calculating a mixing matrix within standard or enhanced covariance combining, or upmixing can be performed by obtaining specific directional information from a directional "codebook" and determining the contributions of time / frequency bins individually based on object identification, which is used to determine gain values ​​for the corresponding contributions. These are summed to obtain the complete contributions per time / frequency bin when there are two or more relevant objects per time / frequency bin. The output of this summation step is equivalent to the output of the mixing matrix application, and final filter bank processing is performed to generate the time-domain output channel signal in the corresponding output format.

[0066] Preferred embodiments of the present invention are described below with reference to the accompanying drawings. [Brief explanation of the drawings]

[0067] [Figure 1a] FIG. 2 illustrates an implementation of an audio encoder according to a first aspect having at least two associated objects per time / frequency bin. [Figure 1b] FIG. 10 illustrates an implementation of an encoder according to a second aspect with downmixing of directionally dependent objects. [Figure 2] FIG. 2 shows a preferred implementation of an encoder according to the second aspect. [Figure 3] FIG. 2 shows a preferred implementation of an encoder according to the first aspect; [Figure 4] FIG. 2 shows a preferred implementation of a decoder according to the first and second aspects; [Figure 5] FIG. 5 illustrates a preferred implementation of the covariance combining process of FIG. 4. [Figure 6a] FIG. 2 illustrates an implementation of a decoder according to a first aspect. [Figure 6b] FIG. 2 shows a decoder according to a second embodiment; [Figure 7a] 1 is a flowchart showing determination of parameter information according to a first embodiment. [Figure 7b] FIG. 10 illustrates a preferred implementation of the further determination of parametric data. [Figure 8] (a) Time / frequency representation of a high-resolution filter bank, (b) Transmission of relevant side information for frame J according to a preferred implementation of the first and second aspects, and (c) "Directional Codebook" included in the encoded audio signal. [Figure 9a] FIG. 10 shows a preferred encoding method according to the second embodiment. [Figure 9b] FIG. 10 illustrates an implementation of a static downmix according to a second embodiment. [Figure 9c] FIG. 10 illustrates an implementation of dynamic downmix according to a second embodiment. [Figure 9d] FIG. 1 shows a further embodiment of the second aspect. [Figure 10a] FIG. 1 shows a flowchart for a preferred decoder-side implementation of the first embodiment. [Figure 10b] FIG. 10b illustrates a preferred implementation of the output channel calculation of FIG. 10a according to an embodiment with a sum of contributions for each output channel. [Figure 10c] FIG. 2 illustrates a preferred method for determining power values ​​according to a first aspect for multiple objects. [Figure 10d]FIG. 10b illustrates an embodiment of the calculation of the output channels of FIG. 10a using covariance combining that relies on the calculation and application of a mixing matrix. [Figure 11] 1 illustrates some embodiments of advanced calculation of mixing matrices for time / frequency bins. [Figure 12a] FIG. 1 illustrates a prior art DirAC encoder. [Figure 12b] FIG. 1 illustrates a prior art DirAC decoder. DETAILED DESCRIPTION OF THE INVENTION

[0068] FIG. 1a shows an apparatus for encoding multiple audio objects, receiving at input the raw audio objects and / or audio object metadata. The encoder includes an object parameter calculator 100 that provides parameter data for at least two related audio objects for a time / frequency bin, which is then transferred to an output interface 200. In particular, the object parameter calculator calculates parameter data for at least two related audio objects for one or more frequency bins among a plurality of frequency bins associated with a time frame, where the number of at least two related audio objects is less than the total number of the plurality of audio objects. Thus, the object parameter calculator 100 actually performs a selection and does not simply indicate that all objects are related. In a preferred embodiment, the selection is by relevance, where relevance is determined by an amplitude-related measure, such as amplitude, power, loudness, or another measure obtained by raising the amplitude to a power different from, and preferably greater than, 1. Then, if a certain number of related objects are available for a time / frequency bin, the object with the most related characteristics, i.e., the object with the highest power among all objects, is selected, and data about these selected objects is included in the parameter data.

[0069] The output interface 200 is configured to output an encoded audio signal containing information about parameter data of at least two related audio objects in one or more frequency bins. Depending on the implementation, the output interface can receive other data, such as a downmix of the objects, or one or more transport channels representing the downmix of the objects, or additional parameter or object waveform data in a mixed representation in which multiple objects are downmixed, or other objects in another representation, to input into the encoded audio signal. In this situation, the objects are directly introduced or "copied" into the corresponding transport channels.

[0070] 1b shows a preferred implementation of an apparatus for encoding multiple audio objects according to the second aspect, in which the audio objects are received together with associated object metadata indicating directional information for the multiple audio objects, i.e., one directional information per object or per group of objects if a group of objects is associated with the same directional information. The audio objects are input to a downmixer 400, which downmixes the multiple audio objects to obtain one or more transport channels. Furthermore, a transport channel encoder 300 is provided, which encodes the one or more transport channels to obtain one or more encoded transport channels that are input to the output interface 200. In particular, the downmixer 400 is connected to an object directional information provider 110, which receives at its input any data from which object metadata can be derived and outputs the directional information actually used by the downmixer 400. The directional information transferred from the object directional information provider 110 to the downmixer 400 is preferably dequantized directional information, i.e., the same directional information that is subsequently available on the decoder side. For this purpose, the object direction information provider 110 is configured to derive or extract or obtain unquantized object metadata and then quantize the object metadata to derive quantized object metadata representing quantization indices which, in a preferred embodiment, are provided to the output interface 200 among the "other data" shown in Fig. 1b. Furthermore, the object direction information provider 110 is configured to dequantize the quantized object direction information to obtain the actual direction information which is forwarded from block 110 to the downmixer 400.

[0071] Preferably, the output interface 200 is further configured to receive parameter data for the audio objects, object waveform data, one or more identifications of single or multiple associated objects per time / frequency bin, and, as described above, quantized directional data.

[0072] Further embodiments are presented below. A parametric approach for coding audio object signals is presented, enabling efficient transmission at low bit rates and high-quality playback at the consumer end. Based on the DirAC principle of considering one directional cue per important frequency band and time (time / frequency tile), a most dominant object is determined for each time / frequency tile of the time / frequency representation of the input signal. This proved insufficient for object inputs, so an additional, second most dominant object is determined for each time / frequency tile. A power ratio is calculated based on these two objects to determine the respective influence of the two objects on the considered time / frequency tile. Note: It is also conceivable to consider more than two most dominant objects per time / frequency unit, especially when the number of input objects is increasing. For simplicity, the following description is mostly based on two dominant objects per time / frequency unit.

[0073] Therefore, the parametric side information transmitted to the decoder includes: · Power ratios calculated for the subset of relevant (dominant) objects in each time / frequency tile (or parameter band). · Object indices that represent the subset of relevant objects for each time / frequency tile (or parameter band). Orientation information associated with the object index and provided for each frame (each time-domain frame contains multiple parameter bands, and each parameter band contains multiple time / frequency tiles).

[0074] The directional information is made available via an input metadata file associated with the audio object signal. The metadata may, for example, be specified on a frame-by-frame basis. Apart from the side information, a downmix signal combining the input object signals is also transmitted to the decoder.

[0075] During the rendering stage, the transmitted directional information (derived via the object index) is used to pan the transmitted downmix signal (or more generally the transport channel) in the appropriate direction. The downmix signal is distributed to the two associated object directions based on the transmit power ratio, which is used as a weighting factor. This process is performed for each time / frequency tile of the time / frequency representation of the decoded downmix signal.

[0076] This section provides an overview of the encoder-side processing, followed by a detailed description of the parameter and downmix calculations. The audio encoder receives one or more audio object signals. Each audio object signal has an associated metadata file that describes object properties. In this embodiment, the object properties described in the associated metadata file correspond to directional information provided on a frame-by-frame basis, with each frame corresponding to 20 milliseconds. Each frame is identified by a frame number, which is also included in the metadata file. The directional information is given as azimuth and elevation angle information, with the azimuth angle taking a value of [-180,180] degrees and the elevation angle taking a value of [-90,90] degrees. Other properties provided in the metadata include distance, spread, gain, etc. These characteristics are not considered in this embodiment.

[0077] The information provided in the metadata file is used together with the actual audio object files to create a set of parameters that are sent to the decoder and used to render the final audio output file. More specifically, the encoder estimates the parameters, i.e., power ratios, of a subset of dominant objects for each given time / frequency tile. The subset of dominant objects is represented by an object index that is also used to identify the object's direction. These parameters are sent to the decoder along with the transport channel and direction metadata.

[0078] An overview of the encoder is shown in FIG. 2, where the transport channels contain a downmix signal calculated from input object files and directional information provided in the input metadata. The number of transport channels is always less than the number of input object files. In one embodiment of the encoder, the encoded audio signal is represented by the encoded transport channels, and the encoded parametric side information is represented by the encoded object index, the encoded power ratio, and the encoded directional information. Together, the encoded transport channels and the encoded parametric side information form a bitstream output by the multiplexer 220. In particular, the encoder comprises a filter bank 102 that receives the input object audio files. Furthermore, the object metadata files are provided to an extractor directional information block 110a. The output of block 110a is input to a quantization directional information block 110b, which outputs directional information to the downmixer 400, which performs the downmix calculation. Furthermore, the quantized directional information, i.e., the quantization index, is transferred from block 110b to an encoding directional information block 202, which preferably performs some kind of entropy encoding to further reduce the required bitrate.

[0079] Furthermore, the output of the filter bank 102 is input to a signal power calculation block 104, the output of which is input to an object selection block 106 and further to a power ratio calculation block 108. The power ratio calculation block 108 is also connected to the object selection block 106 to calculate the power ratio, i.e., the combined value, of only the selected object. In block 210, the calculated power ratio or combined value is quantized and encoded. As outlined below, the power ratio is preferred to save the transmission of one power data item. However, in other embodiments where this savings is not necessary, the actual signal power determined by block 104 or other values ​​derived from the signal power can be input to the quantizer and encoder under the selection of the object selector 106 instead of the power ratio. Then, the power ratio calculation 108 is not required, and the object selection 106 ensures that only the relevant parametric data, i.e., the power-related data of the relevant object, is input to block 210 for quantization and encoding purposes.

[0080] Comparing FIG. 1a with FIG. 2, blocks 102, 104, 110a, 110b, 106, 108 are preferably included in object parameter calculator 100 of FIG. 1a, and blocks 202, 210, 220 are preferably included in output interface block 200 of FIG. 1a.

[0081] Furthermore, the core coder 300 of Figure 2 corresponds to the transport channel encoder 300 of Figure 1b, the downmix calculation block 400 corresponds to the downmixer 400 of Figure 1b, and the object direction information provider 110 of Figure 1b corresponds to blocks 110a and 110b of Figure 2. Furthermore, the output interface 200 of Figure 1b is preferably implemented in the same way as the output interface 200 of Figure 1a and includes blocks 202, 210, and 220 of Figure 2.

[0082] Figure 3 shows a variant of the encoder in which the calculation of the downmix is ​​optional and does not depend on the input metadata. In this variant, the input audio files are fed directly to the core coder, which creates transport channels from them. The number of transport channels therefore corresponds to the number of input object files. This is particularly interesting when the number of input objects is one or two. Even if the number of objects is large, a downmix signal is used to reduce the amount of data to be transmitted.

[0083] In Figure 3, like reference numerals refer to like functions in Figure 2. This is valid not only for Figures 2 and 3, but also for all other figures described in this specification. Unlike Figure 2, Figure 3 performs the downmix calculation 400 without directional information. Therefore, the downmix calculation can be, for example, a static downmix using a known downmix matrix, or an energy-dependent downmix that does not rely on directional information associated with the objects contained in the input object audio file. Nevertheless, the directional information is extracted in block 110a and quantized in block 110b, and the quantized values ​​are transferred to the directional information encoder 202 for the purpose of encoding the directional information in the encoded audio signal, which can be, for example, a binary encoded audio signal forming a bitstream.

[0084] If the number of input audio object files is not very large or if there is sufficient available transmission bandwidth, the downmix calculation block 400 can be omitted and the input audio object files can directly represent the transport channels encoded by the core encoder. In such an implementation, blocks 104, 104, 106, 108, and 210 are also not required. However, a preferred implementation results in a mixed implementation in which some objects are introduced directly into the transport channels and other objects are downmixed into one or more transport channels. In such a situation, all blocks shown in FIG. 3 are required to generate a bitstream with one or more objects directly in the encoded transport channels and one or more transport channels generated by the downmixer 400 of either FIG. 2 or FIG. 3.

[0085] Parameter Calculation The time-domain audio signal, including all input object signals, is transformed into the time / frequency domain using a filter bank. For example, a CLDFB (Composite Low Delay Filter Bank) analysis filter transforms a 20 ms frame (corresponding to 960 samples at a sampling rate of 48 kHz) into a time / frequency tile of size 16x60 with 16 time slots and 60 frequency bands. For each time / frequency unit, the instantaneous signal power is calculated as follows: P i (k,n)=|X i (k,n)| 2 where k is the frequency band index, n is the time slot index, and i is the object index. Since transmitting the parameters for each time / frequency tile would be very costly in terms of the final bit rate, grouping is used to calculate the parameters for a reduced number of time / frequency tiles. For example, 16 time slots can be grouped into one time slot, and 60 frequency bands can be grouped into 11 bands based on the psychoacoustic scale. This reduces the initial size of 16x60 to 1x11, which corresponds to 11 so-called parameter bands. The instantaneous signal power values ​​are summed based on the grouping to obtain the signal power of the reduced dimension.

[0086]

number

[0087] where T corresponds to 15 in this example and B S and B E defines the boundaries of the parameter bands.

[0088] The instantaneous signal power values ​​of all N input audio objects are sorted in descending order to determine a subset of the most dominant objects for which to calculate the parameters. In this embodiment, the two most dominant objects are determined, and their corresponding object indices ranging from 0 to N-1 are stored as part of the transmitted parameters. Furthermore, a power ratio correlating the two dominant object signals is calculated.

[0089]

number

[0090] Or, in more general terms that are not limited to two objects:

[0091]

number

[0092] where, in this context, S denotes the number of dominant objects considered,

[0093]

number

[0094] is.

[0095] In the case of two dominant objects, a power ratio of 0.5 for each of the two objects means that both objects are equally present in the corresponding parameter band, while power ratios of 1 and 0 represent the absence of one of the two objects. These power ratios are stored as the second part of the transmitted parameters. Since the sum of the power ratios is 1, it is sufficient to transmit the value S-1 instead of S.

[0096] In addition to the object index and the power ratio values ​​per parameter band, it is necessary to transmit the direction information of each object extracted from the input metadata file. This is done per frame, since the information is originally provided frame-wise (each frame consists of 11 parameter bands, or a total of 16x60 time / frequency tiles in the described example). Thus, the object index indirectly represents the direction of the object. Note: Since the power ratios sum to 1, the number of power ratios transmitted per parameter band can be reduced by 1. Example: When considering two related objects, it is sufficient to transmit one power ratio value.

[0097] Both the direction information and the power ratio values ​​are quantized and combined with the object index to form the parametric side information. This parametric side information is then encoded and mixed together with the encoded transport channel / downmix signal into the final bitstream representation. A good tradeoff between output quality and consumed bitrate can be achieved by quantizing the power ratios using, for example, 3 bits per value. The direction information can be provided with an angular resolution of 5 degrees and then quantized with 7 bits per azimuth value and 6 bits per elevation value, as shown in the practical example.

[0098] Downmix Calculation All input audio object signals are combined into a downmix signal that includes one or more transport channels. The number of transport channels is less than the number of input object signals. Note: In this embodiment, a single transport channel occurs only when there is only one input object, which means that the downmix calculation is skipped.

[0099] If the downmix contains two transport channels, this stereo downmix is ​​calculated, for example, as a virtual cardioid microphone signal, which is determined by applying the directional information provided for each frame in the metadata file (here, all elevation values ​​are assumed to be zero). w L =0.5+0.5*cos(azimuth-pi / 2) w R =0.5+0.5*cos(azimuth-pi / 2)

[0100] Here, the virtual cardioid is positioned at 90° and -90°, and therefore the weights for each of the two transport channels (left and right) are determined and applied to the corresponding audio object signals.

[0101]

number

[0102] In this context, N is the number of input objects, ≥ 2. If the virtual cardioid weights are updated every frame, a dynamic downmix that adapts to directional information is employed. Another possibility is to employ a fixed downmix, where each object is assumed to be in a static position. This static position may, for example, correspond to the object's initial orientation, resulting in the same static virtual cardioid weight in every frame.

[0103] More than two transport channels are possible if the target bitrate allows. With three transport channels, the cardioids are uniformly positioned, for example, at 0°, 120°, and -120°. With four transport channels, the fourth cardioid can be pointed upward, or all four cardioids can be uniformly positioned horizontally. Object placement can also be adjusted to match the object's location. The resulting downmix signal is processed by the core coder and converted to a bitstream representation, along with the encoded parametric side information.

[0104] Alternatively, the input object signals may be fed to the core coder without being combined into a downmix signal. In this case, the number of resulting transport channels corresponds to the number of input object signals. Typically, a maximum number of transport channels is specified that correlates with the total bit rate. The downmix signal is only used if the number of input object signals exceeds this maximum number of transport channels.

[0105] FIG. 6a shows a decoder for decoding an encoded audio signal, such as the signal output by FIG. 1a, 2, or 3, including one or more transport channels and directional information for multiple audio objects. Furthermore, the encoded audio signal includes parameter data for at least two associated audio objects for one or more frequency bins of a time frame, where the number of the at least two associated objects is less than the total number of audio objects. In particular, the decoder includes an input interface for providing one or more transport channels in a spectral representation having multiple frequency bins within a time frame. This represents the signal transferred from the input interface block 600 to the audio renderer block 700. In particular, the audio renderer 700 is configured to use the directional information included in the encoded audio signal to render the one or more transport channels into multiple audio channels, preferably two for a stereo output format or three or more for a larger number of output formats, such as three, five, or 5.1 channels. In particular, the audio renderer 700 is configured to calculate, for each of one or more frequency bins, contributions from one or more transport channels according to first directional information associated with a first one of the at least two associated audio objects and according to second directional information associated with a second one of the at least two associated objects, In particular, the directional information for the multiple audio objects includes first directional information associated with the first object and second directional information associated with the second object.

[0106] FIG. 8b shows parameter data for a frame, which in a preferred embodiment consists of directional information 810 for multiple audio objects, additionally including power ratios for each of a specific number of parameter bands, as shown in block 812, and one, preferably two, or more object indices for each parameter band, as shown in block 814. In particular, the directional information for multiple audio objects 810 is shown in more detail in FIG. 8c. FIG. 8c shows a table with a first column containing specific object IDs ranging from 1 to N, where N is the number of multiple audio objects. Additionally, a second column is provided containing directional information for each object, preferably as azimuth and elevation values, or, in the case of two-dimensional situations, as azimuth values ​​only. This is shown in 818. Thus, FIG. 8c shows a "directional codebook" included in the encoded audio signal input to input interface 600 of FIG. 6a. The directional information from column 818 is uniquely associated with a specific object ID from column 816 and is valid for the "whole" object in the frame, i.e., for all frequency bands within the frame. Therefore, regardless of the number of frequency bins in the time / frequency tile of the high-resolution representation or the time / parameter band of the low-resolution representation, only a single piece of directional information is transmitted and used by the input interface per object identification.

[0107] In this context, FIG. 8a shows the time / frequency representation generated by the filter bank 102 of FIG. 2 or FIG. 3 when this filter bank is implemented as the aforementioned CLDFB (Complex Low Delay Filter Bank). For a frame in which directional information is provided as previously described with respect to FIG. 8b and FIG. 8c, the filter bank generates 16 time slots, numbered 0 to 15, and 60 frequency bands, numbered 0 to 59, in FIG. 8a. Thus, one time slot and one frequency band represent a time / frequency tile 802 or 804. Nevertheless, to reduce the bit rate of the side information, it is preferable to convert the high-resolution representation to the low-resolution representation shown in FIG. 8b, in which there is only a single time bin and the 60 frequency bands are converted into 11 parameter bands, as shown at 812 in FIG. 8b. Thus, as shown in FIG. 10c, the high-resolution representation is indicated by the time slot index n and the frequency band index k, and the low-resolution representation is given by the grouped time slot index m and the parameter band index l. Nevertheless, in the context of this specification, a time / frequency bin may include high-resolution time / frequency tiles 802, 804 of Figure 8a or low-resolution time / frequency units identified by grouped time slot indices and parameter band indices at the input of block 731c of Figure 10c.

[0108] In the embodiment of Figure 6a, the audio renderer 700 is configured to calculate, for each of one or more frequency bins, contributions from one or more transport channels according to first directional information associated with a first of the at least two related audio objects and according to second directional information associated with a second of the at least two related audio objects. In the embodiment shown in Figure 8b, block 814 has an object index for each related object in the parameter band, i.e., two or more object indices such that there are two contributions per time-frequency bin.

[0109] As outlined below with respect to FIG. 10a, the calculation of contributions can be performed indirectly via a mixing matrix, in which a gain value for each associated object is determined and used to calculate the mixing matrix. Alternatively, as shown in FIG. 10b, the contributions can again be explicitly calculated using the gain values, and the explicitly calculated contributions can be summed for each output channel of a particular time / frequency bin. Thus, regardless of whether the contributions are calculated explicitly or implicitly, the audio renderer nevertheless uses directional information to render one or more transport channels into multiple audio channels. Thus, for each of one or more frequency bins, the contributions from one or more transport channels are included in the number of audio channels according to first directional information associated with a first of at least two associated audio objects and according to second directional information associated with a second of the at least two associated audio objects.

[0110] 6b illustrates a decoder for decoding an encoded audio signal including directional information for one or more transport channels and multiple audio objects and parameter data for the audio objects for one or more frequency bins of a time frame according to a second embodiment. Again, the decoder includes an input interface 600 for receiving the encoded audio signal, and an audio renderer 700 for rendering the one or more transport channels into multiple audio channels using the directional information. In particular, the audio renderer is configured to calculate, for each frequency bin of the multiple frequency bins, direct response information from one or more audio objects and directional information associated with the associated one or more audio objects within the frequency bin. This direct response information preferably includes gain values ​​used for covariance or advanced covariance synthesis, or for explicit calculation of contributions from one or more transport channels.

[0111] Preferably, the audio renderer is configured to use direct response information of one or more associated audio objects in the time / frequency band and to calculate covariance synthesis information using information about the number of audio channels. Further, the covariance synthesis information, which is preferably a mixing matrix, is applied to one or more transport channels to obtain the number of audio channels. In a further implementation, the direct response information is a direct response vector for each of the one or more audio objects, the covariance synthesis information is a covariance synthesis matrix, and the audio renderer is configured to perform a matrix operation for each frequency bin when applying the covariance synthesis information.

[0112] Furthermore, the audio renderer 700 is configured to derive direct response vectors for one or more audio objects in calculating the direct response information and calculate a covariance matrix from each direct response vector for the one or more audio objects. Furthermore, in calculating the covariance synthesis information, a target covariance matrix is ​​calculated. However, instead of the target covariance matrix, related information of the target covariance matrix can be used, i.e., the direct response matrix or vector of one or more most dominant objects and a diagonal matrix of direct powers, denoted as E, determined by applying a power ratio.

[0113] Thus, the target covariance information does not necessarily have to be an explicit target covariance matrix, but is derived from the covariance matrix of one audio object, or the covariance matrices of multiple audio objects in a time / frequency bin, from power information of each one or more audio objects in the time / frequency bin and power information derived from one or more transport channels in one or more time / frequency bins.

[0114] The bitstream representation is read by a decoder and the encoded transport channels and the encoded parametric side information contained therein are made available for further processing. The parametric side information includes: Direction information (per frame) as quantized azimuth and elevation values Object index (for each parameter band) indicating the subset of related objects Quantized power ratios (per parameter band) that correlate related objects

[0115] All processing is done frame by frame, with each frame consisting of one or more subframes. A frame may consist of, for example, four subframes, each lasting 5 milliseconds. Figure 4 shows a simplified overview of the decoder.

[0116] Figure 4 shows an audio decoder implementing the first and second aspects. The input interface 600 shown in Figures 6a and 6b comprises a demultiplexer 602, a core decoder 604, a decoder for decoding object indices 608, a decoder for decoding and dequantizing power ratios 612, and a decoder for decoding and dequantizing directional information indicated at 612. Furthermore, the input interface comprises a filter bank 606 for providing transport channels in a time / frequency representation.

[0117] The audio renderer 700 comprises a direct response calculator 704, a prototype matrix provider 702 controlled by an output configuration received by a user interface, e.g., a covariance synthesis block 706, and a synthesis filter bank 708 to ultimately provide an output audio file containing the number of audio channels in the channel output format.

[0118] Thus, items 602, 604, 606, 608, 610, 612 are preferably included in the input interface of Figures 6a and 6b, and items 702, 704, 706, 708 of Figure 4 are part of the audio renderer of Figure 6a or 6b, indicated by reference numeral 700.

[0119] The encoded parametric side information is decoded to re-obtain the quantized power ratio values, the quantized azimuth and elevation values ​​(directional information), and the object index. The one power ratio value that is not transmitted is obtained by taking the fact that all power ratio values ​​sum to 1. Their resolution (l,m) corresponds to the time / frequency tile group used on the encoder side. In further processing steps where a finer time / frequency resolution (k,n) is used, the parameters of a parameter band are valid for all time / frequency tiles contained in this parameter band, corresponding to the extension (l,m) → (k,n).

[0120] The encoded transport channel is decoded by a core decoder. Using a filter bank (matching the one employed in the encoder), each frame of the audio signal thus decoded is converted into a time / frequency representation whose resolution is typically finer than (but at least equal to) the resolution used for the parametric side information.

[0121] Output signal rendering / compositing The following description applies to one frame of the audio signal, where ^T denotes the transpose operator.

[0122] Decoded transport channel x = x(k,n) = [X1(k,n),X2(k,n)] T , i.e., using the time-frequency representation of the audio signal (consisting of two transport channels in this case) and the parametric side information, the mixing matrix M for each subframe (or frame to reduce computational complexity) is calculated as a time-frequency output signal y = y(k,n) = [Y1(k,n),Y2(k,n),Y3(k,n),…] containing several output channels (e.g., 5.1, 7.1, 7.1+4, etc.). T is derived to synthesize

[0123] For every (input) object, the transmitted object directions are used to determine so-called direct response values, which describe the panning gains to be used for the output channels. These direct response values ​​are specific to the target layout, i.e., the number and positions of loudspeakers (provided as part of the output configuration). Examples of panning methods include Vector-Based Amplitude Panning (VBAP) [Pulkki1997] and Edge-Fading Amplitude Panning (EFAP) [Borss2014]. Each object has a direct response value dr associated with it. i There are vectors (containing as many elements as there are loudspeakers). These vectors are calculated once per frame. Note: If the object's position corresponds to the position of a loudspeaker, the vector contains the value 1 for this loudspeaker, all other values ​​are 0. If the object is located between two (or three) loudspeakers, the number of corresponding non-zero vector elements is 2 (or 3).

[0124] The actual synthesis step (in this embodiment, covariance synthesis [Vilkamo2013]) involves the following substeps (see Figure 5 for visualization): For each parameter band, we use the object index describing the dominant subset of the input objects in the time / frequency tiles grouped in this parameter band to generate the vector dr needed for further processing. i For example, since only two related objects are considered, the two vectors dr associated with these two related objects are extracted. i is necessary. Next, the direct response value dr i From the dimensional output channel covariance matrix C for each output channel i is calculated for each relevant object. C i =dr i *dr i T For each time / frequency tile (within the parameter band), the audio signal power P(k,n) is determined. In the case of two transport channels, the signal power of the first channel is added to the signal power of the second channel. This signal power is multiplied by the value of each power ratio to obtain one direct power value for each relevant / dominant object i. DP i (k,n)=PR i (k,n)*P(k,n) For each frequency band k, the final target covariance matrix C of the output channels of size per output channel Y is obtained by summing over all slots n in the (sub)frame and summing over all relevant objects.

[0125]

number

[0126] Figure 5 shows a detailed overview of the covariance synthesis step performed in block 706 of Figure 4. In particular, the embodiment of Figure 5 includes a signal power calculation block 721, a direct power calculation block 722, a covariance matrix calculation block 73, a target covariance matrix calculation block 724, an input covariance matrix calculation block 726, a mixing matrix calculation block 725, and a rendering block 727, such that the output signal of block 727, which with respect to Figure 5 further includes filter bank block 708 of Figure 4, preferably corresponds to the time-domain output signal. However, if block 708 is not included in the rendering block of Figure 5, the result is a spectral-domain representation of the corresponding audio channel.

[0127] (The following steps are part of the state-of-the-art [Vilkamo2013] and have been added for clarity.) For each (sub)frame and each frequency band, an input covariance matrix C of size per transport channel x =xx Tis calculated from the decoded audio signal. Optionally, only the entries on the main diagonal can be used, in which case the other non-zero entries are set to zero. A prototype matrix of output channels of size per transport channel is defined, describing the mapping of transport channels to output channels (provided as part of the output configuration). The number of output channels is given by the target output format (e.g., the target loudspeaker layout). This prototype matrix can be static or change from frame to frame. For example, if only a single transport channel is transmitted, this transport channel is mapped to each output channel. If two transport channels are transmitted, the left (first) channel is mapped to all output channels located within (+0°, +180°), i.e., the "left" channel. The right (second) channel is mapped to all output channels located within (-0°, -180°), i.e., the "right" channel. (Note: 0° represents a position in front of the listener, positive angles represent positions to the left of the listener, and negative angles represent positions to the right of the listener. If a different convention is adopted, the signs of the angles must be adjusted accordingly.) Input covariance matrix C x , the target covariance matrix C Y Using the prototype matrices, a mixing matrix is ​​calculated for each (sub)frame and each frequency band [Vilkamo2013]. For example, 60 mixing matrices are obtained per (sub)frame. o The mixing matrix is ​​interpolated (e.g. linearly) between (sub)frames to accommodate temporal smoothing. Finally, the output channels y are combined band-wise by multiplying the set of final mixing matrices M of the output channels for each transport channel with the corresponding band of the time / frequency representation of the decoded transport channel x. y=Mx Note that we do not use the residual signal r, as explained in [Vilkamo2013].

[0128] The output signal y is transformed into a time domain representation y(t) using a filter bank.

[0129] Optimized covariance synthesis Input covariance matrix C x and the target covariance matrix C Y Due to how is calculated in this embodiment, specific optimization of the optimal mixing matrix calculation using covariance combining in [Vilkamo2013] can be achieved, significantly reducing the computational complexity of the mixing matrix calculation. Note that in this section, the Hadamard operator ○ represents an element-wise operation on a matrix. That is, instead of following rules such as matrix multiplication, each operation is performed element by element. This operator indicates that the corresponding operation is performed on each element individually, rather than on the entire matrix. For example, the multiplication of matrices A and B does not correspond to the matrix multiplication AB=C, but corresponds to the element-wise operation a_ij * b_ij = c_ij.

[0130] SVD(.) stands for singular value decomposition. The algorithm in [Vilkamo2013] is presented as a Matlab function (Listing 1) and is as follows (Prior Art):

[0131] [Table 1A]

[0132] [Table 1B]

[0133] As mentioned in the previous section, C x Only the leading diagonal elements of are optionally used, and all other entries are set to zero. In this case, C x is a diagonal matrix, and the effective decomposition satisfies equation (3) in [Vilkamo2013]. K x =C x ○1 / 2 The SVD from line 3 of the prior art algorithm is no longer needed.

[0134] Direct response to previous section dr i Considering the equation that generates the target covariance from the direct power (or direct energy),

[0135]

number

[0136] The last equation can be rearranged and written as:

[0137]

number

[0138] Now, if we define

[0139]

number

[0140] and therefore,

[0141]

number

[0142] is obtained. The direct response matrix R = [dr1…dr k ] and place the response directly in e i,i =E i ,C Y can also be expressed as follows: C Y =RER H And, C which satisfies the formula (3) in [Vilkamo2013] Y An effective decomposition of is given by: C y =RE ○1 / 2

[0143] Therefore, the SVD from line 1 of the prior art algorithm is no longer needed.

[0144] This leads to an optimized algorithm for covariance synthesis within the present embodiment, which also always uses the energy compensation option and therefore the residual target covariance C r It also takes into account that it does not require

[0145] [Table 2A]

[0146] [Table 2B]

[0147] A careful comparison of the prior art algorithm with the proposed algorithm shows that the former requires three SVDs of matrices of size m × m, n × n, and m × n, respectively, where m is the number of downmix channels and n is the number of output channels into which the object is rendered.

[0148] The proposed algorithm requires only one SVD of a matrix of size m × k, where k is the number of dominant objects. Moreover, since k is usually much smaller than n, this matrix is ​​smaller than the corresponding matrix in prior art algorithms.

[0149] The complexity of a standard SVD implementation is roughly O(c1m for an m×n matrix. 2 n+c2n 3 ) [Golub2013], where c1 and c2 are constants that depend on the algorithm used. Thus, a significant reduction in the computational complexity of the proposed algorithm is achieved compared to prior art algorithms.

[0150] Also, preferred embodiments relating to the encoder side of the first aspect are discussed with reference to Figures 7a and 7b, and preferred implementations of the encoder side of the second aspect are discussed with reference to Figures 9a to 9d.

[0151] FIG. 7a shows a preferred implementation of the object parameter calculator 100 of FIG. 1a. In block 120, audio objects are converted into a spectral representation. This is performed by the filter bank 102 of FIG. 2 or 3. Next, in block 122, selection information is calculated, for example, as shown in block 104 of FIG. 2 or 3. For this purpose, an amplitude-related measure can be used, such as amplitude itself, power, energy, or any other amplitude-related measure obtained by raising amplitude to a power other than 1. The result of block 122 is a set of selection information for each object in the corresponding time / frequency bin. Next, in block 124, an object ID for each time / frequency bin is derived. In a first embodiment, two or more object IDs are derived for each time / frequency bin. According to a second embodiment, the number of object IDs per time / frequency bin may be only a single object ID, so that the most important, strongest, or most relevant object is identified in block 124 from the information provided by block 122. Block 124 outputs information about the parameter data, including the index or indices of the most relevant object or objects.

[0152] In the case where there are two or more related objects per time / frequency bin, the function of block 126 serves to calculate an amplitude-related measure characterizing the object in the time / frequency bin. This amplitude-related measure may be the same as that calculated for the selection information in block 122, or preferably, a combined value is calculated using the information already calculated by block 102, as indicated by the dashed line between blocks 122 and 126. The amplitude-related measure or one or more combined values ​​are then calculated in block 126 and forwarded to quantizer and encoder block 212 as additional parametric side information to obtain the encoded amplitude-related or encoded combined value in the side information. In the embodiment of FIG. 2 or FIG. 3, these are the "encoded power ratios" included in the bitstream together with the "encoded object index." If there is only one object ID per frequency bin, the power ratio calculation and quantization encoding are not necessary, and the index of the most related object in the time-frequency bin is sufficient to perform decoder-side rendering.

[0153] FIG. 7b shows a preferred implementation of the calculation of the selection information 102 of FIG. 7b. As shown in block 123, signal power is calculated for each object and each time / frequency bin as selection information. Next, in block 125, which shows a preferred embodiment of block 124 of FIG. 7a, the object ID of the single object, or preferably two or more objects, with the highest power is extracted and output. Note that if there are multiple objects of interest, a preferred implementation of block 126 involves calculating a power ratio as shown in block 127, where the power ratio is calculated for the extracted object IDs relative to the power of all extracted objects with corresponding object IDs found by block 125. This procedure is advantageous because, since only one fewer combined value than the number of objects in the time / frequency bin needs to be transmitted, in this embodiment there is a rule known to the decoder stating that the power ratios of all objects must sum to 1. Preferably, the functionality of blocks 120, 122, 124, 126 of FIG. 7a and / or 123, 125, 127 of FIG. 7b is implemented by object parameter calculator 100 of FIG. 1a, and the functionality of block 212 of FIG. 7a is performed by output interface 200 of FIG. 1a.

[0154] Therefore, the apparatus for encoding according to the second aspect shown in FIG. 1b will be described in more detail with respect to several embodiments. In step 110a, directional information is extracted from the input signal or by reading or analyzing metadata information contained in a metadata portion or metadata file, as shown, for example, in FIG. 12a. In step 110b, the directional information and audio objects per frame are quantized, and quantization indices per object per frame are transferred to an encoder or an output interface, such as output interface 200 in FIG. 1b. In step 110c, the directional quantization indices are dequantized to obtain dequantized values, which may also be directly output by block 110b in certain implementations. Next, based on the dequantized directional indices, block 422 calculates weights for each transport channel and each object based on a specific virtual microphone configuration. This virtual microphone configuration may include two virtual microphone signals with different orientations located at the same position, or may be a configuration in which two different positions exist relative to a reference position or orientation, such as a virtual listener position or orientation. Setting two virtual microphone signals results in weights for two transport channels per object.

[0155] When generating three transport channels, the virtual microphone setup can be considered to include three virtual microphone signals from microphones placed at the same position but with different orientations, or from microphones placed at three different positions relative to a reference position or orientation, where the reference position for this orientation can be the position or orientation of the virtual listener.

[0156] Alternatively, the four transport channels can be generated based on a virtual microphone setup that generates four virtual microphone signals from microphones positioned at the same location but with different orientations, or from four virtual microphone signals positioned at four different locations relative to a reference position or direction, where the reference position or direction can be a virtual listener position or direction.

[0157] Furthermore, the weights w for each object and each transport channel are L and w R For purposes of calculating , in the two channel example, a virtual microphone signal is a signal derived from a virtual primary microphone, a virtual cardioid microphone or a virtual figure-eight microphone or a depot microphone, a bidirectional microphone, a virtual directional microphone, a virtual subcardioid microphone, a virtual unidirectional microphone, a virtual hypercardioid microphone, or a virtual omnidirectional microphone.

[0158] It should be noted that in this context, the actual microphone placement is not required for the purpose of calculating the weights. Instead, the rules for calculating the weights change depending on the virtual microphone configuration, i.e., the placement of the virtual microphones and their characteristics.

[0159] In block 404 of Fig. 9a, weights are applied to the objects, and for each object, if the weight is not zero, the object's contribution to a particular transport channel is obtained. Thus, block 404 receives the object signal as input. Then, in block 406, the contributions are summed for each transport channel, e.g., contributions from objects to a first transport channel are added together and contributions from objects to a second transport channel are added together. As shown in block 406, the output of block 406 is, e.g., a transport channel in the time domain.

[0160] Preferably, the object signal input to block 404 is a time-domain object signal with full-band information, and the application in block 404 and the summation in block 406 are performed in the time domain, although, in other words, these steps can also be performed in the spectral domain.

[0161] 9b shows a further embodiment in which a static downmix is ​​implemented. For this purpose, the directional information of the first frame is extracted in block 130 and weights are calculated according to the first frame, as shown in block 403a. Then, to implement a static downmix, the weights are left as they are for the other frames, as shown in block 408.

[0162] FIG. 9c shows an alternative implementation in which a dynamic downmix is ​​calculated. To this end, block 132 extracts directional information for each frame, and weights for each frame are updated as shown in block 403b. Then, in block 405, the updated weights are applied to the frame, implementing a dynamic downmix that changes from frame to frame. Other implementations between the extreme cases of FIG. 9b and FIG. 9c are equally useful, e.g., weights are updated only every second or third frame or every nth frame, and / or weight smoothing over time is performed to prevent antenna characteristics from changing too much from time to time for the purpose of downmixing according to directional information. FIG. 9d shows another implementation of a downmixer 400 controlled by the object direction information provider 110 of FIG. 1b. In block 410, the downmixer is configured to analyze the directional information of all objects in the frame, and in block 112, weights w for the stereo example are calculated. L and w RFor the purpose of calculating , the microphones are positioned according to the analysis results. Microphone placement refers to the microphone's location and / or microphone directionality. In block 414, the microphones are either left for other frames, similar to the static downmix discussed with respect to block 408 of FIG. 9b, or the microphones are updated as discussed with respect to block 405 of FIG. 9c to obtain the function of block 414 of FIG. 9d. Regarding the function of block 412, the microphones can be positioned to provide good separation, such that the first virtual microphone "sees" the first group of objects and the second virtual microphone "sees" the second group of objects. This differs from the first group of objects in that, preferably, objects from one group are not included in the other group, whenever possible. Alternatively, the analysis of block 410 can be enhanced by other parameters, and the placement can also be controlled by other parameters.

[0163] Subsequently, preferred implementations of decoders according to the first or second aspects are discussed, for example, with respect to Figures 6a and 6b and given with respect to Figures 10a, 10b, 10c, 10d and 11 below.

[0164] In block 613, the input interface 600 is configured to retrieve the individual object orientation information associated with the object ID. This procedure corresponds to the function of block 612 of Figure 4 or Figure 5 and results in the "frame codebook" shown and described with respect to Figure 8b, and in particular Figure 8c.

[0165] Furthermore, in block 609, one or more object IDs for each time / frequency bin are retrieved, regardless of whether those data are available for low-resolution parameter bands or high-resolution frequency tiles. The result of block 609, which corresponds to the procedure of block 608 in FIG. 4, is the specific IDs within the time / frequency bins of one or more relevant objects. Next, in block 611, specific object direction information for the specific ID(s) for each time / frequency bin is retrieved from the "frame codebook," i.e., from the exemplary table shown in FIG. 8c. Then, in block 704, gain values ​​are calculated for one or more relevant objects for each output channel, as governed by the output format calculated for each time / frequency bin. Next, in blocks 730, 706, and 708, the output channels are calculated. This function of calculating the output channels can be performed within the explicit calculation of contributions from one or more transport channels, as shown in FIG. 10b, or can be performed by indirectly calculating and using the transport channel contributions, as shown in FIG. 10d or FIG. 11. Figure 10b shows a function in which power values ​​or power ratios are retrieved in block 610, which corresponds to the function in Figure 4. These power values ​​are then applied to the individual transport channels for each associated object, as shown in blocks 733 and 735. Furthermore, because these power values ​​are applied to the individual transport channels in addition to the gain values ​​determined by block 704, blocks 733, 735 result in object-specific contributions of transport channels ch1, ch2, ..., etc. These explicitly calculated channel transport contributions are then summed for each output channel for each time / frequency bin in block 737.

[0166] Depending on the implementation, a diffusion signal calculator 741 may then be provided to generate a diffusion signal in the time / frequency bins corresponding to each output channel ch1, ch2, ..., where the diffusion signal and the contribution results of block 737 are combined to obtain the complete channel contribution in each time / frequency bin. This signal corresponds to the input to the filter bank 708 in FIG. 4 if the covariance synthesis further relies on the diffusion signal. However, if the covariance synthesis 706 does not rely on the diffusion signal and relies only on decorrelator-less processing, the energy of the output signal for at least each time / frequency bin corresponds to the energy of the channel contribution at the output of block 739 in FIG. 10b. Furthermore, if the diffusion signal calculator 741 is not used, the result of block 739 corresponds to the result of block 706, with the complete channel contribution for each time / frequency bin being converted separately for each output channel ch1, ch2. Finally, the time-domain output channels can be saved or transferred to a loudspeaker or any kind of rendering device to obtain an output audio file.

[0167] Figure 10c shows a preferred implementation of the function of block 610 of Figure 10b or Figure 4. In step 610a, a combined (power) value or values ​​are retrieved for a particular time / frequency bin. In block 610b, other values ​​corresponding to other related objects in the time / frequency bin are calculated based on a calculation rule that all combined values ​​must sum to one.

[0168] The result is then preferably a low-resolution representation with two power ratios for each grouped time slot index and each parameter band index, representing the low time / frequency resolution. In block 610c, the time / frequency resolution can be expanded to a high time / frequency resolution, with power values ​​for time / frequency tiles with high-resolution time slot index n and high-resolution frequency band index k. The expansion can include simply using the same low-resolution index for corresponding time slots within the grouped time slots and corresponding frequency bands within the parameter bands.

[0169] FIG. 10d illustrates a preferred implementation of the functionality for calculating covariance combining information in block 706 of FIG. 4, represented by a mixing matrix 725 used to mix two or more input transport channels into two or more output signals. Thus, for example, if there are two transport channels and six output channels, the size of the mixing matrix for each individual time / frequency bin is 6 rows and 2 columns. Block 723, which corresponds to the functionality of block 723 of FIG. 5, receives gain or direct response values ​​for each object in each time / frequency bin and calculates a covariance matrix. Block 722 receives power values ​​or ratios and calculates direct power values ​​for each object in the time / frequency bin; block 722 of FIG. 10d corresponds to block 722 of FIG. 5.

[0170] The results of both blocks 721 and 722 are input to a target covariance matrix calculator 724. Additionally or alternatively, the target covariance matrix C y 5. Instead, the relevant information contained in the target covariance matrix, i.e., the direct response value information represented by matrix R and the direct power values ​​of two or more relevant objects represented by matrix E, is input to block 725a for the mixing matrix calculation for each time / frequency bin. Furthermore, the mixing matrix 725a is calculated by adding together the information about the prototype matrix Q and the input covariance matrix C derived from two or more transport channels shown in block 726, which corresponds to block 726 in FIG. 5.x 5. The mixing matrices for each time / frequency bin and frame can be subjected to time smoothing as shown in block 725b, and in block 727, which corresponds to at least a portion of the rendering block of FIG. 5, the mixing matrix is ​​applied in unsmoothed or smoothed form to the transport channels for the corresponding time / frequency bins to obtain, at the output of block 739, full channel contributions in the time / frequency bins substantially similar to the corresponding full contributions discussed above with respect to FIG. 10b. Thus, FIG. 10b illustrates an implementation of explicit calculation of the transport channel contributions, while FIG. 10d illustrates a procedure for implicitly calculating the transport channel contributions for each time / frequency bin and for each associated object within each time / frequency bin via a target covariance matrix C or via the associated information R and E of blocks 723 and 722, which are introduced directly into the mixing matrix calculation block 725a.

[0171] Next, a preferred optimization algorithm for covariance combining is shown in Figure 11. It is outlined that all steps shown in Figure 11 are calculated within the covariance combining 706 in Figure 4, or within the mixing matrix calculation block 725 in Figure 5, or 725a in Figure 10d. In step 751, the first decomposition result K y is calculated. This decomposition result can be calculated easily without calculating the covariance matrix, as shown in Figure 10d, because the information of the obtained values ​​contained in matrix R and information from two or more related objects, in particular the direct power information contained in matrix ER, is used directly and not explicitly. In this way, a specific singular value decomposition is no longer necessary, so the first decomposition result in block 751 can be calculated directly and without much effort.

[0172] In step 752, the second decomposition result is K x This decomposition can also be computed without explicit singular value decomposition, since the input covariance matrix is ​​treated as a diagonal matrix with off-diagonal elements ignored.

[0173] Next, in step 753, a first regularization result based on the first regularization parameter α is calculated, and in step 754, a second regularization result based on the second regularization parameter β is calculated. x is a diagonal matrix in the preferred implementation, the calculation of the first normalized result 753 is performed as in the prior art using S x is simplified with respect to the prior art because the calculation of is simply a parameter change rather than a decomposition.

[0174] Furthermore, with regard to the calculation of the second regularized result in block 754, the first step is to calculate the matrix U x HS It is not a multiplication with , but just a renaming of the parameter.

[0175] Furthermore, in step 755, the normalized matrix G y is calculated and based on step 755, the unitary matrix P is calculated in step 756 by x , the prototype matrix Q, and the K obtained by block 751 y The calculation of the unitary matrix P is simplified relative to available prior art techniques by the fact that the matrix Λ is not required here.

[0176] Next, in step 757, M opt A mixing matrix without energy compensation, where M is the mixing matrix without energy compensation, is calculated using the unitary matrix P, the result of block 754, and the result of block 751. Energy compensation is then performed in block 758 using the compensation matrix G. Because energy compensation is performed, the residual signal derived from the decorrelator is not needed. However, instead of performing energy compensation, this implementation calculates the mixing matrix M without energy information. opt A residual signal with sufficient energy to fill the energy gap left by ( ) is added. However, for the purposes of the present invention, the decorrelated signal is not relied upon to avoid artifacts introduced by the decorrelator. However, energy compensation as shown in step 758 is preferred.

[0177] Thus, the optimized algorithm for covariance synthesis offers advantages within steps 751, 752, 753, and 754, as well as within step 756 for calculating the unitary matrix P. It should be emphasized that the optimized algorithm even offers advantages over prior art implementations in which only one or a subgroup of steps 755, 752, 753, 754, and 756 are implemented as shown, while the corresponding other steps are implemented as in the prior art. This is because the improvements can be applied independently of each other, rather than being interdependent. However, the more improvements are implemented, the better the procedure is in terms of implementation complexity. Therefore, while the full implementation of the embodiment of FIG. 11 provides the greatest reduction in complexity, it is preferable because even when only one of steps 751, 752, 753, 754, and 756 is implemented according to the optimized algorithm and the other steps are implemented as in the prior art, a reduction in complexity is obtained without a loss of quality.

[0178] An embodiment of the present invention can also be viewed as a procedure for generating comfort noise for a stereo signal by mixing three Gaussian noise sources, one per channel and a third common noise source, to create correlated background noise, or additionally or separately to control the mixing of the noise sources with the coherence values ​​transmitted in the SID frames.

[0179] It is here noted that all alternatives or aspects described above and below, and all aspects defined by the following claims or aspects of the claims, can be used individually, i.e., there are no alternatives or objectives other than those contemplated in the independent claims. However, in other embodiments, two or more of the alternatives or aspects or independent claims can be combined with each other, and in other embodiments, all aspects or alternatives and all independent claims can be combined with each other.

[0180] Signals encoded according to the present invention can be stored on a digital or non-transitory storage medium or transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0181] While some aspects are described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or a function of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or function of a corresponding apparatus.

[0182] Of course, depending on particular implementation requirements, embodiments of the present invention can be implemented in hardware or software. The implementation can be carried out using a digital storage medium, such as a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM or flash memory, having electronically readable control signals stored thereon and cooperating (or being able to cooperating) with a programmable computer system so that the respective methods are performed.

[0183] Some embodiments according to the invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0184] Generally, embodiments of the present invention may be implemented as a computer program product having program code such that when the computer program product runs on a computer, the program code operates to perform one of the methods. The program code may, for example, be stored on a machine-readable carrier.

[0185] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier or a non-transitory storage medium.

[0186] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0187] A further embodiment of the inventive method is, therefore, a data carrier (or digital storage medium, or computer-readable medium) comprising a computer program recorded thereon for performing one of the methods described herein.

[0188] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, the data stream or sequence of signals being adapted to be transmitted via a data communication connection, for example the Internet.

[0189] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0190] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0191] In some embodiments, a programmable logic device (such as a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may be used in conjunction with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0192] The above-described embodiments are merely illustrative of the principles of the present invention. It is to be understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented in the description and explanation.

[0193] Aspects (used independently of each other, together with all other aspects, or only with a subgroup of other aspects)

[0194] An apparatus, method, or computer program comprising one or more of the following features:

[0195] Examples of inventions relating to novel aspects: Combine the multi-wave idea with object coding (using multiple directional cues per T / F tile) An object coding approach that is as close as possible to the DirAC paradigm, allowing any kind of input type in IVAS (object content has not been covered so far).

[0196] Inventive example of parameterization (encoder): For each T / F tile: Selection information of the n most relevant objects in this T / F tile and the power ratio between the contributions of those n most relevant objects Each frame, each object: 1 direction

[0197] Inventive example of rendering (decoder): · Retrieve the direct response value of each associated object from the transmitted object index and orientation information and the target output layout. · Obtain the covariance matrix from the direct response. · Calculate the power for each relevant object directly from the downmix signal power and the transmit power ratio. Obtain the final target covariance matrix from the direct power and covariance matrices. · Use only the diagonal elements of the input covariance matrix. Optimized covariance synthesis

[0198] Additional notes on differences from SAOC: · n dominant objects are considered instead of all objects. → Thus, the power ratio is related to OLD, but calculated in a different way. SAOC does not use orientation in the encoder -> orientation information (rendering matrix) is only introduced in the decoder. → The SAOC-3D decoder receives object metadata for rendering matrices. SAOC adopts downmix matrix and transmits downmix gain. Diffusivity is not considered in embodiments of the present invention.

[0199] Subsequently, further embodiments of the present invention are summarized.

[0200] 1. An apparatus for encoding a plurality of audio objects and associated metadata indicating directional information for the plurality of audio objects, comprising: a downmixer (400) for downmixing a plurality of audio objects to obtain one or more transport channels; a transport channel encoder (300) for encoding one or more transport channels to obtain one or more encoded transport channels; an output interface (200) for outputting an encoded audio signal comprising one or more encoded transport channels; Equipped with The downmixer (400) is configured to downmix the plurality of audio objects in response to the directional information of the plurality of audio objects. Device.

[0201] 2. Down mixer (400) generating two transport channels as two virtual microphone signals that are co-located but oriented differently or at two different positions relative to a reference position or direction, such as the position or direction of a virtual listener; or generating three transport channels as three virtual microphone signals that are co-located and oriented differently or at three different positions relative to a reference position or direction, such as the position or direction of a virtual listener; or generating four transport channels as four virtual microphone signals that are co-located and oriented differently or at four different positions relative to a reference position or direction, such as the position or direction of a virtual listener; It is configured as follows: The virtual microphone signal is a virtual primary microphone signal, or a virtual cardioid microphone signal, or a virtual figure-eight or dipole or bidirectional microphone signal, or a virtual directional microphone signal, or a virtual subcardioid microphone signal, or a virtual unidirectional microphone signal, or a virtual hypercardioid microphone signal, or a virtual omnidirectional microphone signal; The device described in Example 1.

[0202] 3. Down mixer (400) For each of the plurality of audio objects, deriving (402) weighting information for each transport channel using directional information of the corresponding audio object; weighting (404) the corresponding audio object using the weighting information of the audio object of the particular transport channel to obtain an object contribution of the particular transport channel; combining (406) object contributions of a particular transport channel from the plurality of audio objects to obtain a particular transport channel; It is configured as follows: 10. The device of Example 1 or 2.

[0203] 4. The downmixer (400) is configured to calculate one or more transport channels as one or more virtual microphone signals that are co-located, have different orientations, or are at another position with respect to a reference position or direction, such as the position or direction of a virtual listener to which the directional information relates; the different positions or orientations are on or to the left of the center line and on or to the right of the center line, or the different positions or orientations are distributed evenly or unevenly across horizontal positions or orientations such as +90 degrees or -90 degrees relative to the center line, or -120 degrees, 0 degrees, and +120 degrees relative to the center line, or the different positions or orientations include at least one position or orientation that is upward or downward relative to a horizontal plane in which a virtual listener is placed, and the directional information for the plurality of sound objects is associated with the position or reference position or orientation of the virtual listener; The device according to any one of Examples 1 to 3.

[0204] 5. A parameter processor (110) for quantizing metadata indicating directional information for the plurality of audio objects to obtain quantized directional terms for the plurality of audio objects; The downmixer (400) is configured to operate in response to the quantized directional terms as the directional information; the output interface (200) is configured to introduce information about the quantized directional terms into the encoded audio signal; The device according to any one of Examples 1 to 4.

[0205] 6. The downmixer (400) is configured to perform an analysis of directional information of a plurality of audio objects and to position one or more virtual microphones to generate transport channels according to the results of the analysis; 6. The device of any one of Examples 1 to 5.

[0206] 7. The downmixer (400) is configured to downmix (408) using static downmix rules over multiple time periods; or the direction information is variable over a plurality of time periods, and the downmixer (400) is configured to downmix (405) using downmixing rules that are variable over a plurality of time periods; The device according to any one of Examples 1 to 6.

[0207] 8. The apparatus according to any one of embodiments 1 to 7, wherein the downmixer (400) is configured to downmix in the time domain using sample-wise weighting and combining samples of multiple audio objects.

[0208] 9. An object parameter calculator (100) configured to calculate parameter data of at least two associated audio objects for one or more frequency bins of a plurality of frequency bins associated with a time frame, wherein the number of the at least two associated audio objects is less than the total number of the plurality of audio objects; Furthermore, an output interface (200) configured to introduce information about parameter data of at least two related audio objects in one or more frequency bins into the encoded audio signal; 9. The device of any one of Examples 1 to 8.

[0209] 10. The object parameter calculator (100) converting (120) each audio object of the plurality of audio objects into a spectral representation having a plurality of frequency bins; calculating (122) selection information from each audio object in one or more frequency bins; deriving object identifications (124) based on the selection information as parameter data indicative of at least two related audio objects; It is configured as follows: an output interface (200) configured to introduce information regarding object identification into the encoded audio signal; The device described in Example 9.

[0210] 11. The object parameter calculator (100) is configured to quantize and encode (212) one or more amplitude-related measures of the associated audio object or one or more combined values ​​derived from the amplitude-related measures as parameter data in one or more frequency bins; an output interface (200) configured to introduce one or more quantized amplitude-related measures or one or more quantized combined values ​​into the encoded audio signal; 11. The device of Example 9 or 10.

[0211] 12. The selection information is an amplitude-related measurement, such as an amplitude value, a power value, or a loudness value, or an amplitude raised to a power different from the amplitude of the audio object; an object parameter calculator (100) configured to calculate (127) an amplitude-related measure of the associated audio object and a combined value, such as a ratio, from a sum of two or more amplitude-related measures of the associated audio object; an output interface (200) configured to introduce information about the combined values ​​into the encoded audio signal, the number of information items about the combined values ​​of the encoded audio signal being at least equal to one and less than the number of associated audio objects of one or more frequency bins; 12. The device of example 10 or 11.

[0212] 13. The object parameter calculator (100) is configured to select an object identification based on an order of selection information of a plurality of audio objects in one or more frequency bins; The device according to any one of Examples 10 to 12.

[0213] 14. The object parameter calculator (100) Calculate the signal power as the selection information (122); deriving (124) object identities for two or more audio objects having the greatest signal power values ​​in one or more frequency bins corresponding to each of the frequency bins; calculating (126) a power ratio between the sum of the signal powers of the two or more audio objects having the largest signal power values ​​and the signal power of at least one audio object having the derived object identification as parameter data; quantizing and encoding (212) the power ratio; It is structured as follows: The output interface (200) is configured to introduce the quantized and encoded power ratios into the encoded audio signal. 14. The device of any one of Examples 10 to 13.

[0214] 15. The device of any one of Examples 10 to 14, wherein the output interface (200) is configured to introduce into the encoded audio signal one or more encoded transport channels, and as parameter data, for each of one or more frequency bins of a plurality of frequency bins in a time frame, two or more encoded object identifications of associated audio objects and one or more encoded combination values ​​or encoded amplitude-related measures, and quantized and encoded directional data for each audio object in the time frame, which directional data is constant for all frequency bins of the one or more frequency bins.

[0215] 16. The object parameter calculator (100) is configured to calculate parameter data for at least a most dominant object and a second most dominant object in one or more frequency bins; The number of the voice objects of the plurality of voice objects is three or more, and the plurality of voice objects includes a first voice object, a second voice object, and a third voice object; the object parameter calculator (100) is configured to: calculate, for a first frequency bin of the one or more frequency bins, only a first group of audio objects, such as a first audio object and a second audio object, as associated audio objects; and calculate, for a second frequency bin of the one or more frequency bins, only a second group of audio objects, such as the second audio object and a third audio object or the first audio object and a third audio object, as associated audio objects; wherein the first group of audio objects differs from the second group of audio objects with respect to at least one group member; The device according to any one of Examples 9 to 15.

[0216] 17. The object parameter calculator (100) calculating raw parametric data at a first time or frequency resolution, combining the raw parametric data into combined parametric data having a second time or frequency resolution lower than the first time or frequency resolution, and calculating parameter data of at least two related audio objects with respect to the combined parametric data having the second time or frequency resolution; or determining a parameter band having a second time resolution or frequency resolution different from the first time resolution or frequency resolution used in the time decomposition or frequency decomposition of the plurality of audio objects, and calculating parameter data of at least two related audio objects for the parameter band having the second time resolution or frequency resolution; The device according to any one of Examples 9 to 16, configured as follows:

[0217] 18. A decoder for decoding an encoded audio signal comprising one or more transport channels, directional information of a plurality of audio objects, and parameter data of the audio objects for one or more frequency bins of a time frame, comprising: an input interface (600) for providing one or more transport channels in a spectral representation having a plurality of frequency bins within a time frame; an audio renderer (700) for rendering one or more transport channels into multiple audio channels using the direction information; Equipped with an audio renderer (700) configured to calculate direct response information (704) from one or more audio objects for each frequency bin of the plurality of frequency bins, and to calculate directional information (810) associated with the associated one or more audio objects in the frequency bin; decoder.

[0218] 19. An audio renderer (700) configured to use the direct response information and information regarding the number of audio channels (702) to calculate (706) covariance synthesis information and apply (727) the covariance synthesis information to one or more transport channels to obtain the number of audio channels; the direct response information (704) is a direct response vector for each of the one or more audio objects, the covariance synthesis information is a covariance synthesis matrix, and the audio renderer (700) is configured to perform a matrix operation for each frequency bin when applying (727) the covariance synthesis information; 19. The decoder of Example 18.

[0219] 20. The audio renderer (700) Calculating the direct response information (704) includes deriving direct response vectors for one or more audio objects, and calculating a covariance matrix from each direct response vector for the one or more audio objects; In calculating the covariance synthesis information, deriving target covariance information from a covariance matrix of one audio object or from covariance matrices from multiple audio objects, power information for each of the one or more audio objects, and power information derived from one or more transport channels (724); 20. The decoder according to claim 18 or 19, configured as follows:

[0220] 21. The audio renderer (700) In calculating the direct response information, deriving direct response vectors for one or more speech objects and calculating (723) a covariance matrix from each direct response vector for each of the one or more speech objects; Deriving input covariance information from the transport channel (726); Deriving mixing information from the target covariance information, the input covariance information, and information regarding the number of channels (725a, 725b); applying mixing information to the transport channel for each frequency bin within the time window (727); 21. The decoder of claim 20, configured as follows:

[0221] 22. The decoder of embodiment 21, wherein the result of applying the mixing information to each frequency bin in the time window is transformed (708) into the time domain to obtain the number of audio channels in the time domain.

[0222] 23. Audio Renderer (700) using only the main diagonal elements of the input covariance matrix derived from the transport channel in the decomposition (752) of the input covariance matrix; Performing a decomposition (751) of the target covariance matrix using the direct response matrix and the power matrix of the object or transport channel; Performing a decomposition of the input covariance matrix (752) by taking the roots of each main diagonal element of the input covariance matrix; Compute the normalized inverse of the decomposed input covariance matrix (753), Perform singular value decomposition to calculate the optimal matrix used for energy compensation without the extended identity matrix (756) The decoder according to any one of the eighteenth to twenty-second embodiments, configured as described above.

[0223] 24. The parameter data of the one or more audio objects includes parameter data of at least two associated audio objects, and the number of the at least two associated audio objects is less than the total number of the multiple audio objects; the audio renderer (700) is configured to calculate, for each of the one or more frequency bins, contributions from one or more transport channels according to first directional information associated with a first of the at least two associated audio objects and according to second directional information associated with a second of the at least two associated audio objects; 24. The decoder of any one of Examples 18 to 23.

[0224] 25. The decoder of Example 24, wherein the audio renderer (700) is configured to ignore, for one or more frequency bins, directional information of an audio object that is different from at least two related audio objects.

[0225] 26. The encoded audio signal includes amplitude-related measurements for each associated audio object or a combined value associated with at least two associated audio objects in the parameter data; the audio renderer (700) is configured to operate such that contributions from one or more transport channels are taken into account according to first directional information associated with a first of the at least two associated audio objects and according to second directional information associated with a second of the at least two associated audio objects, or to determine quantitative contributions of one or more transport channels according to amplitude-related measurements or combination values; 26. The decoder of Example 24 or 25.

[0226] 27. The encoded signal includes a combined value in the parameter data; an audio renderer (700) configured to determine the contribution of one or more transport channels using the combined value for one of the associated audio objects and the directional information for the one associated audio object; the audio renderer (700) is configured to determine the contribution of one or more transport channels using combined values ​​of other associated audio objects in one or more frequency bins and values ​​derived from directional information of the other associated audio objects; The decoder of Example 26.

[0227] 28. Audio Renderer (700) calculating, for each frequency bin of the plurality of frequency bins, direct response information (704) from the associated sound object and directional information associated with the associated sound object within the frequency bin; 28. The decoder of any one of Examples 24 to 27, configured as follows:

[0228] 29. The audio renderer (700) determines (741) a diffusion signal for each frequency bin of the plurality of frequency bins using diffusion information, such as diffusion parameters or decorrelation rules, included in the metadata, and combines the direct responses to obtain a signal rendered in a spectral region of a channel among the plurality of channels, determined by the direct response information and the diffusion signal; 29. The decoder of Example 28.

[0229] 30. A method of encoding a plurality of audio objects and associated metadata indicating directional information for the plurality of audio objects, comprising: - downmixing a plurality of audio objects to obtain one or more transport channels; encoding one or more transport channels to obtain one or more encoded transport channels; outputting an encoded audio signal comprising one or more encoded transport channels; Including, the downmixing step comprises downmixing the plurality of audio objects in response to directional information relating to the plurality of audio objects. method.

[0230] 31. A method of decoding an encoded audio signal containing one or more transport channel and direction information for a plurality of audio objects and parameter data for the audio objects for one or more frequency bins of a time frame, comprising: providing one or more transport channels in a spectral representation having a plurality of frequency bins within a time frame; audio rendering one or more transport channels into a plurality of audio channels using the directional information; Including, the step of rendering the audio includes calculating, for each frequency bin of the plurality of frequency bins, direct response information from one or more audio objects and directional information associated with the associated one or more audio objects in the frequency bin; method.

[0231] 32. A computer program for performing the method of example 30 or the method of example 31 when running on a computer or processor.

[0232] (References) [Pulkki2009] V. Pulkki, M-V. Laitinen, J. Vilkamo, J. Ahonen, T. Lokki, and T. Pihlajamaeki, “Directional audio coding perception-based reproduction of spatial sound”, International Workshop on the Principles and Application on Spatial Hearing, Nov. 2009, Zao; Miyagi, Japan. [SAOC_STD] ISO / IEC, “MPEG audio technologies Part 2: Spatial Audio Object Coding (SAOC).” ISO / IEC JTC1 / SC29 / WG11 (MPEG) International Standard 23003-2. [SAOC_AES] J. Herre, H. Purnhagen, J. Koppens, O. Hellmuth, J. Engdegaard, J.Hilpert, L. Villemoes, L. Terentiv, C. Falch, A. Hoelzer, M. L. Valero, B. Resch, H. Mundt H, and H. Oh, “MPEG spatial audio object coding - the ISO / MPEG standard for efficient coding of interactive audio scenes,” J. AES, vol. 60, no. 9, pp. 655 - 673, Sep. 2012. [MPEGH_AES] J. Herre, J. Hilpert, A. Kuntz, and J. Plogsties, “MPEG-H audio - the new standard for universal spatial / 3D audio coding,” in Proc. 137th AES Conv., Los Angeles, CA, USA, 2014. [MPEGH_IEEE] J. Herre, J. Hilpert, A. Kuntz, and J. Plogsties, “MPEG-H 3D Audio - The New Standard for Coding of Immersive Spatial Audio“, IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, VOL. 9, NO. 5, AUGUST 2015 [MPEGH_STD] Text of ISO / MPEG 23008 - 3 / DIS 3D Audio, Sapporo, ISO / IEC JTC1 / SC29 / WG11 N14747, Jul. 2014. [SAOC_3D_PAT] APPARATUS AND METHOD FOR ENHANCED SPATAL AUDIO OBJECT CODING, WO 2015 / 011024 A1 [Pulkki1997] V. Pulkki, “Virtual sound source positioning using vector base amplitude panning,” J. Audio Eng. Soc., vol. 45, no. 6, pp. 456 - 466, Jun. 1997. [DELAUNAY] C. B. Barber, D. P. Dobkin, and H. Huhdanpaa, “The quickhull algorithm for convex hulls,” in Proc. ACM Trans. Math. Software (TOMS), New York, NY, USA, Dec. 1996, vol. 22, pp. 469 - 483. [Hirvonen2009] T. Hirvonen, J. Ahonen, and V. Pulkki, “Perceptual compression methods for metadata in Directional Audio Coding applied to audiovisual teleconference”, AES 126th Convention 2009, May 7 - 10, Munich, Germany. [Borss2014] C. Borss, “A Polygon-Based Panning Method for 3D Loudspeaker Setups”, AES 137th Convention 2014, October 9 - 12, Los Angeles, USA. [WO2019068638] Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to DirAC based spatial audio coding, 2018 [WO2020249815] PARAMETER ENCODING AND DECODING FOR MULTICHANNEL AUDIO USING DirAC, 2019 [BCC2001] C. Faller, F. Baumgarte: “Efficient representation of spatial audio using perceptual parametrization”, Proceedings of the 2001 IEEE Workshop on the Applications of Signal Processing to Audio and Acoustics (Cat. No.01TH8575). [JOC_AES] Heiko Purnhagen; Toni Hirvonen; Lars Villemoes; Jonas Samuelsson; Janusz Klejsa: “Immersive Audio Delivery Using Joint Object Coding”, 140th AES Convention, Paper Number: 9587, Paris, May 2016. [AC4_AES] K. Kjoerling, J. Roeden, M. Wolters, J. Riedmiller, A. Biswas, P. Ekstrand, A. Groeschel, P. Hedelin, T. Hirvonen, H. Hoerich, J. Klejsa, J. Koppens, K. Krauss, H-M. Lehtonen, K. Linzmeier, H. Muesch, H. Mundt, S. Norcross, J. Popp, H. Purnhagen, J. Samuelsson, M. Schug, L. Sehlstroem, R. Thesing, L. Villemoes, and M. Vinton: “AC-4 - The Next Generation Audio Codec”, 140th AES Convention, Paper Number: 9491, Paris, May 2016. [Vilkamo2013] J. Vilkamo, T. Baeckstroem, A. Kuntz, “Optimized covariance domain framework for time-frequency processing of spatial audio”, Journal of the Audio Engineering Society, 2013. [Golub2013] Gene H. Golub and Charles F. Van Loan, “Matrix Computations”, Johns Hopkins University Press, 4th edition, 2013. [Explanation of symbols]

[0233] 100 Object Parameter Calculator 110 Parameter Processor 200 output interfaces 212 encoding 300 Transport Channel Encoder 400 Down Mixer 405 Downmix 600 Input Interface 700 Audio Renderer 704 Direct Response Information 810 Direction information 812 Amplitude-related measurements

Claims

1. 1. An apparatus for encoding a plurality of audio objects, comprising: an object parameter calculator (100) configured to calculate parameter data of at least two associated audio objects for one or more frequency bins of a plurality of frequency bins associated with a time window, wherein the number of the at least two associated audio objects is less than a total number of the plurality of audio objects; an output interface (200) for outputting an encoded audio signal containing information about the parameter data of the at least two associated audio objects for the one or more frequency bins; An apparatus comprising:

2. The object parameter calculator (100) converting (120) each sound object of the plurality of sound objects into a spectral representation having the plurality of frequency bins; calculating (122) selection information from each audio object for the one or more frequency bins; deriving (124) object identifications as the parameter data indicative of the at least two related audio objects based on the selection information; It is structured as follows: the output interface (200) is configured to introduce information about the object identification into the encoded audio signal; 10. The apparatus of claim 1.

3. the object parameter calculator (100) is configured to quantize and encode (212) as the parameter data one or more amplitude-related measures of the associated audio objects in the one or more frequency bins or one or more combined values ​​derived from the amplitude-related measures; the output interface (200) is configured to introduce the quantized amplitude-related measures or the quantized combined values ​​into the encoded audio signal; 3. The device according to claim 1 or 2.

4. the selected information is an amplitude-related measure such as an amplitude value, a power value or a loudness value, or an amplitude raised to a power different from the amplitude of the audio object; the object parameter calculator (100) is configured to calculate (127) a combined value, such as a ratio, from a measure related to an associated audio object and a sum of two or more amplitude-related measures of the associated audio object; the output interface (200) is configured to introduce information about the combination values ​​into the encoded audio signal, the number of information items about the combination values ​​of the encoded audio signal being at least equal to one and less than the number of associated audio objects of the one or more frequency bins; 4. The device according to claim 2 or 3.

5. the object parameter calculator (100) is configured to select the object identification based on an order of the selection information of the plurality of audio objects within the one or more frequency bins.

5. An apparatus according to any one of claims 2 to 4.

6. The object parameter calculator (100) Calculating (122) signal power as the selection information; deriving (124) the object identifications for the two or more audio objects having the greatest signal power values ​​in one or more frequency bins corresponding to each of the frequency bins individually; calculating (126) a power ratio between the sum of the signal powers of the two or more audio objects having the largest signal power values ​​and the signal power of each audio object having the derived object identification as its parameter data; quantizing and encoding (212) the power ratio; It is structured as follows: the output interface (200) is configured to introduce the quantized and encoded power ratio into the encoded audio signal; 6. Apparatus according to any one of claims 2 to 5.

7. The output interface (200) outputs to the encoded audio signal: one or more encoded transport channels; as said parameter data, for each of said one or more frequency bins of said plurality of frequency bins within said time window, two or more encoded object identifications of said associated audio objects and one or more encoded combination values ​​or encoded amplitude-related measures; quantized and encoded directional data for each audio object within the time window, the directional data being constant for all frequency bins of the one or more frequency bins; configured to introduce 7. An apparatus according to any one of claims 1 to 6.

8. the object parameter calculator (100) is configured to calculate parameter data for at least a most dominant object and a second most dominant object in the one or more frequency bins; or The number of voice objects of the plurality of voice objects is three or more, and the plurality of voice objects includes a first voice object, a second voice object, and a third voice object; the object parameter calculator (100) is configured to calculate, for a first one of the one or more frequency bins, only a first group of audio objects, such as the first audio object and the second audio object, as the associated audio objects, and to calculate, for a second frequency bin of the one or more frequency bins, only a second group of audio objects, such as the second audio object and the third audio object or the first audio object and the third audio object, as the associated audio objects, wherein the first group of audio objects differs from the second group of audio objects with respect to at least one group member; 8. Apparatus according to any one of claims 1 to 7.

9. The object parameter calculator (100) calculating raw parametric data at a first time or frequency resolution, combining the raw parametric data into combined parametric data having a second time or frequency resolution lower than the first time or frequency resolution, and calculating parameter data of the at least two related audio objects with respect to the combined parametric data having the second time or frequency resolution; or determining parameter bands having a second time resolution or frequency resolution different from the first time resolution or frequency resolution used in the time decomposition or frequency decomposition of the plurality of audio objects, and calculating the parameter data of the at least two related audio objects for the parameter bands having the second time resolution or frequency resolution; 9. The device according to claim 1, configured to:

10. the plurality of audio objects include associated metadata indicating directional information (810) regarding the plurality of audio objects; The device, a downmixer (400) for downmixing the plurality of audio objects to obtain one or more transport channels, the downmixer (400) being configured to downmix the plurality of audio objects in response to the directional information of the plurality of audio objects; a transport channel encoder (300) for encoding one or more transport channels to obtain one or more encoded transport channels; Furthermore, the output interface (200) is configured to introduce the one or more transport channels into the encoded audio signal; 10. Apparatus according to any one of claims 1 to 9.

11. The downmixer (400) generating two transport channels as two virtual microphone signals located at the same position and with different orientations or at two different positions relative to a reference position or orientation, such as the position or orientation of a virtual listener; or generating three transport channels as three virtual microphone signals located at the same position, with different orientations, or at three different positions relative to a reference position or orientation, such as the position or orientation of a virtual listener; or generating four transport channels as four virtual microphone signals positioned at the same position, with different orientations, or at four different positions relative to a reference position or orientation, such as the position or orientation of a virtual listener; It is structured as follows: The virtual microphone signal is a virtual primary microphone signal, or a virtual cardioid microphone signal, or a virtual figure-eight or dipole or bidirectional microphone signal, or a virtual directional microphone signal, or a virtual subcardioid microphone signal, or a virtual unidirectional microphone signal, or a virtual hypercardioid microphone signal, or a virtual omnidirectional microphone signal; 11. The apparatus of claim 10.

12. The downmixer (400) for each of the plurality of audio objects, deriving (402) weighting information for each transport channel using the directional information of the corresponding audio object; weighting (404) the corresponding audio object using the weighting information of the audio object of a particular transport channel to obtain an object contribution of the particular transport channel; combining (406) the object contributions of the particular transport channel from the plurality of audio objects to obtain the particular transport channel; 12. The device according to claim 10 or 11, configured to:

13. the downmixer (400) is configured to calculate the one or more transport channels as one or more virtual microphone signals that are co-located, oriented differently, or located at different positions relative to a reference position or orientation, such as a position or orientation of a virtual listener with which the directional information is associated; The different positions or orientations are on a center line or to the left of the center line, and on a center line or to the right of the center line, or the different positions or orientations are distributed evenly or unevenly among horizontal positions or orientations such as +90 degrees or -90 degrees with respect to the center line, or -120 degrees, 0 degrees, and +120 degrees with respect to the center line, or the different positions or orientations include at least one position or orientation oriented upward or downward with respect to a horizontal plane on which a virtual listener is placed, and the direction information for the plurality of sound objects is associated with the position or reference position or orientation of the virtual listener.

13. Apparatus according to any one of claims 10 to 12.

14. a parameter processor (110) for quantizing the metadata indicating the directional information for the plurality of audio objects to obtain quantized directional terms for the plurality of audio objects; the downmixer (400) is configured to operate in response to the quantized directional terms as the directional information; the output interface (200) is configured to introduce information about the quantized directional items into the encoded audio signal; 14. Apparatus according to any one of claims 10 to 13.

15. the downmixer (400) is configured to perform (410) an analysis of the directional information for the plurality of audio objects and to position (412) one or more virtual microphones to generate the transport channels depending on the results of the analysis.

15. Apparatus according to any one of claims 10 to 14.

16. the downmixer (400) is configured to downmix (408) using static downmix rules across the plurality of time periods; or the directional information is variable across the plurality of time periods, and the downmixer (400) is configured to downmix (405) using a downmixing rule that is variable across the plurality of time periods.

16. Apparatus according to any one of claims 10 to 15.

17. 17. The apparatus of claim 10, wherein the downmixer (400) is configured to downmix in the time domain using sample-by-sample weighting and combining samples of the multiple audio objects.

18. 1. A decoder for decoding an encoded audio signal comprising one or more transport channels, directional information of a plurality of audio objects, and parameter data of at least two associated audio objects for one or more frequency bins of a time frame, wherein the number of the at least two associated audio objects is less than a total number of the plurality of audio objects, the decoder comprising: an input interface (600) for providing the one or more transport channels in a spectral representation having a plurality of frequency bins within the time frame; an audio renderer (700) for using the directional information to render the one or more transport channels into a plurality of audio channels such that contributions from the one or more transport channels are taken into account according to first directional information associated with a first one of the at least two associated audio objects and according to second directional information associated with a second one of the at least two associated audio objects; Equipped with the audio renderer (700) is configured to calculate, for each of the one or more frequency bins, contributions from the one or more transport channels according to first directional information associated with a first of the at least two associated audio objects and according to second directional information associated with a second of the at least two associated audio objects. decoder.

19. the audio renderer (700) is configured to ignore, for the one or more frequency bins, directional information of an audio object that is different from the at least two related audio objects; 19. A decoder according to claim 18.

20. the encoded audio signal includes amplitude-related measures (812) for each associated audio object or a combined value (812) associated with at least two associated audio objects in the parameter data; the audio renderer (700) is configured to determine (704) quantitative contributions of the one or more transport channels according to the amplitude-related measurements or the combined values; Decoder according to claim 18 or 19.

21. the encoded signal includes the combined values ​​in the parameter data; the audio renderer (700) is configured to determine (704, 733) the contributions of the one or more transport channels using the combined value for one of the associated audio objects and the directional information for the one associated audio object; the audio renderer (700) is configured to determine (704, 735) the contribution of the one or more transport channels using a value derived from the combined value of another value of the associated audio object in the one or more frequency bins and the directional information of other of the associated audio objects.

21. A decoder according to claim 20.

22. The audio renderer (700) calculating (704) for each frequency bin of the plurality of frequency bins the direct response information from the associated sound object and the directional information associated with the associated sound object within the frequency bin; 22. A decoder according to any one of claims 18 to 21, configured to:

23. the audio renderer (700) determines (741) a diffuse signal for each frequency bin of the plurality of frequency bins using diffusion information, such as diffusion parameters or decorrelation rules, included in the metadata, and combines the direct responses to obtain a signal determined by the direct response information and the diffuse signal and rendered in the spectral domain of a channel of the plurality of channels; or using the direct response information (704) and information regarding the number of audio channels (702) to calculate (706) synthesis information and applying (727) the covariance synthesis information to the one or more transport channels to obtain the number of audio channels; the direct response information (704) is a direct response vector for each associated audio object, the covariance synthesis information is a covariance synthesis matrix, and the audio renderer (700) is configured to perform a matrix operation for each frequency bin when applying (727) the covariance synthesis information.

23. A decoder according to claim 22.

24. The audio renderer (700): said computing said direct response information (704) comprising deriving a direct response vector for each associated audio object; computing a covariance matrix from each direct response vector for each associated audio object; In said calculating said covariance synthesis information, the covariance matrix from each of the associated audio objects; power information for each of said associated audio objects; power information derived from the one or more transport channels; Derive target covariance information from (724); 24. A decoder according to claim 22 or 23, configured to:

25. The audio renderer (700): In said calculating (704) said direct response information, deriving a direct response vector for each associated audio object, and calculating (723) a covariance matrix from each direct response vector for each associated audio object; deriving (726) input covariance information from the transport channel; deriving (725a, 725b) mixing information from the target covariance information, the input covariance information, and the information regarding the number of channels; applying the mixing information to the transport channels of each frequency bin within the time period (727); 25. A decoder according to claim 24, configured to:

26. 26. The decoder of claim 25, wherein a result of the application of the mixing information to each frequency bin in the time window is transformed (708) into a time domain to obtain a number of audio channels in the time domain.

27. The audio renderer (700): In the decomposition (752) of the input covariance matrix, only the main diagonal elements of the input covariance matrix derived from the transport channel are used; performing a decomposition (751) of the target covariance matrix using the direct response matrix and the power matrix of the object or transport channel; performing (752) a decomposition of the input covariance matrix by taking the roots of each main diagonal element of the input covariance matrix; Compute the regularized inverse of the decomposed input covariance matrix (753), performing singular value decomposition to calculate the optimal matrix used for energy compensation without the extended identity matrix (756); 27. A decoder according to any one of claims 22 to 26, configured to:

28. 1. A method for encoding a plurality of audio objects and associated metadata indicating directional information for said plurality of audio objects, comprising: downmixing the plurality of audio objects to obtain one or more transport channels; encoding the one or more transport channels to obtain one or more encoded transport channels; outputting an encoded audio signal comprising the one or more encoded transport channels; Including, the downmixing step includes downmixing the plurality of audio objects in response to the directional information relating to the plurality of audio objects. method.

29. 1. A method of decoding an encoded audio signal, the encoded audio signal including one or more transport channel and direction information for a plurality of audio objects and parameter data for at least two associated audio objects for one or more frequency bins of a time frame, wherein the number of the at least two associated audio objects is less than a total number of the plurality of objects, the decoding method comprising: providing the one or more transport channels in a spectral representation having a plurality of frequency bins within the time window; audio rendering the one or more transport channels into a plurality of audio channels using the direction information; Including, and wherein the step of rendering the audio comprises calculating, for each of the one or more frequency bins, contributions from the one or more transport channels are taken into account according to first directional information associated with a first of the at least two associated audio objects and according to second directional information associated with a second of the at least two associated audio objects, or according to first directional information associated with the first of the at least two associated audio objects and according to second directional information associated with the second of the at least two associated audio objects. method.

30. 30. A computer program for performing the method of claim 28 or the method of claim 29 when the computer program is run on a computer or processor.

31. An encoded audio signal comprising information about parameter data of at least two audio objects associated with one or more frequency bins.

32. one or more encoded transport channels; the information relating to the parameter data including, for each of the one or more frequency bins of the plurality of frequency bins within a time window, two or more encoded object identifications of the associated audio object and one or more encoded combined values ​​or encoded amplitude-related measures; quantized and encoded directional data for each audio object within the time window, the directional data being constant for all frequency bins of the one or more frequency bins; 32. The encoded audio signal of claim 31, further comprising:

Citation Information

Patent Citations

  • Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to dirac based spatial audio coding

    WO2019068638A1

  • Parameter encoding and decoding

    WO2020249815A2