Apparatus and methods for decoding encoded audio signals

CN122575383APending Publication Date: 2026-08-14FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2021-10-12
Publication Date
2026-08-14

Smart Images

  • Figure CN122575383A_ABST
    Figure CN122575383A_ABST
Patent Text Reader

Abstract

An apparatus for encoding a plurality of audio objects includes: an object parameter calculator (100) configured to: calculate parameter data of at least two related audio objects for one or more frequency intervals of a plurality of frequency intervals associated with a time frame, wherein the number of the at least two related audio objects is less than the total number of the plurality of audio objects; and an output interface (200) configured to output an encoded audio signal including information about the parameter data of the at least two related audio objects in the one or more frequency intervals.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of PCT international application PCT / EP2021 / 078217 filed on October 12, 2021, entitled “Apparatus and method for encoding a plurality of audio objects and apparatus and method for decoding using two or more related audio objects”, which entered the Chinese national phase patent application 202180076553.3. Technical Field

[0002] This invention relates to the encoding of audio signals (e.g., audio objects) and the decoding of encoded audio signals (e.g., encoded audio objects). Background Technology

[0003] introduction

[0004] This document describes a parameterized approach for encoding and decoding object-based audio content at low bit rates using Directional Audio Coding (DirAC). The presented embodiments are used as part of the 3GPP Immersive Speech and Audio Services (IVAS) codec and provide a favorable alternative to the low-bit-rate Independent Stream with Metadata (ISM) mode (a discrete coding approach). Existing technology

[0005] Discrete encoding of objects

[0006] The most straightforward method for encoding object-based audio content is to encode each object separately and send it along with its corresponding metadata. The main drawback of this method is that the bit consumption required to encode each object becomes excessive as the number of objects increases. A simple solution to this problem is to use a "parametric approach," where some relevant parameters are calculated based on the input signal and quantized and sent together with a suitable downmixed signal combining several object waveforms.

[0007] Spatial Audio Object Coding (SAOC)

[0008] Spatial Audio Object Coding [SAOC_STD, SAOC_AES] is a parameterized method in which the encoder is based on a certain downmixing matrix. The downmixed signal is calculated using a set of parameters, and both are sent to the decoder. These parameters represent the psychoacoustic properties and relationships of all individual objects. At the decoder, a rendering matrix is ​​used. Render the downmix to a specific speaker layout.

[0009] The main parameter of SAOC is the object covariance matrix of size N * N. Where N refers to the number of objects. This parameter is transmitted to the decoder as object-level difference (OLD) and optional inter-object covariance (IOC).

[0010] matrix Each element It is given by the following formula:

[0011]

[0012] Object-level differences (OLD) are defined as

[0013]

[0014] in, Absolute object energy (NRG) is described as

[0015]

[0016] as well as

[0017]

[0018] Where i and j are objects and The object index is defined by n, where n is the time index and k is the frequency index. l indicates the time index set, and m indicates the frequency index set. ε is an additional constant to avoid division by zero, e.g., ε = 10.

[0019] The similarity measure of input objects (IOCs) can be given, for example, by cross-correlation:

[0020]

[0021] A downmixing matrix of size N_dmx*N From elements Let be the definition, where i refers to the channel index of the downmix signal, and j refers to the object index. For stereo downmix (N_dmx=2). Based on the parameters DMG and DCLD, it is calculated as follows:

[0022]

[0023]

[0024] in, and It is given by the following formula:

[0025]

[0026]

[0027] For the mono downmixing case (N_dmx=1), Calculated solely based on DMG parameters

[0028]

[0029] in,

[0030]

[0031] Spatial Audio Object Encoding-3D (SAOC-3D)

[0032] Spatial Audio Object Coding 3D Audio Reproduction (SAOC-3D) [MPEGH_AES, MPEGH_IEEE, MPEGH_STD, SAOC_3D_PAT] is an extension of the aforementioned MPEG SAOC technology. MPEG SAOC technology compresses and renders channel and object signals in a very efficient bit rate manner.

[0033] The main differences from SAOC are:

[0034] While the original SAOC only supports up to two downmixing channels, SAOC-3D can map multi-object inputs to any number of downmixing channels (and associated auxiliary information).

[0035] - Direct rendering to multichannel output, compared to the classic SAOC which already uses MPEG surround sound as a multichannel output processor.

[0036] - Some tools, such as residual coding tools, have been discarded.

[0037] Despite these differences, from a parameter perspective, SAOC-3D is identical to SAOC. The SAOC-3D decoder—similar to the SAOC decoder—receives multi-channel downmixing X, covariance matrix E, rendering matrix R, and downmixing matrix D.

[0038] The rendering matrix R is defined by the input channels and the input objects, and is received from the format converter (channels) and the object renderer (objects), respectively.

[0039] The undermixing matrix D consists of elements Let i be the channel index of the downmixer signal, and j be the object index, calculated based on the downmixer gain (DMG):

[0040]

[0041] in,

[0042]

[0043] The output covariance matrix C of size N_out*N_out is defined as:

[0044]

[0045] Related solutions

[0046] Several other schemes exist that are essentially similar to SAOC but have the following subtle differences:

[0047] - Binaural cue coding (BCC) for objects has been described, for example, in [BCC2001] and is a precursor to SAOC technology.

[0048] - Union Object Coding (JOC) and Advanced Union Object Coding (A-JOC) perform similar functions to SAOC, while providing largely separated objects on the decoder side without rendering them to a specific output speaker layout [JOC_AES, AC4_AES]. This technique sends elements of the upmixing matrix from the downmixing as parameters to the separated objects (instead of OLD).

[0049] Directional Audio Coding (DirAC)

[0050] Another parameterization approach is directional audio coding. DirAC [Pulkki2009] is a perception-driven representation of spatial sound. The assumption is that, in a temporal instance and for a critical frequency band, the spatial resolution of the human auditory system is limited to decoding one directional cue and another interaural coherence cue.

[0051] Based on these assumptions, DirAC represents spatial sound in a frequency band through two streams: a non-directional diffuse stream and a directional non-diffuse stream. DirAC processing is performed in two stages: analysis and synthesis, such as... Figure 12a and Figure 12b As shown.

[0052] During the DirAC analysis phase, a first-order coincident microphone in B format is considered as input, and the sound dispersion and direction of arrival are analyzed in the frequency domain.

[0053] In the DirAC synthesizer, sound is divided into two streams: a non-diffuse stream and a diffuse stream. The non-diffuse stream is reproduced as a point source using amplitude panning, which can be done using vector basis amplitude panning (VBAP) [Pulkki 1997]. The diffuse stream is responsible for creating a sense of envelopment and is generated by sending mutually decorrelated signals to the speakers.

[0054] Figure 12aThe analysis stage includes a band filter 1000, an energy estimator 1001, an intensity estimator 1002, time averaging elements 999a and 999b, a diffusivity calculator 1003, and a direction calculator 1004. The calculated spatial parameters are diffusivity values ​​between 0 and 1 for each time / frequency zone and direction of arrival parameters for each time / frequency zone generated by block 1004. Figure 12a In this context, the directional parameters include azimuth and elevation, which indicate the direction of arrival of the sound relative to a reference or listening position and specifically relative to the location of the microphone from which the four-component signal is collected and input to a filter 1000. Figure 12a In this context, these component signals are first-order surround sound components, which include an omnidirectional component W, a directional component X, another directional component Y, and yet another directional component Z.

[0055] Figure 12b The illustrated DirAC synthesis stage includes a bandpass filter 1005 for generating time / frequency representations of B-format microphone signals W, X, Y, and Z. The corresponding signals for each time / frequency zone are input to a virtual microphone stage 1006, which generates a virtual microphone signal for each channel. Specifically, to generate, for example, a virtual microphone signal for the center channel, the virtual microphone is pointed in the direction of the center channel, and the resulting signal is the corresponding component signal of the center channel. This signal is then processed via a direct signal branch 1015 and a spread signal branch 1014. Both branches include corresponding gain adjusters or amplifiers, which are controlled in blocks 1007 and 1008 by spread values ​​derived from the original spread parameters, and are further processed in blocks 1009 and 1010 to obtain certain microphone compensation.

[0056] Gain adjustment is also applied to the component signals in the direct signal branch 1015 using gain parameters derived from direction parameters consisting of azimuth and elevation angles. Specifically, these angles are input into the VBAP (Vector Basis Amplitude Shift) gain table 1011. For each channel, the results are input to a speaker gain averaging stage 1012 and another normalizer 1013, and the resulting gain parameters are then forwarded to the amplifier or gain adjuster in the direct signal branch 1015. In combiner 1017, the diffused signal and the direct signal or non-diffuse stream generated at the output of decorrelector 1016 are combined, and then the other subbands are added in another combiner 1018, which may be, for example, a synthesis filter bank. Thus, a speaker signal is generated for a particular speaker, and the same process is performed for the other channels of other speakers 1019 in a particular speaker setup.

[0057] High-quality versions of DirAC synthesis, such as Figure 12bAs shown, the synthesizer receives all B-format signals and calculates a virtual microphone signal for each speaker direction based on these B-format signals. The radiation pattern used is typically a dipole. The virtual microphone signal is then modified in a non-linear manner, depending on the metadata discussed in branches 1016 and 1015. Figure 12b The low-bitrate version of DirAC is not shown. However, in this low-bitrate version, a single audio channel is sent. The difference in processing is that all virtual microphone signals are replaced by the single received audio channel. The virtual microphone signals are divided into two streams, a diffuse stream and a non-diffuse stream, which are processed separately. The non-diffuse sound is reproduced as a point source using Vector Basis Amplitude Translation (VBAP). During translation, the mono audio signal is applied to a subset of speakers after multiplying the mono audio signal by a speaker-specific gain factor. The gain factor is calculated using speaker setup information and the specified translation direction. In the low-bitrate version, the input signal is simply translated to the direction implied by the metadata. In the high-quality version, each virtual microphone signal is multiplied by its corresponding gain factor, which produces the same effect as translation, but it is less prone to any non-linear artifacts.

[0058] The goal of diffuse sound synthesis is to create a sound perception around the listener. In low bitrate versions, the diffuse stream is reproduced by decorrelating the input signal and reproducing it from each speaker. In high-quality versions, the virtual microphone signals of the diffuse stream are already somewhat incoherent, and only slight decorrelation processing is required.

[0059] DirAC parameters (also known as spatial metadata) consist of tuples of diffusion and orientation, which are represented in spherical coordinates by two angles (i.e., azimuth and elevation). If both the analysis and synthesis stages operate on the decoder side, the time-frequency resolution of the DirAC parameters can be chosen to be the same as the filter bank used for DirAC analysis and synthesis—that is, a different set of parameters for each time slot and frequency interval represented by the filter bank used for the audio signal.

[0060] Some work has been done to reduce the size of metadata so that the DirAC paradigm can be used in spatial audio coding and teleconference scenarios [Hirvonen2009].

[0061] [WO2019068638] introduces a general-purpose spatial audio coding system based on DirAC. Compared to the classic DirAC designed for B-format (first-order surround sound) input, this system can accept first-order or higher-order surround sound, multi-channel or object-based audio input, and also allows mixed-type input signals. All signal types are efficiently encoded and transmitted individually or in combination. The former combines different representations at the renderer (decoder side), while the latter combines different audio representations at the encoder side in the DirAC domain.

[0062] Compatibility with the DirAC framework

[0063] This embodiment builds upon a unified framework for arbitrary input types as proposed in [WO2019068638], and—similar to what [WO2020249815] did for multi-channel content—aims to eliminate the problem of not being able to efficiently apply DirAC parameters (direction and spread) to object input. In fact, the spread parameter is not needed at all, but it has been found that a single direction cue per time / frequency unit is insufficient to reproduce high-quality object content. Therefore, this embodiment proposes using multiple direction cues per time / frequency unit and accordingly introduces an adapted set of parameters that replaces the classic DirAC parameters in the case of object input.

[0064] Flexible systems with low bit rates

[0065] Compared to DirAC, which uses scene-based representations from the listener's perspective, SAOC and SAOC-3D are designed for channel-based and object-based content, where parameters describe the relationship between channels / objects. To use scene-based representations for object input and thus be compatible with the DirAC renderer, while ensuring efficient representation and high-quality reproduction, an adapted set of parameters is needed to also allow for the signaling of multiple directional cues.

[0066] A key objective of this embodiment is to find a method for efficiently encoding object input at a low bit rate that also exhibits good scalability for an increasing number of objects. Discretely encoding the signal for each object does not provide this scalability: each additional object significantly increases the overall bit rate. If the number of objects exceeds the allowed bit rate, this directly leads to a very noticeable degradation of the output signal; this degradation is another argument supporting this embodiment. Summary of the Invention

[0067] The purpose of this invention is to provide an improved concept for encoding multiple audio objects or decoding encoded audio signals.

[0068] This objective is achieved by means of an encoding apparatus, decoder, encoding method, decoding method, computer program, or encoded audio signal as described in embodiments of the present invention.

[0069] In one aspect of the invention, the invention is based on the discovery that for one or more frequency ranges among a plurality of frequency ranges, at least two related audio objects are defined, and parameter data associated with these at least two related objects are included on the encoder side and used on the decoder side to obtain a high-quality and efficient audio encoding / decoding concept.

[0070] According to another aspect of the invention, the invention is based on the discovery that performing specific downmixing suitable for the directional information associated with each object, such that each object having associated directional information valid for the entire object (i.e., for all frequency ranges in a time frame) is used to downmix that object into multiple transmission channels. The use of directional information is, for example, equivalent to generating transmission channels as virtual microphone signals with certain adjustable characteristics.

[0071] On the decoder side, specific synthesis dependent on covariance synthesis is performed, which in certain embodiments is particularly suitable for high-quality covariance synthesis unaffected by artifacts introduced by the decorrelator. In other embodiments, advanced covariance synthesis relying on specific improvements associated with standard covariance synthesis is used to improve audio quality and / or reduce the computational cost required to compute the mixing matrix used within the covariance synthesis.

[0072] However, even in more classical synthesis where audio rendering is performed by explicitly determining individual contributions within time / frequency intervals based on the sent selection information, the audio quality is superior to existing object coding or channel downmixing methods. In this case, each time / frequency interval has object identification information, and when performing audio rendering—that is, when considering the directional contribution of each object—the object identifier is used to find the direction associated with that object information in order to determine the gain value of each output channel for each time / frequency interval. Therefore, when only a single associated object exists in a time / frequency interval, the gain value of that single object is determined for each time / frequency interval based on a "codebook" of the object ID and the directional information of the associated object.

[0073] However, when there are more than one relevant object in a time / frequency interval, the gain value for each relevant object is calculated to assign the corresponding time / frequency interval of the transmission channel to the corresponding output channel controlled by the user-provided output format (e.g., a channel format is stereo, 5.1, etc.). Regardless of whether the gain value is used for covariance synthesis (i.e., for applying a mixing matrix to mix the transmission channel into the output channel) or whether the gain value is used to explicitly determine the individual contribution of each object in the time / frequency interval (possibly enhanced by adding a spread signal component) by multiplying the gain value by the corresponding time / frequency interval of one or more transmission channels and then summing the contributions of each output channel in the corresponding time / frequency interval, the output audio quality is still enhanced due to the flexibility provided by determining one or more relevant objects in each frequency interval.

[0074] This determination can be highly efficient because only one or more object IDs for the time / frequency range must be encoded and sent to the decoder along with the orientation information for each object. However, it can also be highly efficient because, for a frame, only a single orientation information exists for all frequency ranges.

[0075] Therefore, regardless of whether the synthesis is performed using preferred enhanced covariance synthesis or a combination of explicit transmission channel contributions for each object, efficient and high-quality object downmixing is achieved, which is preferably enhanced by using object-specific orientation-dependent downmixing that depends on downmixing weights that reflect the generated transmission channel as a virtual microphone signal.

[0076] The aspects associated with two or more related objects for each time / frequency interval can preferably be combined with aspects of performing object-to-transmission channel-specific correlation downmixing. However, these two aspects can also be applied independently of each other. Furthermore, although covariance synthesis of two or more related objects for each time / frequency interval is performed in some embodiments, advanced covariance synthesis and advanced transmission channel-to-output channel upmixing can also be performed by sending only a single object identifier for each time / frequency interval.

[0077] Furthermore, regardless of whether there are single or multiple related objects in each time / frequency interval, upmixing can be performed by calculating the mixing matrix in standard or enhanced covariance synthesis, or by determining the gain value of the corresponding contribution for each time / frequency interval individually (based on object identifiers used to obtain specific directional information from the directional "codebook"). These are then summed to obtain the full contribution for each time / frequency interval if two or more related objects are present. The output of this summation step is then equivalent to the output of the mixing matrix application, and the final filter bank processing is performed to generate the time-domain output channel signal in the corresponding output format. Attached Figure Description

[0078] Preferred embodiments of the invention will then be described with reference to the accompanying drawings, in which:

[0079] Figure 1a It is an implementation of an audio encoder based on the first aspect of having at least two related objects in each time / frequency interval;

[0080] Figure 1b It is based on the implementation of the second aspect of the encoder with direction-dependent object downmixing;

[0081] Figure 2 It is a preferred implementation of the encoder based on the second aspect;

[0082] Figure 3 It is a preferred implementation of the encoder based on the first aspect;

[0083] Figure 4 It is a preferred implementation of the decoder based on the first and second aspects;

[0084] Figure 5 yes Figure 4 Optimal implementation of covariance synthesis processing;

[0085] Figure 6a It is based on the implementation of the decoder in the first aspect;

[0086] Figure 6b It is based on the decoder of the second aspect;

[0087] Figure 7a It is a flowchart used to illustrate the determination of parameter information based on the first aspect;

[0088] Figure 7b This further determines the optimal implementation of parameterized data;

[0089] Figure 8a The time / frequency representation of the high-resolution filter bank is shown;

[0090] Figure 8bThe transmission of relevant auxiliary information for frame J is shown according to a preferred implementation of the first and second aspects;

[0091] Figure 8c The “direction codebook” included in the coded audio signal is shown;

[0092] Figure 9a The preferred encoding method according to the second aspect is shown;

[0093] Figure 9b The implementation of static undermixing according to the second aspect is shown;

[0094] Figure 9c The implementation of dynamic undermixing according to the second aspect is shown;

[0095] Figure 9d Another embodiment of the second aspect is shown;

[0096] Figure 10a A flowchart illustrating a preferred implementation of the decoder side of the first aspect is shown;

[0097] Figure 10b An example of summing based on the contribution to each output channel is shown. Figure 10a The optimal implementation of output channel calculation;

[0098] Figure 10c A preferred method for determining power values ​​for multiple relevant objects according to the first aspect is shown;

[0099] Figure 10d This demonstrates the use of covariance synthesis, which depends on the calculation and application of the mixture matrix, to compute... Figure 10a An example of the output channel;

[0100] Figure 11 Several embodiments for high-level computation of hybrid matrices over time / frequency intervals are shown;

[0101] Figure 12a The prior art DirAC encoder is shown; and

[0102] Figure 12b The existing DirAC decoder is shown. Detailed Implementation

[0103] Figure 1aAn apparatus for encoding multiple audio objects is illustrated, which receives, at input, the audio objects as is and / or metadata of the audio objects. The encoder includes an object parameter calculator 100 that provides parameter data for at least two relevant audio objects for a time / frequency interval and forwards this data to an output interface 200. Specifically, the object parameter calculator calculates parameter data for at least two relevant audio objects for one or more frequency intervals associated with a time frame, wherein, specifically, the number of at least two relevant audio objects is less than the total number of the multiple audio objects. Therefore, the object parameter calculator 100 actually performs a selection, rather than simply indicating that all objects are relevant. In a preferred embodiment, this selection is performed by correlation, and the correlation is determined by amplitude correlation measurements (e.g., amplitude, power, loudness, or another measurement obtained by increasing the amplitude to a power other than 1 and preferably greater than 1). Then, if a certain number of relevant objects are available for the time / frequency interval, the object with the most relevant characteristics (i.e., the object with the highest power among all objects) is selected, and data about these selected objects is included in the parameter data.

[0104] Output interface 200 is configured to output an encoded audio signal that includes information about parameter data for at least two related audio objects across one or more frequency ranges. Depending on the implementation, the output interface may receive other data (e.g., object downmixing, or one or more transport channels representing object downmixing, or additional parameters, or object waveform data in a mixed representation of several downmixed objects, or other objects in individual representations) and input them into the encoded audio signal. In this case, the objects are directly introduced or "copied" into the corresponding transport channels.

[0105] Figure 1bA preferred implementation of an apparatus for encoding multiple audio objects according to the second aspect is shown, wherein audio objects and associated object metadata are received, the associated object metadata indicating directional information about the multiple audio objects, i.e., directional information for each object or group of objects (if the group of objects has the same directional information associated with it). The audio objects are input to a downmixer 400 to downmix the multiple audio objects to obtain one or more transport channels. Furthermore, a transport channel encoder 300 is provided, which encodes the one or more transport channels to obtain one or more encoded transport channels, which are then input to an output interface 200. Specifically, the downmixer 400 is connected to an object directional information provider 110, which receives at its input any data from which object metadata can be derived and outputs directional information actually used by the downmixer 400. The directional information forwarded from the object directional information provider 110 to the downmixer 400 is preferably dequantized directional information, i.e., the same directional information then available on the decoder side. To this end, the object orientation information provider 110 is configured to export, extract, or acquire unquantized object metadata, and then quantize the object metadata to derive a representation. Figure 1b The quantized object metadata in the "Other Data" section is provided to the output interface 200 in a preferred embodiment. Furthermore, the object orientation information provider 110 is configured to dequantize the quantized object orientation information to obtain the actual orientation information forwarded from block 110 to downmixer 400.

[0106] Preferably, the output interface 200 is configured to additionally receive parameter data of the audio object, object waveform data, one or more identifiers of a single or multiple related objects for each time / frequency interval, and quantization direction data as discussed previously.

[0107] Subsequently, other embodiments are shown. A parameterization method for encoding audio object signals is proposed, which allows for efficient transmission at low bit rates and high-quality reproduction on the consumer side. Based on the DirAC principle, which considers a directional cue for each key frequency band and time instance (time / frequency zone), the most dominant object is determined for each such time / frequency zone of the time / frequency representation of the input signal. Since this proves insufficient for the object input, an additional, second most dominant object is determined for each time / frequency zone, and based on these two objects, a power ratio is calculated to determine the impact of each of the two objects on the time / frequency zone considered. Notice:It is conceivable to consider more than two primary objects for each time / frequency unit, especially with an increasing number of input objects. For simplicity, the following description is primarily based on two primary objects per time / frequency unit.

[0108] Therefore, the parameterized auxiliary information sent to the decoder includes:

[0109] - Power ratio calculated for the relevant (primary) subset of objects for each time / frequency zone (or parameter band).

[0110] - An object index representing a subset of relevant objects for each time / frequency zone (or parameter band).

[0111] - Directional information associated with the object index and provided for each frame (where each time-domain frame includes multiple parameter bands, and each parameter band includes multiple time / frequency zones).

[0112] Directional information can be obtained via an input metadata file associated with the audio object signal. For example, metadata can be specified based on frames. In addition to auxiliary information, a downmixed signal combining the input object signal is also sent to the decoder.

[0113] During the rendering phase, the transmitted direction information (derived via the object index) is used to translate the transmitted downmix signal (or more generally: the transmission channel) to the appropriate direction. The downmix signal is assigned to two relevant object directions based on the transmitted power ratio, which is used as a weighting factor. This processing is performed for each time / frequency zone of the time / frequency representation of the decoded downmix signal.

[0114] This section summarizes the encoder-side processing and then describes the parameters and downmixing calculations in detail. The audio encoder receives one or more audio object signals. For each audio object signal, a metadata file describing the object's attributes is associated. In this embodiment, the object attributes described in the associated metadata file correspond to direction information provided based on frames, where one frame corresponds to 20 milliseconds. Each frame is identified by a frame number, which is also included in the metadata file. The direction information is given in the form of azimuth and elevation information, where the azimuth is taken from (-180, 180) degrees, and the elevation is taken from [-90, 90] degrees. Other attributes provided in the metadata may include, for example, distance, propagation, and gain; these attributes are not considered in this embodiment.

[0115] The information provided in the metadata file, along with the actual audio object files, is used to create a parameter set, which is sent to the decoder for rendering the final audio output file. More specifically, the encoder estimates the parameters, or power ratios, of the principal object subset for each given time / frequency region. The principal object subset is represented by an object index, which also identifies the object orientation. These parameters, along with transport channel and orientation metadata, are sent to the decoder.

[0116] Figure 2 An overview of the encoder is provided, wherein the transmission channels include a downmixed signal calculated based on the input object file and direction information provided in the input metadata. The number of transmission channels is always less than the number of input object files. In the encoder of this embodiment, the encoded audio signal is represented by the encoded transmission channels, and the encoded parameterization auxiliary information is indicated by the encoded object index, the encoded power ratio, and the encoded direction information. The encoded transmission channels and the encoded parameterization auxiliary information together form a bitstream output by the multiplexer 220. Specifically, the encoder includes a filter bank 102 that receives the input object audio file. Furthermore, the object metadata file is provided to the extractor direction information block 110a. The output of block 110a is input to the quantization direction information block 110b, which outputs the direction information to the downmixer 400 that performs the downmixing calculation. In addition, the quantized direction information (i.e., the quantization index) is forwarded from block 110b to the encoded direction information block 202, which preferably performs some kind of entropy coding to further reduce the required bit rate.

[0117] Furthermore, the output of filter bank 102 is input to signal power calculation block 104, and the output of signal power calculation block 104 is input to object selection block 106 and additionally to power ratio calculation block 108. Power ratio calculation block 108 is also connected to object selection block 106 to calculate the power ratio (i.e., a combined value for only the selected object). In block 210, the calculated power ratio or combined value is quantized and encoded. As will be outlined later, the power ratio is preferred to preserve the transmission of a power data item. However, in other embodiments where such preservation is not required, instead of the power ratio, the actual signal power or other values ​​derived from the signal power determined by block 104 can be input to the quantizer and encoder at the selection of object selector 106. Then, power ratio calculation 108 is not required, and object selection 106 ensures that only relevant parameterized data (i.e., power-related data of the relevant object) is input to block 210 for quantization and encoding purposes.

[0118] Will Figure 1a and Figure 2 For comparison, blocks 102, 104, 110a, 110b, 106, and 108 are preferably included. Figure 1a The object parameter calculator 100 is included, and blocks 202, 210, and 220 are preferably included. Figure 1a The output interface block 200.

[0119] also, Figure 2 The core encoder 300 in the middle corresponds to Figure 1b The transmission channel encoder 300, the downmixing calculation block 400 corresponds to Figure 1b The downmixer 400, and Figure 1b The object orientation information provider 110 corresponds to Figure 2 Blocks 110a and 110b. Furthermore... Figure 1b The output interface 200 is preferably connected to... Figure 1a The output interface 200 is implemented in the same way and includes Figure 2 Blocks 202, 210, and 220.

[0120] Figure 3 A variant of the encoder is shown where downmixing is optional and independent of the input metadata. In this variant, input audio files can be fed directly into the core encoder, which creates transport channels from them; therefore, the number of transport channels corresponds to the number of input object files. This is particularly interesting if the number of input objects is one or two. For a large number of objects, the downmixing signal will still be used to reduce the amount of data to be transmitted.

[0121] exist Figure 3 In the figures, similar reference numerals refer to Figure 2 Similar functionality. This is not only for... Figure 2 and Figure 3 It is valid, and also valid for all other figures described in this specification. Figure 2 different, Figure 3 Downmixing calculation 400 is performed without any directional information. Therefore, the downmixing calculation can be, for example, static downmixing using a pre-known downmixing matrix, or energy-dependent downmixing that does not depend on any directional information associated with the object included in the input object audio file. However, directional information is extracted in block 110a and quantized in block 110b, and the quantized value is forwarded to the directional information encoder 202 to have encoded directional information in the encoded audio signal (e.g., a binary encoded audio signal) forming the bitstream.

[0122] When there are a limited number of input audio object files or sufficient available transmission bandwidth, the downmixing calculation block 400 can be omitted, allowing the input audio object files to directly represent the transport channels encoded by the core encoder. In this implementation, blocks 104, 104, 106, 108, and 210 are also unnecessary. However, a preferred implementation results in a mixing implementation where some objects are directly introduced into the transport channels, while others are downmixed into one or more transport channels. In this case, then... Figure 3 All the blocks shown will be necessary to directly generate a bitstream having one or more objects within the encoded transport channel, and by... Figure 2 or Figure 3 The downmixer 400 generates one or more transmission channels.

[0123] Parameter calculation

[0124] A filter bank is used to transform the time-domain audio signal, which includes all input object signals, into the time / frequency domain. For example, the CLDFB (Complex Low-Delay Filter Bank) analysis filter transforms a 20-millisecond frame (corresponding to 960 samples at a sampling rate of 48kHz) into a 16x60 time / frequency unit with 16 time slots and 60 frequency bands. For each time / frequency unit, the instantaneous signal power is calculated as:

[0125]

[0126] Here, k represents the band index, n represents the slot index, and i represents the object index. Since the transmission parameters for each time / frequency band are very costly in terms of the final bit rate, grouping is employed to compute parameters for a reduced number of time / frequency bands. For example, 16 slots can be combined into a single slot, and 60 bands can be combined into 11 bands based on a psychoacoustic scale. This reduces the initial size of 16x60 to 1x11, corresponding to 11 so-called parameter bands. The instantaneous signal power values ​​are summed based on the grouping to obtain the signal power with reduced dimensionality:

[0127]

[0128] Where T corresponds to 15 in this example, and and Define parameters with boundaries.

[0129] To determine the subset of the most principal objects for its computational parameters, the instantaneous signal power values ​​of all N input audio objects are sorted in descending order. In this embodiment, we identify two most principal objects, and the corresponding object indices, ranging from 0 to N-1, are stored as part of the parameters to be sent. Furthermore, the power ratio that correlates the signals of the two principal objects is calculated:

[0130]

[0131]

[0132] Or, in a more general way, not limited to two objects:

[0133]

[0134] In this context, Indicates the number of main objects to be considered, and:

[0135]

[0136] In the case of two primary objects, for each of the two objects, a power ratio of 0.5 means that both objects exist within the corresponding parameter band, while power ratios of 1 and 0 indicate that one of the two objects does not exist. These power ratios are stored as the second part of the parameters to be sent. Since the sum of the power ratios is 1, the transmission... One rather than One value is enough.

[0137] In addition to the object index and power ratio value for each parameter band, orientation information for each object extracted from the input metadata file must also be sent. Since the information is initially provided frame-based, this is done per frame (where each frame in the described example includes 11 parameter bands or a total of 16x60 time / frequency zones). Therefore, the object index indirectly indicates the object orientation. Note: Since the sum of the power ratios is 1, the number of power ratios to be sent per parameter band can be reduced by 1; for example, considering 2 related objects, sending only 1 power ratio value is sufficient.

[0138] Both the direction information and the power ratio are quantized and combined with the object index to form parameterized auxiliary information. This parameterized auxiliary information is then encoded and mixed with the encoded transmission channel / downmixer signal into the final bitstream representation. For example, a good trade-off between output quality and extended bit rate is achieved by quantizing the power ratio using 3 bits per value. The direction information can be provided with an angular resolution of 5 degrees and subsequently quantized with 7 bits per azimuth value and 6 bits per elevation value, to give a practical example.

[0139] Lower mixing calculation

[0140] All input audio object signals are combined into a downmixed signal, which includes one or more transmission channels, wherein the number of transmission channels is less than the number of input object signals. Note: In this embodiment, if there is only one input object, only a single transmission channel exists, which therefore means that the downmixing calculation is skipped.

[0141] If the downmixer includes two transmission channels, the stereo downmixer can be calculated, for example, as a virtual cardioid microphone signal. The virtual cardioid microphone signal is determined by applying direction information provided for each frame to the metadata file (here, it is assumed that all elevation angle values ​​are zero):

[0142]

[0143]

[0144] Here, the virtual cardioid is located at 90° and -90°. Therefore, individual weights are determined for each of the two transmission channels (left and right), and these individual weights are applied to the corresponding audio object signals:

[0145]

[0146]

[0147] In this context, This is the number of input objects greater than or equal to 2. If the virtual cardioid weights are updated for each frame, dynamic downmixing adapted to the orientation information is used. Alternatively, fixed downmixing is used, where each object is assumed to be at a static position. This static position could, for example, correspond to the object's initial orientation, which then results in static virtual cardioid weights that are the same for all frames.

[0148] If the target bit rate allows, more than two transmission channels can be envisioned. With three transmission channels, the cards can be evenly arranged at, for example, 0°, 120°, and -120°. If four transmission channels are used, the fourth card can face upwards, or the four cards can again be evenly arranged horizontally. If the object location is, for example, only part of a hemisphere, the arrangement can also be customized for the object location. The resulting downmixed signal is processed by the core encoder and, along with encoding parameterization auxiliary information, is converted into a bitstream representation.

[0149] Alternatively, the input object signals can be fed into the core encoder without being combined into a downmixed signal. In this case, the number of transmission channels corresponds to the number of input object signals. Typically, a maximum number of transmission channels related to the total bit rate is given. Then, the downmixed signal is used only if the number of input object signals exceeds this maximum number of transmission channels.

[0150] Figure 6a This illustrates the method for encoding audio signals (e.g., by...) Figure 1a or Figure 2 or Figure 3 A decoder decodes the output signal, which includes directional information of multiple audio objects and one or more transmission channels. Furthermore, for one or more frequency intervals of a time frame, the encoded audio signal includes parameter data of at least two related audio objects, wherein the number of at least two related objects is less than the total number of audio objects. Specifically, the decoder includes an input interface for providing a spectral representation of one or more transmission channels having multiple frequency intervals in the time frame. This represents the signal forwarded from input interface block 600 to audio renderer block 700. Specifically, audio renderer 700 is configured to render one or more transmission channels into multiple audio channels using the directional information included in the encoded audio signal, the number of audio channels preferably being two channels for stereo output formats, or preferably more than two channels for a larger number of output formats (e.g., 3-channel, 5-channel, 5.1-channel, etc.). Specifically, audio renderer 700 is configured to: for each of the one or more frequency intervals, calculate the contribution of one or more transmission channels based on first directional information associated with a first related audio object among the at least two related audio objects and second directional information associated with a second related audio object among the at least two related audio objects. Specifically, the directional information of the multiple audio objects includes first directional information associated with the first object and second directional information associated with the second object.

[0151] Figure 8b The parameter data of the frame is shown. In a preferred embodiment, this parameter data includes direction information 810 for a plurality of audio objects, as well as the power ratio of each parameter band in a number of parameter bands shown at 812, and one (preferably two or more) object indexes for each parameter band indicated at block 814. Specifically, Figure 8c The text shows more detailed directional information for multiple audio objects 810. Figure 8cA table is shown with a first column containing the IDs of a given object, ranging from 1 to N, where N is the number of audio objects. A second column is also provided, containing orientation information for each object, preferably as azimuth and elevation values, or, in the case of two dimensions, only as the azimuth value. This is shown at 818. Therefore, Figure 8c The input to is shown Figure 6a The "direction codebook" included in the encoded audio signal is located in the input interface 600. Direction information from column 818 is uniquely associated with a specific object ID from column 816 and is valid for the "entire" object in the frame (i.e., for all frequency bands in the frame). Therefore, regardless of whether the number of frequency bands is a time / frequency band in a high-resolution representation or a time / parameter band in a lower-resolution representation, only a single direction piece of information is sent and used by the input interface for each object identifier.

[0152] In this context, Figure 8a It shows when Figure 2 or Figure 3 When filter bank 102 is implemented as the previously discussed CLDFB (Complex Low-Delay Filter Bank), the time / frequency representation generated by this filter bank is... Figure 8b and Figure 8c Discuss frames that provide directional information, and generate filter banks. Figure 8a The time slots range from 0 to 15, comprising 16 time slots and 60 frequency bands range from 0 to 59. Therefore, one time slot and one frequency band represent time / frequency region 802 or 804. However, to reduce the bit rate of auxiliary information, it is preferable to convert the high-resolution representation to a low-resolution representation, such as... Figure 8b As shown, there is only a single time interval, and 60 frequency bands are converted as follows: Figure 8b The 11 parameter bands are shown at position 812. Therefore, as... Figure 10c As shown, high resolution is indicated by the slot index n and the band index k, while low resolution is given by the grouped slot index m and the parameter band index l. However, in the context of this specification, time / frequency intervals may include... Figure 8a High-resolution time / frequency zones 802, 804 or by Figure 10c The input of block 731c contains the grouped slot index and the low-resolution time / frequency unit with indexed parameters.

[0153] exist Figure 6aIn one embodiment, the audio renderer 700 is configured to: for each of one or more frequency ranges, calculate the contribution of one or more transmission channels based on first directional information associated with a first related audio object among at least two related audio objects and second directional information associated with a second related audio object among at least two related audio objects. Figure 8b In the embodiment shown, block 814 has an object index for each relevant object in the parameter band, i.e., it has two or more object indices, such that there are two contributions for each time frequency interval.

[0154] As will be discussed later Figure 10a In summary, the calculation of contributions can be performed indirectly via a mixing matrix, where the gain value of each relevant object is determined and used to calculate the mixing matrix. Alternatively, such as Figure 10b As shown, the contribution can be explicitly calculated again using the gain value, and then the explicitly calculated contribution is summed for each output channel within a certain time / frequency interval. Therefore, regardless of whether the contribution is explicitly or implicitly calculated, the audio renderer still uses directional information to render one or more transport channels into multiple audio channels, such that for each of the one or more frequency intervals, the contribution of one or more transport channels is included in the multiple audio channels based on first directional information associated with a first related audio object among at least two related audio objects and second directional information associated with a second related audio object among at least two related audio objects.

[0155] Figure 6b A decoder for decoding an encoded audio signal according to the second aspect is shown. The encoded audio signal includes: directional information of a plurality of audio objects and one or more transmission channels; and parametric data of the audio objects for one or more frequency intervals of a time frame. Similarly, the decoder includes an input interface 600 for receiving the encoded audio signal, and the decoder includes an audio renderer 700 for rendering the one or more transmission channels into a plurality of audio channels using the directional information. Specifically, the audio renderer is configured to calculate direct response information based on one or more audio objects in each of the plurality of frequency intervals and the directional information associated with one or more related audio objects in the frequency intervals. This direct response information preferably includes gain values ​​for covariance synthesis or advanced covariance synthesis, or for explicitly calculating the contribution of one or more transmission channels.

[0156] Preferably, the audio renderer is configured to use direct response information of one or more related audio objects in time / frequency band and to compute covariance synthesis information using information about multiple audio channels. Furthermore, the covariance synthesis information (preferably a blending matrix) is applied to one or more transmission channels to obtain multiple audio channels. In another implementation, the direct response information is the direct response vector of each of the one or more audio objects, and the covariance synthesis information is a covariance synthesis matrix; the audio renderer is configured to perform matrix operations for each frequency range when applying the covariance synthesis information.

[0157] Furthermore, the audio renderer 700 is configured to: derive direct response vectors for one or more audio objects when calculating direct response information, and calculate a covariance matrix for each direct response vector for one or more audio objects. Additionally, a target covariance matrix is ​​calculated when calculating covariance synthesis information. However, relevant information about the target covariance matrix (i.e., the direct response matrices or vectors of one or more primary objects, and a diagonal matrix of direct power indicated as E, determined by applying a power ratio) can be used instead of the target covariance matrix.

[0158] Therefore, the target covariance information does not necessarily have to be an explicit target covariance matrix, but can be derived from the covariance matrix of an audio object or the covariance matrix of more audio objects in a time / frequency interval, from the power information of the corresponding one or more audio objects in the time / frequency interval, and from the power information derived from one or more transmission channels for one or more time / frequency intervals.

[0159] The bitstream representation is read by the decoder, and the encoded transport channel and the encoded parameterization auxiliary information contained therein can be used for further processing. The parameterization auxiliary information includes:

[0160] - Directional information used to quantize azimuth and elevation values ​​(for each frame)

[0161] - The object index representing the subset of related objects (for each parameter band).

[0162] - The quantization power ratio that correlates related objects with each other (for each parameter band)

[0163] All processing is performed frame-by-frame, where each frame contains one or more subframes. For example, a frame may consist of four subframes, in which case each subframe would have a duration of 5 milliseconds. Figure 4 A simplified overview of the decoder is shown.

[0164] Figure 4 An audio decoder implementing the first and second aspects is shown. Figure 6a and Figure 6b The input interface 600 shown includes a demultiplexer 602, a core decoder 604, a decoder 608 for decoding the object index, a decoder 612 for decoding and dequantizing the power ratio, and a decoder for decoding and dequantizing the direction information indicated at 612. Furthermore, the input interface includes a filter bank 606 for providing a transmission channel in time / frequency representation form.

[0165] The audio renderer 700 includes: a direct response calculator 704; a prototype matrix provider 702, controlled by output configuration received, for example, from a user interface; a covariance synthesis block 706; and a synthesis filter bank 708, to ultimately provide an output audio file including the number of audio channels in a channel output format.

[0166] Therefore, items 602, 604, 606, 608, 610, and 612 are preferably included. Figure 6a and Figure 6b In the input interface, and Figure 4 Projects 702, 704, 706, and 708 are... Figure 6a or Figure 6b The portion of the audio renderer indicated by the figure mark 700.

[0167] The encoded parameterization auxiliary information is decoded to re-obtain the quantized power ratio, quantized azimuth angle, and quantized elevation angle (direction information), as well as the object index. An untransmitted power ratio is obtained by utilizing the fact that the sum of all power ratios is 1. Their resolution... This corresponds to the grouping of time / frequency zones used on the encoder side. A finer time / frequency resolution is then used. During further processing steps, the parameters of the parameter band are valid for all time / frequency regions contained in that parameter band, corresponding to making An extension of.

[0168] The encoded transmission path is decoded by the core decoder. Each frame of the thus decoded audio signal is converted into a time / frequency representation using a filter bank (matched to the filter bank used in the encoder), the resolution of which is typically better than (but at least equal to) the resolution used for parameterizing auxiliary information.

[0169] Output signal rendering / compositing

[0170] The following description applies to one frame of an audio signal; The transpose operator is represented as follows:

[0171] Use the decoding transmission channel That is, the audio signal in time-frequency representation (in this case, including two transmission channels) and parameterized auxiliary information, deriving the mixing matrix for each subframe (or frames used to reduce computational complexity). To synthesize a time-frequency output signal that includes multiple output channels (e.g., 5.1, 7.1, 7.1+4, etc.). :

[0172] For all (input) objects, the so-called direct response values ​​are determined using the direction of the transmitted object. These direct response values ​​describe the translation gain to be used for the output channel. These direct response values ​​are specific to the target layout, i.e., the number and location of the speakers (provided as part of the output configuration). Examples of translation methods include Vector Basis Amplitude Translation (VBAP) [Pulkki 1997] and Edge Fading Amplitude Translation (EFAP) [Borß 2014]. Each object has a vector of direct response values ​​associated with it. (Contains as many elements as the speakers). These vectors are calculated once per frame. Note: If the object's position corresponds to a speaker's position, the vector contains a value of 1 for that speaker; all others contain 0. If the object is located between two (or three) speakers, the corresponding number of non-zero vector elements is two (or three).

[0173] The actual synthesis steps (in this embodiment, covariance synthesis [Vilkamo2013]) include the following sub-steps (see...) Figure 5 (Visualization)

[0174] For each parameter band, the object index describing the principal subset of input objects grouped into the time / frequency region of that parameter band is used to extract the vector subset required for further processing. Since we are only considering, for example, two related objects, we need two vectors associated with these two related objects. .

[0175] o Based on direct response value Then, for each relevant object, calculate the covariance matrix with a size of output channel * output channel. :

[0176]

[0177] For each time / frequency zone (within the parameter band), determine the audio signal power. In the case of two transmission channels, the signal power of the first channel is added to the signal power of the second channel. Each of these power ratios is multiplied by the signal power to produce a direct power value for each relevant / primary object i:

[0178]

[0179] For each frequency band k, the final target covariance matrix is ​​of size (output channel * output channel). It is obtained by summing all time slots n within the (sub)frame and summing all related objects:

[0180]

[0181] Figure 5 It shows in Figure 4 A detailed overview of the covariance synthesis steps performed in block 706. Specifically, Figure 5 The embodiment includes a signal power calculation block 721, a direct power calculation block 722, a covariance matrix calculation block 73, a target covariance matrix calculation block 724, an input covariance matrix calculation block 726, a hybrid matrix calculation block 725, and a rendering block 727, wherein the rendering block 727 is for... Figure 5 Additional land includes Figure 4 The filter block 708 is configured such that the output signal of block 727 preferably corresponds to the time-domain output signal. However, when block 708 is not included... Figure 5 When rendered in a block, the result is the spectral domain representation of the corresponding audio channel.

[0182] (The following steps are part of the state-of-the-art [Vilkamo2013] and are added for clarity.)

[0183] For each (sub)frame and each frequency band, calculate the input covariance matrix of size * transmission channel based on the decoded audio signal. Alternatively, only the entries on the main diagonal can be used, in which case the other non-zero entries are set to zero.

[0184] `o` defines a prototype matrix of size `output channels * transmission channels` that describes the mapping from transmission channels to output channels (provided as part of the output configuration), the number of which is given by the target output format (e.g., target speaker layout). This prototype matrix can be static or vary frame-by-frame. Example: If only a single transmission channel is transmitted, that channel is mapped to each output channel. If two transmission channels are transmitted, the left (first) channel is mapped to all output channels located within (+0°, +180°), i.e., the "left" channel. The right (second) channel is correspondingly mapped to all output channels located within (-0°, -180°), i.e., the "right" channel. (Note: 0° describes the position in front of the listener, a positive angle describes the position to the left of the listener, and a negative angle describes the position to the right of the listener. If different conventions are used, the signs of the angles need to be adjusted accordingly.)

[0185] o Using the input covariance matrix Target covariance matrix The prototype matrix is ​​used to compute the mixing matrix for each (sub)frame and each frequency band [Vilkamo2013], resulting in, for example, 60 mixing matrices per (sub)frame.

[0186] The mixing matrix is ​​interpolated (e.g., linearly) between (sub)frames, corresponding to temporal smoothing.

[0187] Finally, by using the final blending matrix The output channel is the product of the size of each of the transmission channels. Band-by-band synthesis into the decoding transmission channel The corresponding frequency band represented by time / frequency:

[0188]

[0189] Note that we did not use the residual signal as described in [Vilkamo2013]. .

[0190] - Use a filter bank to output the signal Convert back to time domain representation .

[0191] Optimized Covariance Synthesis

[0192] Regarding how to calculate the input covariance matrix in this embodiment and target covariance matrix Therefore, it is possible to implement certain optimizations to the computation of the optimal mixture matrix using covariance synthesis from [Vilkamo2013], resulting in a significant reduction in the computational complexity of the mixture matrix computation. Note that in this section, the Hadamard operator... This operator represents an element-wise operation on a matrix, meaning it doesn't follow the rules of matrix multiplication, but rather performs the operation element by element. This operator implies that the operation is not performed on the entire matrix, but on each element separately. For example, multiplying matrices A and B will not correspond to matrix multiplication AB=C, but rather to the element-wise operation a_ij*b_ij=c_ij.

[0193] SVD() denotes Singular Value Decomposition. The algorithm from [Vilkamo2013], presented as a Matlab function (List 1), is as follows (prior art):

[0194]

[0195]

[0196]

[0197] As mentioned in the previous section, it can only be used optionally. The main diagonal element is set, and all other entries are set to zero. In this case, It is a diagonal matrix, and the efficient decomposition of equation (3) of [Vilkamo2013] is

[0198]

[0199] Furthermore, SVD for row 3 from existing technology algorithms is no longer required.

[0200] Consider using it from a direct response The formula for generating the target covariance using direct power (or direct energy) from the previous section.

[0201]

[0202]

[0203]

[0204] The last formula can be rearranged and written as

[0205]

[0206] If we define now

[0207]

[0208] Thus obtain

[0209]

[0210] It is easy to see that if we arrange The direct response matrix of the most important objects The direct response in, and the creation of a diagonal matrix of direct power as ,in , It can also be expressed as

[0211]

[0212] And it satisfies equation (3) of [Vilkamo2013]. The efficient decomposition of is given by the following formula:

[0213]

[0214] Therefore, SVD for row 1 from existing technology algorithms is no longer needed.

[0215] This leads to the optimization algorithm used for covariance synthesis in this embodiment, which also takes into account that we always use the energy compensation option and therefore do not require a residual target covariance. :

[0216]

[0217]

[0218] A careful comparison of existing algorithms and the proposed algorithm shows that: the former requires sizes of respectively , and The three SVDs of the matrix, where, It is the number of downmixing channels, and It represents the number of output channels to which the object is rendered.

[0219] The proposed algorithm only requires a size of An SVD of the matrix, where, This refers to the number of primary objects. Furthermore, due to... Usually more Much smaller, therefore the matrix is ​​smaller than the corresponding matrix from existing technology algorithms.

[0220] for The matrix [Golub2013] has a standard SVD implementation complexity of approximately O(n). ,in, and The computational complexity depends on the constants of the algorithm used. Therefore, compared with existing algorithms, the proposed algorithm has a significantly reduced computational complexity.

[0221] Subsequently, regarding Figure 7a , Figure 7b Preferred embodiments related to the encoder side of the first aspect were discussed. Furthermore, regarding... Figures 9a to 9d The preferred implementation of the encoder-side implementation of the second aspect is discussed.

[0222] Figure 7a It shows Figure 1a A preferred implementation of the object parameter calculator 100. In block 120, the audio object is converted into a spectrum representation. This is done by... Figure 2 or Figure 3 The filter bank 102 is implemented. Then, in block 122, the selection information is calculated, for example as follows: Figure 2 or Figure 3Block 104 is shown. For this purpose, amplitude-related measurements can be used, such as amplitude itself, power, energy, or any other amplitude-related measurement obtained by increasing the amplitude to a power other than 1. The result of Block 122 is a set of selection information for each object in the corresponding time / frequency interval. Then, in Block 124, the object IDs for each time / frequency interval are derived. In the first aspect, two or more object IDs for each time / frequency interval are derived. According to the second aspect, the number of object IDs for each time / frequency interval can even be only a single object ID, allowing the identification of the most important, strongest, or most relevant object from the information provided by Block 122 in Block 124. Block 124 outputs information about the parameter data and includes one or more indices of the most relevant one or more objects.

[0223] In cases where there are two or more related objects in each time / frequency interval, the function of block 126 is useful for calculating amplitude-related measurements characterizing the objects in the time / frequency interval. This amplitude-related measurement can be the same as the amplitude-related measurement already calculated in block 122 for the selected information, or preferably, the combined value is calculated using information already calculated by block 102, as shown by the dashed line between blocks 122 and 126. The amplitude-related measurement or one or more combined values ​​are then calculated in block 126 and forwarded to quantizer and encoder block 212 so that the encoded amplitude-related value or encoded combined value is included in the auxiliary information as additional parameterized auxiliary information. Figure 2 or Figure 3 In this embodiment, these values ​​are "coded power ratios" included in the bitstream along with the "coded object index". With only a single object ID per frequency interval, power ratio calculation and quantization encoding are not necessary, and the index of the most relevant object in the time-frequency interval is sufficient to perform decoder-side rendering.

[0224] Figure 7b It shows Figure 7b A preferred implementation of the calculation of selection information 102. As shown in block 123, signal power is calculated as selection information for each object and each time / frequency interval. Then, in the example shown... Figure 7aIn block 125 of a preferred implementation of block 124, the object IDs of a single object, or preferably two or more objects, with the highest power are extracted and output. Furthermore, in the case of two or more related objects, the power ratio is calculated as shown in block 127 of a preferred implementation of block 126, where the power ratio is calculated relative to the extracted object IDs, which are related to the power of all extracted objects having the corresponding object IDs found in block 125. This process is advantageous because only a combination of values ​​less than the number of objects in the time / frequency interval must be sent, since a rule known to the decoder exists in this embodiment that dictates the power ratios of all objects must sum to 1. Preferably, Figure 7a Blocks 120, 122, 124, 126 and / or Figure 7b The functions of 123, 125, and 127 are provided by Figure 1a The object parameter calculator 100 is used to implement this, and Figure 7a The function of block 212 is determined by Figure 1a The output interface 200 is used to implement this.

[0225] Subsequently, several embodiments are described in more detail. Figure 1b The apparatus shown is for encoding according to the second aspect. In step 110a, the direction information is either extracted from the input signal, for example, as... Figure 12a As shown, it can be extracted either by reading or parsing metadata information included in the metadata section or metadata file. In step 110b, the direction information of each frame and audio object is quantized, and the quantization index of each frame per object is forwarded to the encoder or output interface, for example... Figure 1b The output interface 200. In step 110c, the direction quantization index is dequantized to have a dequantized value that can also be directly output by block 110b in some implementations. Then, based on the dequantized direction index, block 422 calculates the weights for each transmission channel and each object based on a certain virtual microphone setup. This virtual microphone setup may include two virtual microphone signals arranged in the same location but with different orientations, or it may be a setup where there are two different positions relative to a reference position or orientation (e.g., the virtual listener position or orientation). The setup with two virtual microphone signals will result in weights for the two transmission channels of each object.

[0226] In the case of generating three transmission channels, the virtual microphone setup can be considered to include three virtual microphone signals from microphones arranged at the same location but with different orientations, or at three different locations relative to a reference location or orientation, wherein the reference location or orientation may be the virtual listener location or orientation.

[0227] Alternatively, four transmission channels can be generated based on a virtual microphone setup that generates four virtual microphone signals from microphones arranged in the same location but with different orientations, or from microphones arranged in four different locations relative to a reference location or reference orientation, wherein the reference location or orientation can be a virtual listener location or virtual listener orientation.

[0228] In addition, in order to calculate the weight w for each object and each transmission channel (taking two channels as an example) L and w R The virtual microphone signal is a signal derived from a virtual first-order microphone, a virtual cardioid microphone, a virtual figure-eight microphone, a dipole microphone, or a two-way microphone, or a signal derived from a virtual directional microphone, a virtual subcardioid microphone, a virtual unidirectional microphone, a virtual supercardioid microphone, or a virtual omnidirectional microphone.

[0229] In this context, it should be noted that no actual microphone placement is required for calculating the weights. Instead, the rules used to calculate the weights vary depending on the virtual microphone setup (i.e., the placement and characteristics of the virtual microphone).

[0230] exist Figure 9a In block 404, weights are applied to the objects such that, for each object, the contribution of the object to a specific transmission channel is obtained if the weight is not zero. Therefore, block 404 receives the object signal as input. Then, in block 406, the contributions for each transmission channel are summed, such that, for example, the contributions of the objects to the first transmission channel are summed and the contributions of the objects to the second transmission channel are summed, and so on. As shown in block 406, the output of block 406 is, for example, the transmission channel in the time domain.

[0231] Preferably, the object signal input to block 404 is a time-domain object signal with full-band information, and the application in block 404 and the summation in block 406 are performed in the time domain. However, in other embodiments, these steps may also be performed in the spectral domain.

[0232] Figure 9b Another embodiment for implementing static downmixing is shown. For this purpose, the orientation information of the first frame is extracted in block 130, and weights are calculated based on the first frame, as shown in block 403a. Then, for the other frames indicated in block 408, the weights remain unchanged to achieve static downmixing.

[0233] Figure 9cAnother implementation of dynamic downmixing is shown. To do this, block 132 extracts the orientation information for each frame and updates the weights for each frame, as shown in block 403b. Then, in block 405, the updated weights are applied to the frames to achieve dynamic downmixing from frame to frame. Figure 9b and Figure 9c Other implementations between those extreme cases are also useful, such as updating weights only for every second, third, or nth frame, and / or performing weight smoothing over time so that the antenna characteristics do not change too much from time to time for the purpose of downmixing based on directional information. Figure 9d It shows the result of Figure 1b Another implementation of the downmixer 400 controlled by the object orientation information provider 110. In block 410, the downmixer is configured to analyze the orientation information of all objects in the frame, and in block 112, for calculating the weight w of the stereo example. L and w R The purpose is to position the microphone in accordance with the analysis results, where microphone placement refers to microphone position and / or microphone orientation. In block 414, similar to the section about... Figure 9b The static undermixing discussed in block 408 leaves the microphone available for other frames, or according to... Figure 9c Block 405 discusses updating the microphone in order to obtain... Figure 9d The function of block 414. Regarding the function of block 412, the microphones can be positioned to achieve good separation, such that the first virtual microphone "looks" at the first group of objects and the second virtual microphone "looks" at the second group of objects, which are different from the first group of objects, and preferably, the difference is that any object in one group is not included in the other group as much as possible. Alternatively, the analysis of block 410 can be enhanced by other parameters, and the placement can also be controlled by other parameters.

[0234] Subsequently, based on the first or second aspect and regarding, for example Figure 6a and Figure 6b The preferred implementation of the decoder discussed is as follows: Figure 10a , Figure 10b , Figure 10c , Figure 10d and Figure 11 Provided.

[0235] In block 613, input interface 600 is configured to obtain individual object orientation information associated with the object ID. This process corresponds to... Figure 4 or Figure 5 The function of block 612, and leads to, as about Figure 8b And especially the “frame codebook” shown and discussed in 8c.

[0236] Furthermore, in block 609, one or more object IDs are obtained for each time / frequency interval, regardless of whether this data is available for low-resolution parameter bands or high-resolution frequency bands. Corresponding to Figure 4 The result of block 609 of the process in block 608 is a specific ID of one or more related objects in the time / frequency interval. Then, in block 611, from the "frame codebook" (i.e., from...) Figure 8c The example table shown retrieves specific object orientation information for one or more IDs for each time / frequency interval. Then, in block 704, for each time / frequency interval, the gain value of one or more related objects for each output channel controlled by the output format is calculated. Then, in block 730 or 706, 708, the output channel is calculated. The function for calculating the output channel can be as follows: Figure 10b The calculation of the contribution to one or more transmission channels is shown in the diagram and can be performed as follows: Figure 10d or Figure 11 The method shown is to perform this through indirect calculation and use of the contribution to the transmission channel. Figure 10b It shows that in relation to Figure 4 The corresponding function in block 610 obtains the power value or power ratio. These power values ​​are then applied to the respective transmission channels for each relevant object, as shown in blocks 733 and 735. Furthermore, in addition to the gain value determined by block 704, these power values ​​are also applied to the respective transmission channels, causing blocks 733 and 735 to generate object-specific contributions for each transmission channel (e.g., transmission channels ch1, ch2, etc.). Then, in block 737, these explicitly calculated channel transmission contributions are summed for each output channel in each time / frequency interval.

[0237] Then, depending on the implementation, a diffusion signal calculator 741 can be provided, which generates a diffusion signal for each output channel ch1, ch2, ... in the corresponding time / frequency interval, and the combination of the diffusion signal and the contribution result of block 737 is combined to obtain the complete channel contribution in each time / frequency interval. When the covariance synthesis additionally depends on the diffusion signal, the signal corresponds to Figure 4 The input of filter bank 708. However, when covariance synthesis 706 depends not on the diffused signal but only on processing without any decorrelator, then at least the energy of the output signal in each time / frequency interval corresponds to the energy in Figure 10bThe energy contributed by the channel at the output of block 739. Furthermore, without using the diffuse signal calculator 741, the result of block 739 corresponds to the result of block 706, which is a complete channel contribution with each time / frequency interval that can be individually converted for each output channel ch1, ch2, so as to finally obtain an output audio file with time-domain output channels, which can be stored or forwarded to speakers or any type of rendering device.

[0238] Figure 10c It shows Figure 10b or Figure 4 The preferred implementation of the function of block 610. In step 610a, a combined (power) value or several values ​​are obtained for a certain time / frequency interval. In block 610b, based on the calculation rule that all combined values ​​must sum to 1, the corresponding other values ​​of other related objects in the time / frequency interval are calculated.

[0239] The result will then preferably be a low-resolution representation, wherein, for each group slot index and each parameter band index, the low-resolution representation has two power ratios. These power ratios represent the low time / frequency resolution. In block 610c, the time / frequency resolution can be extended to a high time / frequency resolution, such that it has power values ​​for the time / frequency regions of the high-resolution slot index n and the high-resolution band index k. This extension may include directly using the same low-resolution index for the corresponding slots within the group slots and the corresponding bands within the parameter bands.

[0240] Figure 10d The calculation is shown Figure 4 A preferred implementation of the covariance synthesis information function in block 706 is represented by a mixing matrix 725 for mixing two or more input transmission channels into two or more output signals. Therefore, when there are, for example, two transmission channels and six output channels, the size of the mixing matrix for each individual time / frequency interval will be six rows and two columns. In conjunction with... Figure 5 In block 723, corresponding to the function of block 723, the gain value or direct response value of each object in each time / frequency interval is received, and the covariance matrix is ​​calculated. In block 722, the power value or power ratio is received, and the direct power value of each object in the time / frequency interval is calculated, and... Figure 10d Block 722 in the middle corresponds to Figure 5 Block 722.

[0241] Input the results from blocks 721 and 722 into the target covariance matrix calculator 724. Additionally or alternatively, the target covariance matrix C yExplicit calculation is not required. Instead, relevant information included in the target covariance matrix (i.e., direct response value information indicated in matrix R and direct power values ​​indicated in matrix E for two or more related objects) is input into block 725a for calculating the mixing matrix for each time / frequency interval. Furthermore, the mixing matrix 725a receives information about the results from the target covariance matrix. Figure 5 The input covariance matrix C derived from the two or more transmission channels shown in block 726 corresponding to block 726 x And information about the prototype matrix Q. Temporal smoothing can be performed on the mixing matrix for each time / frequency interval and frame, as shown in block 725b, and in conjunction with... Figure 5 In at least a portion of the rendering block corresponding to block 727, a blending matrix is ​​applied in a non-smooth or smooth form to the transmission channels in the corresponding time / frequency interval to obtain the full channel contribution in the time / frequency interval, which is essentially similar to the previous discussion on... Figure 10b The corresponding full contribution discussed at the output of block 739. Therefore, Figure 10b An implementation of explicit calculation of the transmission channel contribution is shown, while Figure 10d The diagram shows the effect of the target covariance matrix C. y Alternatively, the transmission channel contribution of each relevant object in each time / frequency interval and in each time / frequency interval can be implicitly calculated by directly introducing the relevant information R and E into the hybrid matrix calculation block 725a via blocks 723 and 722.

[0242] Subsequently, regarding Figure 11 A preferred optimization algorithm for covariance synthesis is shown. It should be summarized as follows: Figure 11 All the steps shown are in Figure 4 The covariance of the composite is within 706 or within Figure 5 Hybrid matrix computation block 725 or Figure 10d The calculation is performed within 725a. In step 751, the first decomposition result K is calculated. y Due to the following facts: Figure 10d As shown, the decomposition result can be easily computed by directly using the gain value information included in matrix R and the direct power information from two or more related objects, specifically included in matrix ER, without explicitly calculating the covariance matrix. Therefore, the first decomposition result in block 751 can be computed directly without much effort, since a specific singular value decomposition is no longer required.

[0243] In step 752, the second decomposition result is calculated as K. x Since the input covariance matrix is ​​treated as a diagonal matrix ignoring off-diagonal elements, the decomposition result can also be computed without explicit singular value decomposition.

[0244] Then, in step 753, a first regularization result based on the first regularization parameter α is calculated, and in step 754, a second regularization result is calculated based on the second regularization parameter β. Since K x In the preferred implementation, it is a diagonal matrix, thus simplifying the calculation of the first regularization result 753 compared to the prior art, because S x The calculation involves only parameter variations, rather than decomposition as in existing technologies.

[0245] Furthermore, for the calculation of the second regularization result in block 754, the first step additionally involves only parameter renaming instead of the prior art involving matrix U. x HS Multiplication.

[0246] Furthermore, in step 755, the normalization matrix G is calculated. y And based on step 755, in step 756 based on K x and the prototype matrix Q and K obtained from block 751 y The information is used to calculate the unitary matrix P. Since no matrix Λ is needed here, the calculation of the unitary matrix P is simplified compared to existing available techniques.

[0247] Then, in step 757, the mixing matrix without energy compensation, i.e., M, is calculated. opt And for this purpose, the unitary matrix P, the results of block 754, and the results of block 751 are used. Then, in block 758, energy compensation is performed using the compensation matrix G. Performing energy compensation eliminates the need for any residual signal derived from the decorcorrelator. However, instead of performing energy compensation, this implementation adds energy with a sufficiently large capacity to fill the gap created by the mixing matrix M. opt The remaining bandgap contains a residual signal without energy information. However, for the purposes of this invention, the decorrelation signal is not relied upon to avoid any artifacts introduced by the decorrelation unit. However, energy compensation as shown in step 758 is preferred.

[0248] Therefore, the optimized algorithm for covariance synthesis offers advantages in the computation of the unitary matrix P in steps 751, 752, 753, 754, and within step 756. It is important to emphasize that the optimized algorithm even offers advantages over existing techniques, where only one or a subgroup of steps 755, 752, 753, 754, and 756 is implemented, as shown in the figure, while the corresponding other steps are implemented in the same way as existing techniques. This is because these improvements are not interdependent but can be applied independently. However, in terms of implementation complexity, the more improvements implemented, the better the process will be. Therefore, Figure 11The complete implementation of the embodiment is preferred because it provides the greatest possible reduction in complexity, but even if only one of steps 751, 752, 753, 754, and 756 is implemented according to the optimization algorithm and the other steps are implemented as in the prior art, a reduction in complexity is achieved without any degradation in quality.

[0249] Embodiments of the present invention can also be viewed as the following process: generating comfortable noise for a stereo signal by mixing three Gaussian noise sources (one Gaussian noise source per channel) and a third common noise source for creating relevant background noise, or additionally or separately, using coherent values ​​transmitted with the SID frame to control the mixing of the noise sources.

[0250] It should be noted here that all alternatives or aspects discussed above and below, as well as all aspects defined by the claims in the appended claims, can be used individually; that is, there are no other alternatives or objectives different from the contemplated alternatives, objectives, or independent claims. However, in other embodiments, two or more alternatives or aspects or independent claims can be combined with each other, and in other embodiments, all aspects or alternatives and all independent claims can be combined with each other.

[0251] The encoded signal of the present invention can be stored on a digital storage medium or a non-transitory storage medium, or it can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0252] Although some aspects have been described in the context of the apparatus, it will be clear that these aspects also represent a description of the corresponding method, where a block or apparatus corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also indicate a description of the features of the corresponding block or item or the corresponding apparatus.

[0253] Depending on certain implementation requirements, embodiments of the invention may be implemented in hardware or software. Implementation may be performed using a digital storage medium (e.g., floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM, or flash memory) on which electronically readable control signals are stored, in cooperation with (or capable of cooperating with) a programmable computer system, such that the corresponding methods are executed.

[0254] Some embodiments of the invention include a data carrier having electronically readable control signals, capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0255] Typically, embodiments of the present invention can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product is run on a computer. This program code may, for example, be stored on a machine-readable medium.

[0256] Other embodiments include a computer program stored on a machine-readable carrier or non-transitory storage medium for performing one of the methods described herein.

[0257] In other words, embodiments of the method of the present invention are therefore computer programs having program code for performing one of the methods described herein when the computer program is run on a computer.

[0258] Therefore, other embodiments of the method of the present invention are data carriers or digital storage media or computer-readable media on which a computer program is recorded for performing one of the methods described herein.

[0259] Therefore, other embodiments of the method of the present invention represent a data stream or signal sequence of a computer program for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).

[0260] Another embodiment includes a processing means, such as a computer or a programmable logic device, which is configured or adapted to perform one of the methods described herein.

[0261] Another embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.

[0262] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0263] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, it is intended to be limited only by the scope of the appended claims and not by the specific details given by way of the description and explanation of the embodiments herein.

[0264] Aspects (used independently of each other, or used in conjunction with all other aspects or only as subgroups of other aspects)

[0265] An apparatus, method, or computer program includes one or more of the features listed below:

[0266] Examples of novel inventions:

[0267] - Combining multi-wave concepts with object encoding (using more than one directional cue per T / F zone)

[0268] - An object encoding method that is as close as possible to the DirAC paradigm to allow for any kind of input type in IVAS (object content not currently covered).

[0269] Examples of inventions related to parametric (encoder) systems:

[0270] - For each T / F zone: the selection information of the n most relevant objects in that T / F zone plus the power ratio between the contributions of these n most relevant objects.

[0271] - For each frame, for each object: one direction

[0272] Examples of inventions related to rendering (decoders):

[0273] - Obtain the direct response value of each relevant object from the sent object index and direction information, as well as the target output layout.

[0274] - Obtain the covariance matrix from the direct response

[0275] - Calculate direct power based on the downmixed signal power and transmit power ratio of each relevant object.

[0276] - Obtain the final target covariance matrix from the direct power and covariance matrices.

[0277] - Use only the diagonal elements of the input covariance matrix

[0278] Optimized covariance synthesis

[0279] Some side notes regarding the differences from SAOC:

[0280] - Consider n primary objects instead of all objects

[0281] The power ratio is therefore related to OLD, but the calculation method is different.

[0282] SAOC does not use orientation -> orientation information at the encoder; it is only introduced at the decoder (rendering matrix).

[0283] -> The SAOC-3D decoder receives object metadata used to render the matrix.

[0284] - SAOC uses a downmixing matrix and sends the downmixing gain.

[0285] - This invention does not consider diffusivity in its embodiments.

[0286] Subsequently, other examples of the invention are summarized.

[0287] 1. An apparatus for encoding a plurality of audio objects and related metadata indicating directional information about the plurality of audio objects, comprising:

[0288] A downmixer (400) is used to downmix the plurality of audio objects to obtain one or more transmission channels;

[0289] Transmission channel encoder (300), used to encode one or more transmission channels to obtain one or more encoded transmission channels; and

[0290] Output interface (200) is used to output encoded audio signals including the one or more encoded transmission channels.

[0291] The downmixer (400) is configured to downmix the plurality of audio objects in response to directional information about the plurality of audio objects.

[0292] 2. The apparatus according to Example 1, wherein the downmixer (400) is configured as follows:

[0293] Two transmission channels are generated as two virtual microphone signals, which are arranged at the same location but with different orientations, or arranged at two different locations relative to a reference location or orientation, such as the virtual listener's position or orientation.

[0294] Three transmission channels are generated as three virtual microphone signals, which are arranged at the same location but with different orientations, or arranged at three different locations relative to a reference location or orientation, such as the virtual listener's position or orientation.

[0295] Four transmission channels are generated as four virtual microphone signals, which are arranged at the same location but with different orientations, or arranged at four different locations relative to a reference location or orientation such as the virtual listener's position or orientation.

[0296] The virtual microphone signal is a virtual first-order microphone signal, or a virtual cardioid microphone signal, or a virtual figure-eight, dipole, or bidirectional microphone signal, or a virtual directional microphone signal, or a virtual subcardioid microphone signal, or a virtual unidirectional microphone signal, or a virtual supercardioid microphone signal, or a virtual omnidirectional microphone signal.

[0297] 3. The apparatus according to Example 1 or 2, wherein the downmixer (400) is configured as follows:

[0298] For each of the plurality of audio objects, the direction information of the corresponding audio object is used to derive (402) the weighted information for each transmission channel;

[0299] The corresponding audio object is weighted using weighting information for the audio object of a specific transmission channel (404) to obtain the object contribution for the specific transmission channel, and

[0300] Combine (406) the object contributions of the multiple audio objects to the specific transmission channel to obtain the specific transmission channel.

[0301] 4. The apparatus according to one of the foregoing examples,

[0302] The downmixer (400) is configured to: calculate the one or more transmission channels as one or more virtual microphone signals, the one or more virtual microphone signals being arranged at the same location and having different orientations, or arranged at different locations relative to a reference location or orientation such as a virtual listener location or orientation, the orientation information being related to the reference location or orientation.

[0303] Wherein, the different positions or orientations are on the center line or to the left of the center line and on the center line or to the right of the center line, or wherein the different positions or orientations are evenly or unevenly distributed to horizontal positions or orientations, for example, at +90 degrees or -90 degrees relative to the center line, or at -120 degrees, 0 degrees and +120 degrees relative to the center line, or wherein the different positions or orientations include at least one position or orientation pointing upward or downward relative to the horizontal plane where the virtual listener is located, wherein the directional information of the plurality of audio objects is related to the position or reference position or orientation of the virtual listener.

[0304] 5. The apparatus according to one of the foregoing examples further includes:

[0305] Parameter processor (110) is configured to quantize metadata indicating directional information about the plurality of audio objects to obtain quantized directional terms for the plurality of audio objects.

[0306] The downmixer (400) is configured to operate in response to the quantization direction term, which is the direction information.

[0307] The output interface (200) is configured to introduce information about the quantization direction term into the encoded audio signal.

[0308] 6. The apparatus according to one of the foregoing examples,

[0309] The downmixer (400) is configured to perform analysis on directional information about the plurality of audio objects and place one or more virtual microphones based on the results of the analysis to generate the transmission channel.

[0310] 7. The apparatus according to one of the foregoing examples,

[0311] The downmixer (400) is configured to perform downmixing (408) using downmixing rules that are static across multiple time frames, or

[0312] The direction information is variable across multiple time frames, and the downmixer (400) is configured to downmix (405) using downmixing rules that are variable across the multiple time frames.

[0313] 8. The apparatus according to one of the foregoing examples,

[0314] The downmixer (400) is configured to perform downmixing in the time domain using a sample-by-sample weighted sum of samples from the plurality of audio objects.

[0315] 9. The apparatus according to any of the foregoing examples further includes:

[0316] An object parameter calculator (100) is configured to: calculate parameter data for at least two related audio objects for one or more frequency intervals among a plurality of frequency intervals related to a time frame, wherein the number of the at least two related audio objects is less than the total number of the plurality of audio objects, and

[0317] The output interface (200) is configured to introduce information about parameter data of at least two related audio objects for one or more frequency ranges into the encoded audio signal.

[0318] 10. The apparatus according to Example 9, wherein the object parameter calculator (100) is configured to:

[0319] Each of the plurality of audio objects is converted (120) into a spectral representation having the plurality of frequency ranges.

[0320] Calculate (122) selection information for each audio object for the one or more frequency ranges, and

[0321] Based on the selection information, derive (124) object identifiers as parameter data indicating the at least two related audio objects, and

[0322] The output interface (200) is configured to incorporate information about the object identifier into the encoded audio signal.

[0323] 11. The apparatus according to Example 9 or 10, wherein the object parameter calculator (100) is configured to: quantize and encode (212) one or more amplitude-related measurements or one or more combined values ​​derived from amplitude-related measurements of the relevant audio object in the one or more frequency ranges, as the parameter data, and

[0324] The output interface (200) is configured to introduce one or more quantized amplitude-related measurements or one or more quantized combined values ​​into the encoded audio signal.

[0325] 12. The apparatus according to Example 10 or 11,

[0326] The selection information refers to amplitude-related measurements of the audio object, such as amplitude, power, or loudness values, or the amplitude at which power is increased to a level other than 1.

[0327] The object parameter calculator (100) is configured to: calculate (127) a combined value, such as the ratio of the amplitude-related measurement of the relevant audio object to the sum of two or more amplitude-related measurements of the relevant audio object, and

[0328] The output interface (200) is configured to introduce information about the combined value into the encoded audio signal, wherein the number of information items about the combined value in the encoded audio signal is at least equal to 1 and less than the number of related audio objects in the one or more frequency ranges.

[0329] 13. The apparatus according to any one of Examples 10 to 12,

[0330] The object parameter calculator (100) is configured to select the object identifier based on the order of selection information of the plurality of audio objects in the one or more frequency ranges.

[0331] 14. The apparatus according to any one of Examples 10 to 13, wherein the object parameter calculator (100) is configured to:

[0332] Calculate the signal power (122) as the selection information.

[0333] For each frequency range, derive (124) object identifiers for two or more audio objects with the maximum signal power value corresponding to one or more frequency ranges.

[0334] Calculate (126) the power ratio between the sum of the signal powers of two or more audio objects having the maximum signal power value and the signal power of at least one audio object having the derived object identifier, as the parameter data, and

[0335] The power ratio is quantized and encoded (212), and

[0336] The output interface (200) is configured to introduce the quantized and encoded power ratio into the encoded audio signal.

[0337] 15. The apparatus according to any one of Examples 10 to 14, wherein the output interface (200) is configured to introduce the following into the encoded audio signal:

[0338] One or more encoded transmission channels

[0339] The parameter data includes two or more coded object identifiers of the relevant audio object in each of one or more frequency intervals within the multiple frequency intervals of the time frame, and one or more coded combination values ​​or coded amplitude related measurements.

[0340] The quantization and encoding direction data for each audio object in the time frame, wherein the direction data is constant for all frequency ranges in the one or more frequency ranges.

[0341] 16. The apparatus according to any one of Examples 9 to 15, wherein the object parameter calculator (100) is configured to: calculate parameter data of at least the most important object and the second most important object in the one or more frequency intervals, or

[0342] The plurality of audio objects comprises three or more audio objects, including a first audio object, a second audio object, and a third audio object.

[0343] The object parameter calculator (100) is configured to: calculate only a first group of audio objects, such as the first audio object and the second audio object, as the relevant audio objects for a first frequency interval in the one or more frequency intervals; and calculate only a second group of audio objects, such as the second audio object and the third audio object or the first audio object and the third audio object, as the relevant audio objects for a second frequency interval in the one or more frequency intervals, wherein the first group of audio objects differs from the second group of audio objects in at least one group member.

[0344] 17. The apparatus according to any one of Examples 9 to 16, wherein the object parameter calculator (100) is configured to:

[0345] Calculate raw parameterized data with a first time or frequency resolution, combine the raw parameterized data into combined parameterized data with a second time or frequency resolution lower than the first time or frequency resolution, and calculate parameter data for at least two related audio objects relative to the combined parameterized data with the second time or frequency resolution.

[0346] A parameter band having a second time or frequency resolution different from the first time or frequency resolution used in the time or frequency decomposition of the plurality of audio objects is determined, and parameter data of at least two related audio objects are calculated for the parameter band having the second time or frequency resolution.

[0347] 18. A decoder for decoding an encoded audio signal, the encoded audio signal comprising: one or more transmission channels and direction information of a plurality of audio objects; and parameter data of the audio objects for one or more frequency intervals of a time frame, the decoder comprising:

[0348] Input interface (600) for providing the one or more transmission channels in a spectral representation having multiple frequency ranges in the time frame; and

[0349] An audio renderer (700) is used to render the one or more transmission channels into multiple audio channels using the direction information.

[0350] The audio renderer (700) is configured to calculate direct response information (704) based on one or more audio objects in each of the plurality of frequency ranges and directional information (810) associated with one or more related audio objects in the frequency ranges.

[0351] 19. The decoder according to Example 18,

[0352] The audio renderer (700) is configured to: calculate (706) covariance synthesis information using the direct response information and information about the plurality of audio channels (702), and apply (727) the covariance synthesis information to the one or more transmission channels to obtain the plurality of audio channels, or

[0353] Wherein, the direct response information (704) is the direct response vector of each of one or more audio objects, and wherein, the covariance synthesis information is a covariance synthesis matrix, and wherein, the audio renderer (700) is configured to perform matrix operations for each frequency range when applying the covariance synthesis information (727).

[0354] 20. The decoder according to Example 18 or 19, wherein the audio renderer (700) is configured to:

[0355] In calculating the direct response information (704), direct response vectors of the one or more audio objects are derived, and for each of the one or more audio objects, a covariance matrix is ​​calculated based on each direct response vector.

[0356] When calculating the covariance synthesis information, target covariance information is derived from the following: the covariance matrix of an audio object or the covariance matrix of multiple audio objects, power information about the corresponding one or more audio objects, and power information derived from the one or more transmission channels.

[0357] 21. The decoder according to Example 20, wherein the audio renderer (700) is configured to:

[0358] When calculating the direct response information, direct response vectors of the one or more audio objects are derived, and for each of the one or more audio objects, a (723) covariance matrix is ​​calculated based on each direct response vector.

[0359] The input covariance information (726) is derived from the transmission channel, and

[0360] The mixed information (725a, 725b) is derived from the target covariance information, the input covariance information, and information about multiple channels, and

[0361] The hybrid information is applied (727) to the transmission channel of each frequency interval in the time frame.

[0362] 22. The decoder according to Example 21, wherein the result of applying mixing information to each frequency range in the time frame is converted (708) to the time domain to obtain multiple audio channels in the time domain.

[0363] 23. The decoder according to any one of Examples 18 to 22, wherein the audio renderer (700) is configured to:

[0364] When decomposing (752) the input covariance matrix derived from the transmission channel, only the main diagonal elements of the input covariance matrix are used, or

[0365] The decomposition of the target covariance matrix is ​​performed using the direct response matrix and the power matrix of the object or transmission channel (751), or

[0366] The decomposition of the input covariance matrix is ​​performed by taking the root of each main diagonal element of the input covariance matrix, or...

[0367] Calculate the regularized inverse matrix of the decomposed input covariance matrix (753), or

[0368] Singular value decomposition (756) is performed when calculating the optimal matrix to be used for energy compensation, without expanding the identity matrix.

[0369] 24. A decoder according to any one of Examples 18 to 23, wherein the parameter data of the one or more audio objects includes parameter data of at least two related audio objects, wherein the number of the at least two related audio objects is less than the total number of the plurality of audio objects, and

[0370] The audio renderer (700) is configured to: for each of the one or more frequency ranges, calculate the contribution from the one or more transmission channels based on first direction information associated with a first related audio object among the at least two related audio objects and second direction information associated with a second related audio object among the at least two related audio objects.

[0371] 25. The decoder according to Example 24,

[0372] The audio renderer (700) is configured to ignore the directional information of audio objects that are different from the at least two related audio objects for the one or more frequency ranges.

[0373] 26. The decoder according to Example 24 or 25,

[0374] The encoded audio signal includes amplitude-related measurements of each relevant audio object in the parameter data, or combined values ​​related to at least two relevant audio objects.

[0375] The audio renderer (700) is configured to: take into account the contributions from the one or more transmission channels based on first direction information associated with a first related audio object among the at least two related audio objects and second direction information associated with a second related audio object among the at least two related audio objects, or to determine the quantitative contribution of the one or more transmission channels based on the amplitude-related measurement or the combined value.

[0376] 27. The decoder according to Example 26, wherein the encoded signal includes combined values ​​from the parameter data, and

[0377] The audio renderer (700) is configured to: determine the contribution of the one or more transmission channels using a combined value of one of the relevant audio objects and the direction information of that relevant audio object; and

[0378] The audio renderer (700) is configured to determine the contribution of the one or more transmission channels using a combined value derived from the direction information of another related audio object in the one or more frequency ranges.

[0379] 28. The decoder according to any one of Examples 24 to 27, wherein the audio renderer (700) is configured to:

[0380] The direct response information (704) is calculated based on the relevant audio object in each of the plurality of frequency intervals and the directional information associated with the relevant audio object in the frequency interval.

[0381] 29. The decoder according to Example 28,

[0382] The audio renderer (700) is configured to: use diffusion information such as diffusion parameters or decorrelation rules included in the metadata to determine (741) the diffusion signal of each of the plurality of frequency intervals, and combine the direct response determined by the direct response information and the diffusion signal to obtain the spectral domain rendering signal of the channel among the plurality of channels.

[0383] 30. A method for encoding a plurality of audio objects and related metadata indicating directional information about the plurality of audio objects, comprising:

[0384] Downmix the multiple audio objects to obtain one or more transmission channels;

[0385] Encode the one or more transmission channels to obtain one or more encoded transmission channels; and

[0386] The output includes coded audio signals from the one or more coded transmission channels.

[0387] The downmixing includes downmixing the plurality of audio objects in response to directional information about the plurality of audio objects.

[0388] 31. A method for decoding an encoded audio signal, the encoded audio signal comprising: one or more transmission channels and direction information of a plurality of audio objects; and parameter data of the audio objects for one or more frequency intervals of a time frame, the method comprising:

[0389] The one or more transmission channels provide a spectral representation having multiple frequency ranges within the time frame; and

[0390] The direction information is used to render the audio from the one or more transmission channels into multiple audio channels.

[0391] The audio rendering includes: calculating direct response information based on one or more audio objects in each of the plurality of frequency ranges and directional information associated with one or more related audio objects in the frequency ranges.

[0392] 32. A computer program, when run on a computer or processor, for performing the method described according to Example 30 or the method described according to Example 31.

[0393] This application also includes the following examples:

[0394] A1. An apparatus for encoding a plurality of audio objects, comprising:

[0395] An object parameter calculator (100) is configured to: calculate parameter data for at least two related audio objects for one or more frequency intervals among a plurality of frequency intervals related to a time frame, wherein the number of the at least two related audio objects is less than the total number of the plurality of audio objects, and

[0396] The output interface (200) is configured to output an encoded audio signal, the encoded audio signal including information about parameter data of the at least two related audio objects in the one or more frequency ranges.

[0397] A2. The apparatus according to Example A1, wherein the object parameter calculator (100) is configured to:

[0398] Each of the plurality of audio objects is converted (120) into a spectral representation having multiple frequency ranges.

[0399] Calculate the selection information for each audio object in one or more frequency ranges as described in (122), and

[0400] Based on the selection information, object identifiers (124) are derived as parameter data indicating the at least two related audio objects, and

[0401] The output interface (200) is configured to incorporate information about the object identifier into the encoded audio signal.

[0402] A3. The apparatus according to example A1 or A2, wherein the object parameter calculator (100) is configured to: quantize and encode (212) one or more amplitude-related measurements or one or more combined values ​​derived from amplitude-related measurements of a relevant audio object in the one or more frequency ranges, as the parameter data, and

[0403] The output interface (200) is configured to introduce one or more quantized amplitude-related measurements or one or more quantized combined values ​​into the encoded audio signal.

[0404] A4. The apparatus according to example A2 or A3

[0405] The selection information refers to amplitude-related measurements of the audio object, such as amplitude, power, or loudness values, or the amplitude at which power is increased to a level other than 1.

[0406] The object parameter calculator (100) is configured to calculate (127) combined values, such as the ratio of the amplitude-related measurement of the relevant audio object to the sum of two or more amplitude-related measurements of the relevant audio object, and

[0407] The output interface (200) is configured to introduce information about the combined value into the encoded audio signal, wherein the number of information items about the combined value in the encoded audio signal is at least equal to 1 and less than the number of related audio objects in the one or more frequency ranges.

[0408] A5. The apparatus according to any one of Examples A2 to A4

[0409] The object parameter calculator (100) is configured to select the object identifier based on the order of selection information of the plurality of audio objects in the one or more frequency ranges.

[0410] A6. The apparatus according to any one of Examples A2 to A5, wherein the object parameter calculator (100) is configured to:

[0411] Calculate the signal power (122) as the selection information.

[0412] For each frequency range, derive (124) object identifiers for two or more audio objects with the maximum signal power value corresponding to one or more frequency ranges.

[0413] Calculate (126) the power ratio between the sum of the signal powers of two or more audio objects having the maximum signal power value and the signal power of each audio object having the derived object identifier as the parameter data, and

[0414] The power ratio is quantized and encoded (212), and

[0415] The output interface (200) is configured to introduce the quantized and encoded power ratio into the encoded audio signal.

[0416] A7. The apparatus according to any one of Examples A1 to A6, wherein the output interface (200) is configured to introduce the following into the encoded audio signal:

[0417] One or more encoded transmission channels

[0418] As the parameter data, two or more coded object identifiers of the relevant audio object in each of one or more frequency intervals in the multiple frequency intervals of the time frame, and one or more coded combination values ​​or coded amplitude related measurements, and

[0419] The quantized and encoded directional data for each audio object in the time frame, wherein the directional data is constant for all frequency ranges in the one or more frequency ranges.

[0420] A8. The apparatus according to any one of Examples A1 to A7, wherein the object parameter calculator (100) is configured to: calculate parameter data of at least the most important object and the second most important object in the one or more frequency intervals, or

[0421] The plurality of audio objects comprises three or more audio objects, including a first audio object, a second audio object, and a third audio object.

[0422] The object parameter calculator (100) is configured to: calculate only a first group of audio objects, such as the first audio object and the second audio object, as the relevant audio objects for a first frequency interval in the one or more frequency intervals; and calculate only a second group of audio objects, such as the second audio object and the third audio object or the first audio object and the third audio object, as the relevant audio objects for a second frequency interval in the one or more frequency intervals, wherein the first group of audio objects differs from the second group of audio objects in at least one group member.

[0423] A9. The apparatus according to any one of Examples A1 to A8, wherein the object parameter calculator (100) is configured to:

[0424] Calculate raw parametric data with a first time or frequency resolution, combine the raw parametric data into combined parametric data with a second time or frequency resolution lower than the first time or frequency resolution, and calculate parametric data for the at least two related audio objects relative to the combined parametric data with the second time or frequency resolution.

[0425] A parameter band having a second time or frequency resolution different from the first time or frequency resolution used in the time or frequency decomposition of the plurality of audio objects is determined, and parameter data of the at least two related audio objects are calculated for the parameter band having the second time or frequency resolution.

[0426] A10. The apparatus according to one of the foregoing examples, wherein the plurality of audio objects includes relevant metadata indicating directional information (810) about the plurality of audio objects, and

[0427] The device further includes:

[0428] A downmixer (400) is configured to downmix the plurality of audio objects to obtain one or more transmission channels, wherein the downmixer (400) is configured to: downmix the plurality of audio objects in response to directional information about the plurality of audio objects; and

[0429] Transmission channel encoder (300), used to encode one or more transmission channels to obtain one or more encoded transmission channels; and

[0430] The output interface (200) is configured to introduce one or more transmission channels into the encoded audio signal.

[0431] A11. The apparatus according to Example A10, wherein the downmixer (400) is configured as follows:

[0432] Two transmission channels are generated as two virtual microphone signals, which are arranged at the same location but with different orientations, or arranged at two different locations relative to a reference location or orientation, such as the virtual listener's position or orientation.

[0433] Three transmission channels are generated as three virtual microphone signals, which are arranged at the same location but with different orientations, or arranged at three different locations relative to a reference location or orientation, such as the virtual listener's position or orientation.

[0434] Four transmission channels are generated as four virtual microphone signals, which are arranged at the same location but with different orientations, or arranged at four different locations relative to a reference location or orientation such as the virtual listener's position or orientation.

[0435] The virtual microphone signal is a virtual first-order microphone signal, or a virtual cardioid microphone signal, or a virtual figure-eight, dipole, or bidirectional microphone signal, or a virtual directional microphone signal, or a virtual subcardioid microphone signal, or a virtual unidirectional microphone signal, or a virtual supercardioid microphone signal, or a virtual omnidirectional microphone signal.

[0436] A12. The apparatus according to example A10 or A11, wherein the downmixer (400) is configured as follows:

[0437] For each of the plurality of audio objects, the direction information of the corresponding audio object is used to derive (402) the weighted information for each transmission channel;

[0438] The corresponding audio object is weighted using weighting information for the audio object of a specific transmission channel (404) to obtain the object contribution for the specific transmission channel, and

[0439] Combine (406) the object contributions of the multiple audio objects to the specific transmission channel to obtain the specific transmission channel.

[0440] A13. The apparatus according to any one of Examples A10 to A12

[0441] The downmixer (400) is configured to: calculate the one or more transmission channels as one or more virtual microphone signals, the one or more virtual microphone signals being arranged at the same location and having different orientations, or arranged at different locations relative to a reference location or orientation such as a virtual listener location or orientation, the orientation information being related to the reference location or orientation.

[0442] Wherein, the different positions or orientations are on the center line or to the left and right of the center line, or wherein the different positions or orientations are evenly or unevenly distributed to horizontal positions or orientations, for example, at +90 degrees or -90 degrees relative to the center line, or at -120 degrees, 0 degrees and +120 degrees relative to the center line, or wherein the different positions or orientations include at least one position or orientation pointing upward or downward relative to the horizontal plane where the virtual listener is located, wherein the directional information of the plurality of audio objects is related to the position or reference position or orientation of the virtual listener.

[0443] A14. The apparatus according to any one of Examples A10 to A13 further includes:

[0444] Parameter processor (110) is configured to quantize metadata indicating directional information about the plurality of audio objects to obtain quantized directional terms for the plurality of audio objects.

[0445] The downmixer (400) is configured to operate in response to a quantized direction term that serves as the direction information, and

[0446] The output interface (200) is configured to introduce information about the quantization direction term into the encoded audio signal.

[0447] A15. The apparatus according to any one of Examples A10 to A14

[0448] The downmixer (400) is configured to perform (410) analysis on directional information about the plurality of audio objects and place (412) one or more virtual microphones to generate a transmission channel based on the results of the analysis.

[0449] A16. The apparatus according to any one of Examples A10 to A15

[0450] The downmixer (400) is configured to perform downmixing (408) using downmixing rules that are static across multiple time frames, or

[0451] The direction information is variable across multiple time frames, and the downmixer (400) is configured to downmix (405) using downmixing rules that are variable across the multiple time frames.

[0452] A17. According to the apparatus described in any one of Examples A10 to A16, the downmixer (400) is configured to perform downmixing in the time domain using a sample-by-sample weighted sum of samples of the plurality of audio objects.

[0453] A18. A decoder for decoding an encoded audio signal, the encoded audio signal comprising: one or more transmission channels and direction information of a plurality of audio objects, and parameter data of at least two related audio objects for one or more frequency intervals of a time frame, wherein the number of the at least two related audio objects is less than the total number of the plurality of audio objects, the decoder comprising:

[0454] Input interface (600) for providing the one or more transmission channels in a spectral representation having multiple frequency ranges in the time frame; and

[0455] An audio renderer (700) is configured to render the one or more transmission channels into a plurality of audio channels using the direction information, such that contributions from the one or more transmission channels are considered based on first direction information associated with a first related audio object among the at least two related audio objects and second direction information associated with a second related audio object among the at least two related audio objects.

[0456] The audio renderer (700) is configured to: for each of the one or more frequency ranges, calculate the contribution from the one or more transmission channels based on first direction information associated with a first related audio object among the at least two related audio objects and second direction information associated with a second related audio object among the at least two related audio objects.

[0457] A19. The decoder according to Example A18

[0458] The audio renderer (700) is configured to ignore the directional information of audio objects that are different from the at least two related audio objects for the one or more frequency ranges.

[0459] A20. The decoder according to examples A18 or A19.

[0460] The encoded audio signal includes amplitude-related measurements (812) of each relevant audio object in the parameter data or combined values ​​(812) related to at least two relevant audio objects, and

[0461] The audio renderer (700) is configured to determine (704) the quantitative contribution of the one or more transmission channels based on the amplitude-related measurement or the combined value.

[0462] A21. The decoder according to Example A20, wherein the encoded signal includes the combined values ​​in the parameter data, and

[0463] The audio renderer (700) is configured to: determine the contribution of the one or more transmission channels (704, 733) using a combined value of one of the relevant audio objects and the direction information of that relevant audio object; and

[0464] The audio renderer (700) is configured to determine (704, 735) the contribution of the one or more transmission channels using a combined value derived from the direction information of another related audio object in the one or more frequency ranges.

[0465] A22. The decoder according to one of examples A18 to A21, wherein the audio renderer (700) is configured as follows:

[0466] (704) Direct response information is calculated based on the relevant audio objects in each of the plurality of frequency intervals and the directional information associated with the relevant audio objects in the frequency intervals.

[0467] A23. The decoder according to Example A22

[0468] The audio renderer (700) is configured to: determine (741) the spread signal for each of the plurality of frequency intervals using spread information such as spread parameters or decorrelation rules included in the metadata, and combine the direct response determined by the direct response information and the spread signal to obtain the spectral domain rendered signal for the channel of the plurality of channels, or

[0469] The direct response information (704) and information (702) about the plurality of audio channels are used to calculate (706) synthesis information, and the covariance synthesis information is applied (727) to the one or more transmission channels to obtain the plurality of audio channels, or

[0470] Wherein, the direct response information (704) is the direct response vector of each relevant audio object, and wherein, the covariance synthesis information is a covariance synthesis matrix, and wherein, the audio renderer (700) is configured to perform matrix operations for each frequency range when applying the covariance synthesis information (727).

[0471] A24. The decoder according to example A22 or A23, wherein the audio renderer (700) is configured as follows:

[0472] In calculating the direct response information (704), a direct response vector for each relevant audio object is derived; and for each relevant audio object, a covariance matrix is ​​calculated based on each direct response vector.

[0473] When calculating the covariance synthesis information, the target covariance information (724) is derived from the following:

[0474] The covariance matrix of each of the related audio objects.

[0475] Regarding the power information of the relevant audio objects, and

[0476] Power information derived from the one or more transmission channels.

[0477] A25. The decoder according to Example A24, wherein the audio renderer (700) is configured as follows:

[0478] In calculating the direct response information (704), a direct response vector for each relevant audio object is derived; and for each relevant audio object, a covariance matrix (723) is calculated based on each direct response vector.

[0479] The input covariance information (726) is derived from the transmission channel, and

[0480] From the target covariance information, the input covariance information, and information about the multiple channels, derive the (725a, 725b) mixed information, and

[0481] The hybrid information is applied (727) to the transmission channel of each frequency interval in the time frame.

[0482] A26. The decoder according to Example A25, wherein the result of applying the mixed information to each frequency range in the time frame is transformed (708) into the time domain to obtain multiple audio channels in the time domain.

[0483] A27. The decoder according to one of examples A22 to A26, wherein the audio renderer (700) is configured as follows:

[0484] When decomposing (752) the input covariance matrix derived from the transmission channel, only the main diagonal elements of the input covariance matrix are used, or

[0485] The decomposition of the target covariance matrix is ​​performed using the direct response matrix and the power matrix of the object or transmission channel (751), or

[0486] The decomposition of the input covariance matrix is ​​performed by taking the root of each main diagonal element of the input covariance matrix, or...

[0487] Calculate the regularized inverse matrix of the decomposed input covariance matrix (753), or

[0488] Singular value decomposition (756) is performed when calculating the optimal matrix to be used for energy compensation, without expanding the identity matrix.

[0489] A28. A method for encoding a plurality of audio objects and related metadata indicating directional information about the plurality of audio objects, comprising:

[0490] Downmix the multiple audio objects to obtain one or more transmission channels;

[0491] Encode the one or more transmission channels to obtain one or more encoded transmission channels; and

[0492] The output includes coded audio signals from the one or more coded transmission channels.

[0493] The downmixing includes downmixing the plurality of audio objects in response to directional information about the plurality of audio objects.

[0494] A29. A method for decoding an encoded audio signal, the encoded audio signal comprising: one or more transmission channels and direction information of a plurality of audio objects, and parameter data of at least two related audio objects for one or more frequency intervals of a time frame, wherein the number of the at least two related audio objects is less than the total number of the plurality of audio objects, the decoding method comprising:

[0495] The one or more transmission channels provide a spectral representation having multiple frequency ranges within the time frame; and

[0496] The direction information is used to render the audio from the one or more transmission channels into multiple audio channels.

[0497] The audio rendering includes: for each of the one or more frequency intervals, calculating the contribution from the one or more transmission channels based on first direction information associated with a first related audio object among the at least two related audio objects and second direction information associated with a second related audio object among the at least two related audio objects, or considering the contribution from the one or more transmission channels based on the first direction information associated with the first related audio object among the at least two related audio objects and second direction information associated with the second related audio object among the at least two related audio objects.

[0498] A30. A computer program, when run on a computer or processor, for performing the method described according to Example A28 or the method described according to Example A29.

[0499] A31. An encoded audio signal, including information about parametric data of at least two related audio objects in one or more frequency ranges.

[0500] A32. The encoded audio signal according to Example A31 further includes:

[0501] One or more encoded transmission channels

[0502] Two or more coded object identifiers of the relevant audio object in each of one or more frequency intervals within a time frame, and one or more coded combination values ​​or coded amplitude-related measurements, serve as information about the parameter data.

[0503] The quantized and encoded directional data for each audio object in the time frame, wherein the directional data is constant for all frequency ranges in the one or more frequency ranges.

[0504] 2 References

[0505]

[0506]

[0507]

[0508]

[0509]

[0510]

[0511]

[0512]

[0513]

[0514]

[0515]

[0516]

[0517]

[0518]

[0519]

[0520]

[0521]

[0522]

Claims

1. A decoder for decoding an encoded audio signal, the encoded audio signal comprising: The decoder includes: one or more transmission channels and direction information of a plurality of audio objects, and parameter data of at least two related audio objects for one or more frequency intervals of a time frame, wherein the number of the at least two related audio objects is less than the total number of the plurality of audio objects, wherein the number of at least two related audio objects is selected from the total number of audio objects, wherein the total number of audio objects is not indicated as related, and the decoder comprises: Input interface (600) for providing the one or more transmission channels in a spectral representation having multiple frequency ranges in the time frame; and An audio renderer (700) is configured to render the one or more transmission channels into a plurality of audio channels using the direction information, such that contributions from the one or more transmission channels are considered based on first direction information associated with a first related audio object among the at least two related audio objects and second direction information associated with a second related audio object among the at least two related audio objects. The audio renderer (700) is configured to: for each of the one or more frequency ranges, calculate the contribution from the one or more transmission channels based on first direction information associated with a first related audio object among the at least two related audio objects and second direction information associated with a second related audio object among the at least two related audio objects.

2. The decoder according to claim 1, in, The audio renderer (700) is configured to ignore the directional information of audio objects that are different from the at least two related audio objects for the one or more frequency ranges.

3. The decoder according to claim 1, in, The encoded audio signal includes amplitude-related measurements (812) of each relevant audio object in the parameter data or combined values ​​(812) related to at least two relevant audio objects, and The audio renderer (700) is configured to determine (704) the quantitative contribution of the one or more transmission channels based on the amplitude-related measurement or the combined value.

4. The decoder according to claim 3, wherein, The encoded signal includes the combined values ​​in the parameter data, and The audio renderer (700) is configured to: determine the contribution of the one or more transmission channels (704, 733) using a combined value of one of the relevant audio objects and the direction information of that relevant audio object; and The audio renderer (700) is configured to determine (704, 735) the contribution of the one or more transmission channels using a combined value derived from the direction information of another related audio object in the one or more frequency ranges.

5. The decoder according to claim 1, wherein, The audio renderer (700) is configured as follows: (704) Direct response information is calculated based on the relevant audio objects in each of the plurality of frequency intervals and the directional information associated with the relevant audio objects in the frequency intervals.

6. The decoder according to claim 5, wherein, The audio renderer (700) is configured as follows: The diffusion information, such as diffusion parameters, or the decorrelation rules included in the metadata are used to determine (741) the diffusion signal for each of the plurality of frequency intervals, and the direct response determined by the direct response information and the diffusion signal are combined to obtain the spectral domain rendering signal of the channel among the plurality of channels.

7. The decoder according to claim 5, wherein, The audio renderer (700) is configured to: use the direct response information (704) and information (702) about the plurality of audio channels to calculate (706) synthesis information, and apply (727) the covariance synthesis information to the one or more transmission channels to obtain the plurality of audio channels.

8. The decoder according to claim 5, wherein, The direct response information (704) is a direct response vector for each relevant audio object, and wherein the covariance synthesis information is a covariance synthesis matrix, and wherein the audio renderer (700) is configured to perform matrix operations for each frequency range when applying the covariance synthesis information (727).

9. The decoder according to claim 5, wherein, The audio renderer (700) is configured as follows: In calculating the direct response information (704), a direct response vector for each relevant audio object is derived; and for each relevant audio object, a covariance matrix is ​​calculated based on each direct response vector. When calculating the covariance synthesis information, the target covariance information (724) is derived from the following: The covariance matrix of each of the related audio objects. Regarding the power information of the relevant audio objects, and Power information derived from the one or more transmission channels.

10. The decoder according to claim 9, wherein, The audio renderer (700) is configured as follows: In calculating the direct response information (704), a direct response vector for each relevant audio object is derived; and for each relevant audio object, a covariance matrix (723) is calculated based on each direct response vector. The input covariance information (726) is derived from the transmission channel, and From the target covariance information, the input covariance information, and information about the multiple channels, derive the (725a, 725b) mixed information, and The hybrid information is applied (727) to the transmission channel of each frequency interval in the time frame.

11. The decoder according to claim 10, wherein, The result of applying the mixed information to each frequency range in the time frame is transformed (708) into the time domain to obtain multiple audio channels in the time domain.

12. The decoder according to claim 5, wherein, The audio renderer (700) is configured as follows: When decomposing (752) the input covariance matrix derived from the transmission channel, only the main diagonal elements of the input covariance matrix are used.

13. The decoder according to claim 5, wherein, The audio renderer (700) is configured as follows: The decomposition of the target covariance matrix is ​​performed using the direct response matrix and the power matrix of the object or transmission channel (751).

14. The decoder according to claim 5, wherein, The audio renderer (700) is configured as follows: The decomposition of the input covariance matrix is ​​performed by taking the root of each diagonal element of the input covariance matrix (752).

15. The decoder according to claim 5, wherein, The audio renderer (700) is configured as follows: Calculate the regularized inverse matrix of the decomposed input covariance matrix (753).

16. The decoder according to claim 5, wherein, The audio renderer (700) is configured as follows: Singular value decomposition (756) is performed when calculating the optimal matrix to be used for energy compensation, without expanding the identity matrix.

17. A method for decoding an encoded audio signal, the encoded audio signal comprising: The decoding method comprises: one or more transmission channels and direction information of a plurality of audio objects, and parameter data of at least two related audio objects for one or more frequency intervals of a time frame, wherein the number of the at least two related audio objects is less than the total number of the plurality of audio objects, wherein the number of at least two related audio objects is selected from the total number of audio objects, wherein the total number of audio objects is not indicated as related; and the decoding method includes: The one or more transmission channels provide a spectral representation having multiple frequency ranges within the time frame; and The direction information is used to render the audio from the one or more transmission channels into multiple audio channels. The audio rendering includes: for each of the one or more frequency intervals, calculating the contribution from the one or more transmission channels based on first direction information associated with a first related audio object among the at least two related audio objects and second direction information associated with a second related audio object among the at least two related audio objects, or considering the contribution from the one or more transmission channels based on the first direction information associated with the first related audio object among the at least two related audio objects and second direction information associated with the second related audio object among the at least two related audio objects.

18. A computer program, when run on a computer or processor, for performing the method of claim 17.

19. An encoded audio signal comprising information about parametric data of at least two related audio objects regarding one or more frequency intervals of a plurality of frequency intervals associated with a time frame, wherein, The number of at least two related audio objects is selected from the total number of audio objects, wherein the total number of audio objects is not indicated as related.

20. The encoded audio signal according to claim 19, further comprising: One or more encoded transmission channels Two or more coded object identifiers of the relevant audio object in each of one or more frequency intervals within a time frame, and one or more coded combination values ​​or coded amplitude-related measurements, serve as information about the parameter data. The quantized and encoded directional data for each audio object in the time frame, wherein the directional data is constant for all frequency ranges in the one or more frequency ranges.

Citation Information

Patent Citations

  • Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to dirac based spatial audio coding

    WO2019068638A1

  • Parameter encoding and decoding

    WO2020249815A2