Title - AUDIO DATA CONVERTER, METHOD FOR CARRYING OUT THE CONVERSION, APPARATUS AND METHOD FOR CARRYING OUT AN AUDIO SYNTHESIS

AR125562B2Active Publication Date: 2026-08-26FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
ARP20220100655
Authority / Receiving Office
AR · AR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-10-01
Filing Date
2022-03-21
Publication Date
2026-08-26
Estimated Expiration
2038-10-04

AI Technical Summary

Technical Problem

Existing audio scene representations, such as channel-based, object-based, and scene-based formats, lack a universal encoding scheme that efficiently combines and manipulates complex audio scenes, particularly in 3D audio, and do not effectively support audio objects.

Method used

A universal parametric coding scheme based on Directional Audio Coding (DirAC) that processes and combines multi-channel signals, Ambisonics, and audio objects, allowing for interactive manipulation and efficient encoding of complex audio scenes.

Benefits of technology

Enables high-quality, interactive manipulation and efficient encoding of complex audio scenes by combining different audio representations into a single parametric format, supporting various playback configurations and allowing for selective manipulation of audio objects.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

An apparatus is provided for generating a combined audio scene description, comprising: an input interface (100) for receiving a first description of a first scene in a first format and a second description of a second scene in a second format, wherein the second format is different from the first format; a format converter (120) for converting the first description into a common format and for converting the second description into the common format, where the second format is different from the common format; and a format combiner (140) for combining the first description in the common format and the second description in the common format to obtain the combined audio scene.
Need to check novelty before this filing date? Find Prior Art

Description

APPARATUS AND METHOD FOR GENERATING A DESCRIPTION OF A COMBINED AUDIO SCENE Field of invention The present invention relates to audio signal processing and in particular to the processing of audio signals of audio descriptions of audio scenes. Introduction and state of the art: Transmitting a three-dimensional audio scene requires managing multiple channels, which typically generates a large amount of data to transmit. Furthermore, 3D sound can be represented in different ways: traditional channel-based sound, where each transmission channel is associated with a speaker position; sound carried through audio objects, which can be positioned in three dimensions independently of speaker positions; and scene-based (or Ambisonics), where the audio scene is represented by a set of coefficient signals that are the linear weights of spatially orthogonal basis functions, e.g., spherical harmonics.In contrast to channel-based rendering, scene-based rendering is independent of a specific speaker setup, and can be played back on any speaker configuration, at the expense of additional rendering processing in the decoder. For each of these formats, specific coding schemes were developed for the efficient storage or transmission of audio signals at low bitrates. For example, MPEG Surround is a parametric coding scheme for surround sound based on 1 1724882 of 109 channels, while MPEG Spatial Audio Object Coding (SAOC) is a parametric coding method dedicated to object-based audio. A parametric coding technique for higher-order Ambisonics was also provided in the recent MPEG-H Phase 2 standard. In this context, where all three audio scene representations—channel-based, object-based, and scene-based—are used and need to be supported, there is a need to design a universal scheme that allows for efficient parametric encoding of all three 3D audio representations. Furthermore, there is a need to be able to encode, transmit, and play back complex audio scenes composed of a mix of these different audio representations. Directional Audio Coding (DirAC) [1] is an effective approach for spatial sound analysis and reproduction. DirAC uses a perceptually motivated representation of the sound field based on direction of arrival (DOA) and diffusivity measured by frequency band. It is based on the assumption that at a given instant and within a critical band, the spatial resolution of the auditory system is limited to decoding one signal for direction and another for interaural coherence. Spatial sound is then represented in the frequency domain by the crossfading of two streams: a non-directional diffuse stream and a directional non-diffuse stream. DirAC was originally intended for recorded B-format sound, but it could also serve as a common format for mixing different audio formats. DirAC has already been extended to process the conventional 5.1 surround sound format in [3]. 1724882 of 109 proposes the merging of multiple DirAC streams in [4]. Furthermore, DirAC was extended to also support microphone inputs other than B format [6]. However, a universal concept is missing to make DirAC a universal representation of 3D audio scenes that is also capable of supporting the concept of audio objects. Some considerations were previously made for handling audio objects in DirAC. DirAC was employed in [5] as an acoustic front end for the Spatial Audio Encoder, SAOC, as a blind source separation for extracting multiple transmitters from a source mix. However, it was not envisioned to use DirAC itself as the spatial audio encoding scheme and to process audio objects directly along with their metadata and to potentially combine them with each other and with other audio representations. An objective of the present invention is to provide an improved concept for handling and processing audio scenes and audio scene descriptions. This objective is achieved by means of an apparatus for generating a description of a combined audio scene according to claim 1, a method for generating a description of a combined audio scene according to claim 14, or a related computer program according to claim 15. Furthermore, this objective is achieved by means of an apparatus for performing a synthesis of a plurality of audio scenes according to claim 16, a method for performing a synthesis of a plurality of audio scenes according to claim 20, or a related computer program according to claim 21. 1724882 of 109 This objective is further achieved by means of an audio data converter according to claim 22, a method for performing an audio data conversion according to claim 28, or a related computer program according to claim 29. Furthermore, this objective is achieved by means of an audio scene encoder according to claim 30, an audio scene encoding method according to claim 34, or a related computer program according to claim 35. Furthermore, this objective is achieved by means of an apparatus for performing an audio data synthesis according to claim 36, a method for performing an audio data synthesis according to claim 40, or a related computer program according to claim 41. The embodiments of the invention relate to a universal parametric coding scheme for 3D audio scenes based on the Directional Audio Coding (DirAC) paradigm, a perceptually motivated technique for spatial audio processing. Originally, DirAC was designed to analyze B-format recordings of audio scenes. The present invention aims to extend its capability to efficiently process any spatial audio format, such as channel-based audio, Ambisonics, audio objects, or a mixture thereof. DirAC playback can be easily generated for arbitrary speaker and headphone designs. The present invention also extends this output capability to Ambisonics, audio objects, or a mix of formats. More importantly, the invention allows the 1724882 of 109 possibility for the user to manipulate audio objects and to achieve, for example, an improvement of the dialogue at the decoder end. Context: System overview of a DirAC Spatial Audio Encoder The following is an overview of a novel immersive DirAC-based spatial audio coding system designed for Voice and Audio Services (IVAS). The goal of this system is to handle different spatial audio formats representing the audio scene, encode them at low bitrates, and reproduce the original audio scene as faithfully as possible after transmission. The system can accept different representations of audio scenes as input. The input audio scene can be captured by multi-channel signals intended to be played back at different speaker positions, by listening objects along with metadata describing the positions of the objects over time, or by a first-order or higher-order Ambisonics format representing the sound field at the listening or reference position. Preferably, the system is based on 3GPP Enhanced Voice Services (EVS), as the solution is expected to operate with low latency to enable conversation services on mobile networks. Figure 9 shows the encoder side of the DirAC-based spatial audio encoding system that supports different audio formats. As shown in Figure 9, the encoder (IVAS encoder) is capable of supporting different audio formats presented to the system separately or simultaneously. The audio signals can be 1724882 of 109, either acoustic in nature, picked up by microphones, or electrical in nature, intended to be transmitted to loudspeakers. Supported audio formats include multi-channel signals, first- and higher-order Ambisonics components, and audio objects. A complex audio scene can also be described by combining different input formats. All audio formats are then passed to DirAC analysis 180, which extracts a parametric representation of the entire audio scene. An arrival direction and a diffusivity measured per unit time-frequency form the parameters. DirAC analysis is followed by a spatial metadata encoder 190, which quantifies and encodes the DirAC parameters to obtain a low-bitrate parametric representation. Along with the parameters, a downmix signal derived from the various audio input sources or signals is encoded for transmission by a conventional audio core encoder. In this case, an EVS-based audio encoder is used for encoding the downmix signal. The downmix signal consists of different channels, called transport channels: the signal can be, for example, the four coefficient signals that make up a B-format signal, a stereo pair, or a mono downmix, depending on the target bit rate. The encoded spatial parameters and the encoded audio bitstream are multiplexed before being transmitted over the communication channel. Figure 10 shows a DirAC-based spatial audio coding decoder that outputs different audio formats. In the decoder shown in Figure 10, the transport channels are decoded by 1724882 of 109 the core decoder 1020, while the DirAC metadata is first decoded 1060 before being transported with the decoded transport channels to DirAC synthesis 220, 240. At this stage (1040), different options can be considered. It is possible to request that the audio scene be played directly on the speaker or headphone configurations, as is usually possible in a conventional DirAC system (MC in Fig. 10). In addition, it is also possible to request that the scene be rendered to an Ambisonics format for further manipulations, such as rotation, reflection, or movement of the scene (FOA / HOA in Fig. 10). Finally, the decoder can deliver the individual objects as they were presented on the encoder side (Objects in Fig. 10). Audio objects can also be restored, but it's more engaging for the listener to adjust the mix through interactive object manipulation. Common object manipulations include level adjustment, equalization, and spatial object placement. Object-based dialogue enhancement, for example, becomes a possibility offered by this interactivity feature. Finally, it's possible to output the original formats as they were presented at the encoder input. This could be a combination of audio channels and objects, or Ambisonics and objects. To achieve separate transmission of multiple channels and Ambisonics components, several instances of the described system could be used. The present invention is advantageous in that, especially with regard to the first aspect, it establishes a framework for combining different scene descriptions into a combined audio scene by means of a common format, which allows combining the different descriptions of 1724882 of 109 audio scenes. This common format can, for example, be format B, or it can be the pressure / speed signal representation format, or, preferably, it can also be the DirAC parameter representation format. This format is a compact format that also allows a significant amount of user interaction, on the one hand, and, on the other hand, is useful with respect to a bit rate required for the representation of an audio signal. According to a further aspect of the present invention, a synthesis of a plurality of audio scenes can be advantageously carried out by combining two or more different DirAC descriptions. Both of these different DirAC descriptions can be processed by combining the scenes in the parameter domain or, alternatively, by rendering each audio scene separately and then combining the audio scenes rendered from the individual DirAC descriptions in the spectral domain or, alternatively, in the time domain. This procedure allows for very efficient yet high-quality processing of different audio scenes that are to be combined into a single scene representation and, in particular, a single time-domain audio signal. An additional aspect of the invention is advantageous in that a particular set of useful audio data converted for the conversion of object metadata into DirAC metadata is derived, wherein this audio data converter can be used within the framework of the first, second, or third aspect or can also be applied independently of one another. The audio data converter 1724882 of 109 allows for the efficient conversion of audio object data, for example, a waveform signal for an audio object, and the corresponding data position, typically with respect to time, for the representation of a certain path of an audio object within a playback configuration into a very useful and compact audio scene description, and, in particular, the DirAC audio scene description format.While a typical audio object description with an audio object waveform signal and audio object position metadata is related to a particular playback setup or, more generally, to a certain playback coordinate system, the DirAC description is particularly useful because it is related to a listener or microphone position and is completely free of any limitations regarding a speaker configuration or playback setup. Therefore, the DirAC description generated from audio object metadata signals also allows a very useful, compact, and high-quality combination of audio objects, unlike other audio object combination technologies such as spatial audio object encoding or by means of amplitude panning of objects in a playback setup. An audio scene encoder according to a further aspect of the present invention is particularly useful for providing a combined representation of an audio scene having DirAC metadata and, additionally, an audio object with audio object metadata. In particular, in this situation, it is especially useful and advantageous for high interactivity in order to generate a combined description of metadata that have 1724882 of 109 DirAC metadata on one hand and, in parallel, object metadata on the other. Thus, in this respect, the object metadata is not combined with the DirAC metadata, but is converted into DirAC-like metadata such that the object metadata includes the direction or, additionally, a distance and / or diffusivity of the individual object along with the object signal. Therefore, the object signal is converted into a DirAC-like representation in such a way that very flexible handling of a DirAC representation for a first audio scene and an additional object within this first audio scene becomes possible. Thus, for example, specific objects can be processed very selectively because their corresponding transport channel is still available, on the one hand, and the DirAC style parameters, on the other. According to a further aspect of the invention, an apparatus or method for performing audio data synthesis is particularly useful because it provides a manipulator for manipulating a DirAC description of one or more audio objects, a DirAC description of a multi-channel signal, or a DirAC description of first-order Ambisonics signals or higher-order Ambisonics signals. The manipulated DirAC description is then synthesized using a DirAC synthesizer. This aspect has the particular advantage that any specific manipulation of any audio signal is carried out very usefully and efficiently within the DirAC domain; that is, by manipulating either the transport channel of the DirAC description or, alternatively, by manipulating the parametric data of the DirAC description. This modification is substantially more efficient and 1724882 of 109 is more practical to perform in the DirAC domain compared to manipulation in other domains. In particular, position-dependent weighting operations, as preferred manipulation operations, can be carried out specifically in the DirAC domain. Therefore, in a specific embodiment, converting a representation of the corresponding signal into the DirAC domain and then performing the manipulation within the DirAC domain is a particularly useful application scenario for processing and manipulating modern audio scenes. The preferred forms of realization are discussed later with regard to the accompanying figures, in which: Fig. 1a is a block diagram of a preferred implementation of an apparatus or method for generating a description of a combined audio scene according to a first aspect of the invention; Fig. 1b is an implementation of the generation of a combined audio scene, where the common format is the pressure / velocity representation; Fig. 1c is a preferred implementation of the generation of a combined audio scene, where the DirAC parameters and the DirAC description are in the common format; Fig. 1d is a preferred implementation of the combiner; Fig. 1c illustrates two different alternatives for implementing the DirAC parameter combiner for different audio scenes or audio scene descriptions. Fig. 1e is a preferred implementation of the generation of a combined audio scene where the common format is format B as an example for a 1724882 of 109 Ambisonics representation; Fig. 1f is an illustration of an audio / DirAC object converter useful in the context of, for example, Fig. 1c or 1d, or useful in the context of the third aspect relating to a metadata converter; Fig. 1g is an example illustration of a 5.1 multi-channel signal in a DirAC description; Fig. 1h is a further illustration of the conversion of a multi-channel format to the DirAC format in the context of an encoder and one side of the decoder; Fig. 2a illustrates an embodiment of an apparatus or a method for carrying out a synthesis of a plurality of audio scenes according to a second aspect of the present invention; Fig. 2b illustrates a preferred implementation of the DirAC synthesizer of Fig. 2a; Fig. 2c illustrates a further implementation of the DirAC synthesizer with a combination of rendered signals; Fig. 2d illustrates an implementation of a selective manipulator connected either before the scene combiner 221 of Fig. 2b or before the combiner 225 of Fig. 2c; Fig. 3a is a preferred implementation of an apparatus or method for realizing and converting audio data according to a third aspect of the present invention; Fig. 3b is a preferred implementation of the metadata converter also illustrated in Fig. 1f; Fig. 3c is a flowchart for the realization of a further implementation of an audio data conversion through the pressure / velocity domain; 1724882 of 109 Fig. 3d illustrates a flowchart for performing a merge within the DirAC domain; Fig. 3e illustrates a preferred implementation for realizing different DirAC descriptions, for example, as illustrated in Fig. 1d with respect to the first aspect of the present invention; Fig. 3f illustrates the conversion of an object position data into a DirAC parametric representation; Fig. 4a illustrates a preferred implementation of an audio scene encoder according to a fourth aspect of the present invention for generating a combined metadata description comprising DirAC metadata and object metadata; Fig. 4b illustrates a preferred embodiment with respect to the fourth aspect of the present invention; Fig. 5a illustrates a preferred implementation of an apparatus for performing an audio data synthesis or a corresponding procedure according to a fifth aspect of the present invention; Fig. 5b illustrates a preferred implementation of the DirAC synthesizer of Fig. 5A; Fig. 5c illustrates an additional alternative to the manipulator procedure of Fig. 5A; Fig. 5d illustrates an additional procedure for implementing the manipulator of Fig. 5A; Fig. 6 illustrates an audio signal converter for generating, from a mono signal and one direction of arrival information, i.e., from an example DirAC description, where the diffusivity, for example, is set to zero, a B-format representation comprising an omnidirectional component and directional components in X, Y, and Z; Fig. 7a illustrates an implementation of a 1724882 of 109 DirAC analysis of a microphone format B signal; Fig. 7b illustrates an implementation of a DirAC synthesis according to a known procedure; Fig. 8 illustrates a flowchart for the illustration of additional embodiments of, in particular, the embodiment of Fig. 1a; Fig. 9 is the encoder side of the DirAC-based spatial audio coding that supports different audio formats; Fig. 10 is a DirAC-based spatial audio coding decoder that delivers different audio formats; Fig. 11 is an overview of the DirAC-based encoder / decoder system that combines different input formats into a combined B format; Fig. 12 is an overview of the DirAC-based encoder / decoder system that combines in the pressure / velocity domain; Fig. 13 is an overview of the DirAC-based encoder / decoder system that combines different input formats in the DirAC domain with the possibility of object manipulation on the decoder side; Fig. 14 is an overview of the DirAC-based encoder / decoder system that combines different input formats on the decoder side via a DirAC metadata combiner; Figure 15 is an overview of the DirAC-based encoder / decoder system, the combination of different input formats on the decoder side in DirAC synthesis; and Fig. 16a af illustrates several representations of audio formats useful in the context of the first to fifth aspects of the present invention. 1724882 of 109 Figure 1a illustrates a preferred embodiment of an apparatus for generating a combined audio scene description. The apparatus comprises an input interface 100 for receiving a first description of a first scene in a first format and a second description of a second scene in a second format, wherein the second format differs from the first format. The format can be any audio scene format, such as any of the formats or scene descriptions illustrated in Figures 16a to 16f. Fig. 16a, for example, illustrates an object description that typically consists of a waveform signal of an object (encoded) 1 such as a mono channel and corresponding metadata relating to the position of object 1, where this information is typically given for each time frame or group of time frames, and the waveform signal of object 1 is encoded. Corresponding representations for a second or more objects may be included as illustrated in Fig. 16a. Another alternative is an object description consisting of a downmix of objects that are either a mono signal, a two-channel stereo signal, or a signal with three or more channels, along with related object metadata such as object energies, time / frequency compartment correlation information, and optionally, object positions. However, object positions can also be provided on the decoder side as typical rendering information and can therefore be modified by a user. The format in Fig. 16b can, for example, be implemented as the well-known SAOC (Spatial Audio Object Coding) format. Another description of a scene is illustrated in Fig. 1724882 of 109 16c as a description of multi-channel audio that has both encoded and unencoded representations of a first channel, a second channel, a third channel, a fourth channel, and a fifth channel, where the first channel can be the left channel (L), the second channel can be the right channel (R), the third channel can be the center channel (C), the fourth channel can be the left surround channel (LS), and the fifth channel can be the right surround channel (RS). Naturally, the multi-channel signal can have a smaller or larger number of channels, such as only two channels for a stereo channel, six channels for a 5.1 format, eight channels for a 7.1 format, and so on. A more efficient representation of a multi-channel signal is illustrated in Fig. 16d, in which the channel downmix, such as a mono downmix, a stereo downmix, or a downmix with more than two channels, is associated with parametric side information as channel metadata for, typically, each time and / or frequency compartment. Such a parametric representation can, for example, be implemented according to the MPEG surround sound standard. Another representation of an audio scene could be, for example, format B, which consists of an omnidirectional W signal and directional X, Y, and Z components, as shown in Fig. 16e. This would be a first-order or FoA signal. A higher-order Ambisonics signal, i.e., a HoA signal, may have additional components, as is known in the art. The representation in Fig. 16e, in contrast to the representations in Fig. 16c and Fig. 16d, is not dependent on a specific speaker configuration, but rather describes a sound field based on the experience at a given position. 1724882 of 109 determined (microphone or the listener). Another such description of the sound field is the DirAC format, for example, as illustrated in Fig. 16f. The DirAC format typically comprises a DirAC downmix signal, which is either a mono or stereo downmix signal, or a transport signal, and the corresponding parametric side information. This parametric side information is, for example, arrival direction information per time / frequency compartment and, optionally, diffusivity information per time / frequency compartment. The input at input interface 100 in Fig. 1a can be, for example, in any of the formats illustrated with respect to Fig. 16a to Fig. 16f. Input interface 100 forwards the corresponding format descriptions to a format converter 120. The format converter 120 is configured to convert the first description to a common format and to convert the second description to the same common format when the second format is different from the common format. When, however, the second format is already in the common format, the format converter only converts the first description to the common format, since the first description is already in a different format. Therefore, at the output of the format converter, or more generally at the input of a format combiner, there is indeed a representation of the first scene in the common format and a representation of the second scene in the same common format. Because both descriptions are now contained within one and the same common format, the format combiner can now combine the first and second descriptions. 1724882 of 109 to obtain a combined audio scene. According to an embodiment illustrated in Fig. 1e, the format converter 120 is configured to convert the first description into a first B format signal, for example, as illustrated in 127 in Fig. 1e, and to calculate the B format representation for the second description as illustrated in 128 in Fig. 1e. Then, the format combiner 140 is implemented as a signal component adder illustrated in 146a for the W adder component, 146b for the X adder component, illustrated in 146c for the Y adder component, and illustrated in 146d for the Z adder component. Therefore, in the embodiment of Fig. 1e, the combined audio scene can be a B-format representation, and the B-format signals can then function as the transport channels and be encoded via a transport channel encoder 170 of Fig. 1a. Thus, the combined audio scene with respect to the B-format signal can be directly input to the encoder 170 of Fig. 1a to generate an encoded B-format signal, which could then be output via the output interface 200. In this case, no spatial metadata is required, but at the cost of an encoded representation of four audio signals, namely the omnidirectional component W and the directional components X, Y, and Z. Alternatively, the common format is the pressure / velocity format, as illustrated in Fig. 1b. For this purpose, the format converter 120 comprises a time / frequency analyzer 121 for the first audio scene and the time / frequency analyzer 122 for the second audio scene or, in general, the audio scene with number N, where N is a 1724882 of 109 whole number. Then, for each such spectral representation generated by the spectral converters 121, 122, the pressure and velocity are calculated according to what is illustrated in 123 and 124, and the format combiner below is configured to calculate a summed pressure signal on one side, by means of the sum of the corresponding pressure signals generated by blocks 123, 124. And, in addition, an individual velocity signal is also calculated for each of the blocks 123, 124 and the velocity signals can be summed together in order to obtain a combined pressure / velocity signal. Depending on the application, the procedures in blocks 142 and 143 do not necessarily have to be carried out. Instead, the combined or summed pressure signal and the combined or summed velocity signal can be encoded in an analogy according to the format B signal illustrated in Fig. 1e. This pressure / velocity representation could then be encoded again through the encoder 170 of Fig. 1a and transmitted to the decoder without any additional information regarding spatial parameters, since the combined pressure / velocity representation already includes the spatial information necessary to obtain a high-quality sound field that is ultimately rendered on one side of the decoder. In one embodiment, however, a DirAC analysis is preferred to the pressure / velocity representation generated by block 141. To this end, the intensity vector 142 is calculated, and in block 143, the DirAC parameters are calculated from the intensity vector, and then the DirAC parameters 19 The combined 1724882 of 109 are obtained as a parametric representation of the combined audio scene. To this end, the DirAC analyzer 180 of Fig. 1a is implemented to perform the functionality of blocks 142 and 143 of Fig. 1b. Preferably, the DirAC data are further subjected to a metadata encoding operation in the metadata encoder 190. The metadata encoder 190 typically comprises a quantizer and entropy encoder in order to reduce the bit rate required for transmitting the DirAC parameters. Along with the encoded DirAC parameters, an encoded transport channel is also transmitted. The encoded transport channel is generated by the transport channel generator 160 of Fig. 1a, which can, for example, be implemented as illustrated in Fig. 1b by a first downmix generator 161 for generating a downmix of the first audio scene and an Nth downmix generator 162 for generating a downmix of the Nth audio scene. Next, the downmix channels are combined in combiner 163 in a typical way by direct addition, and the combined downmix signal is then the transport channel that is encoded by encoder 170 in Fig. 1a. The combined downmix can, for example, be a stereo pair, i.e., a first and second channel from a stereo representation, or it can be a mono channel, i.e., a single channel signal. According to a further embodiment illustrated in Fig. 1c, a format conversion in the 120 format converter is performed to directly convert each of the input audio formats into the DirAC format as the common format. For this purpose, 1724882 of 109 The format converter 120 again performs a time-frequency conversion or time / frequency analysis in the corresponding blocks 121 for the first scene and block 122 for a second or additional scene. The DirAC parameters are then derived from the spectral representations of the corresponding audio scenes illustrated in 125 and 126. The result of the procedure in blocks 125 and 126 are DirAC parameters consisting of energy information per time / frequency tile, an arrival direction information eDOA per time / frequency tile, and diffusivity information ψ for each time / frequency tile. The format combiner 140 is then configured to perform a combination directly in the DirAC parameter domain in order to generate combined DirAC parameters Ψ for diffusivity and eDOA for arrival direction.In particular, the E1 and EN energy information is required by combiner 144, but is not part of the final combined parametric representation generated by format combiner 140. Therefore, comparing Fig. 1c to Fig. 1e reveals that when the format combiner 140 already performs a combination in the DirAC parameter domain, the DirAC analyzer 180 is unnecessary and is not implemented. Instead, the output of the format combiner 140, which is the output of block 144 in Fig. 1c, is forwarded directly to the metadata encoder 190 in Fig. 1a and from there to the output interface 200, such that the encoded spatial metadata, and in particular the encoded combined DirAC parameters, are included in the output signal encoded by the output interface 200. Furthermore, the transport channel generator 160 in Fig. 1a can receive, from the input interface 100, 1724882 of 109 a waveform signal representation for the first scene and the waveform signal representation for the second scene. These representations are fed into the downmix generator blocks 161, 162 and the results are summed in block 163 to obtain a combined downmix as illustrated with respect to Fig. 1b. Figure 1d illustrates a similar representation to that shown in Figure 1c. However, in Figure 1d, the audio object waveform is input into time / frequency representation converter 121 for audio object 1 and 122 for audio object N. In addition, the metadata, along with the spectral representation, is input into the DirAC parameter calculators 125 and 126, as also illustrated in Figure 1c. However, Fig. 1D provides a more detailed representation of how the preferred implementations of combiner 144 operate. In a first alternative, the combiner performs an energy-weighted sum of the individual diffusivity for each individual object or scene, and a corresponding energy-weighted calculation of a combined DoA for each time / frequency tile is carried out as illustrated in the lower equation of alternative 1. However, other implementations are also possible. In particular, another highly efficient calculation sets the diffusivity to zero for the combined DirAC metadata and selects, as the arrival direction for each time / frequency tile, the arrival direction calculated from a given audio object that has the highest energy within that specific time / frequency tile. The procedure in Fig. 1d is preferably more appropriate when the input to the input interface consists of audio objects. 1724882 of 109 individuals correspondingly representing a mono waveform or signal for each object and corresponding metadata, such as position information illustrated with respect to Fig. 16a or 16b. However, in the embodiment of Fig. 1c, the audio scene can be any of the representations illustrated in Figs. 16c, 16d, 16e, or 16f. Metadata may or may not be present; that is, the metadata in Fig. 1c is optional. A typically useful diffusivity is then calculated for a given scene description, such as an Ambisonics scene description in Fig. 16e, and the first alternative for combining the parameters is preferred over the second alternative in Fig. 1d. Therefore, according to the invention, the format converter 120 is configured to convert a higher-order Ambisonics format or a first-order Ambisonics format into format B, wherein the higher-order Ambisonics format is truncated before being converted into format B. In a further embodiment, the format converter is configured to project an object or channel onto a reference position in spherical harmonics to obtain the projected signals, and the format combiner is configured to combine the projected signals to obtain coefficients in B format, wherein the object or channel is located in space at a specified position and has an optional individual distance from a reference position. This procedure works particularly well for converting object signals or multi-channel signals into first-order or higher-order Ambisonics signals. Alternatively, the format converter 1724882 of 109 120 is configured to perform a DirAC analysis comprising a time-frequency analysis of the components in B format and a determination of the pressure and velocity vectors, wherein the format combiner is then configured for combining different pressure / velocity vectors, and wherein the format combiner further comprises the DirAC analyzer 180 for deriving DirAC metadata from the combined pressure / velocity data. In a further alternative embodiment, the format converter is configured to extract the DirAC parameters directly from the object metadata of an audio object format such as the first or second format, in which the pressure vector for DirAC rendering is the object waveform signal and the direction is derived from the object's position in space or the diffusivity is directly given in the object metadata or set to a default value such as zero. In a further embodiment, the format converter is configured to convert the DirAC parameters derived from the object data format into pressure / speed data, and the format combiner is configured to combine the pressure / speed data with pressure / speed data derived from different descriptions of one or more different audio objects. However, in a preferred implementation illustrated with respect to Fig. 1c and 1d, the format combiner is configured to directly combine the DirAC parameters derived by the format converter 120 such that the combined audio scene generated by block 140 in Fig. 1a is either the final result and a DirAC analyzer 180 illustrated in Fig. 1a is not required, since the data output by the combiner 1724882 of 109 formats 140 is already in DirAC format. In a further implementation, the format converter 120 already comprises a DirAC analyzer for a first-order Ambisonics or higher-order Ambisonics input format or a multi-channel signal format. Furthermore, the format converter comprises a metadata converter for converting object metadata into DirAC metadata. Such a metadata converter is illustrated, for example, in Fig. 1f in 150, which operates again in the time / frequency analysis in block 121 and calculates the energy per band per time frame, as illustrated in 147. The arrival direction is illustrated in block 148 of Fig. 1f, and the diffusivity is illustrated in block 149 of Fig. 1f.And, the metadata are combined by the combiner 144 for the combination of the individual DirAC metadata streams, preferably by means of a weighted sum as illustrated by example by one of the two embodiment alternatives in Fig. 1d. Multi-channel signals can be directly converted to B format. The resulting B format can then be processed by a conventional DirAC. Fig. 1g illustrates a conversion to B format and subsequent DirAC processing. Reference [3] describes methods for performing multi-channel signal conversion to B-format. In principle, converting multi-channel audio signals to B-format is simple: virtual loudspeakers are defined to be in different speaker layout positions. For example, for a 5.0 layout, the loudspeakers are positioned in the horizontal plane at azimuth angles of + / - 30 and + / - 110 degrees. A virtual B-format microphone is then defined to be in the center of the loudspeakers, and a virtual recording is made. 1724882 of 109. Therefore, channel W is created by summing all the speaker channels of the 5.0 audio file. The process for obtaining W and other coefficients in B format can be summarized below: k Ύ = Wj S; (cos(g¿) É = 1 k V = ^wÉsÉ(sm($)cos(pÉ)) É = 1 k Z = ^wÉsÉ(sin(pÉ)) É = 1 where si are the multi-channel signals located in space at the speaker positions defined by the azimuth angle θι and the elevation angle φι of each speaker, and wi are weights based on distance. If the distance is unavailable or simply ignored, then wi = 1. However, this simple technique is limited, as it is an irreversible process. Furthermore, since speakers are generally non-uniformly distributed, there is also a bias in the estimate made by a subsequent DirAC analysis toward the direction with the highest speaker density. For example, in a 5.1 design, there will be a bias toward the front, since there are more speakers at the front than at the rear. To address this problem, an additional technique was proposed in [3] for multi-channel 5.1 signal processing with DirAC. The final encoding scheme will then look like that illustrated in Fig. 1h, which shows the format converter B 127, the DirAC analyzer 180 as generally described with respect to element 180 in Fig. 1, and 1724882 of 109 the other elements 190, 1000, 160, 170, 1020, and / or 220, 240. In a further embodiment, the 200 output interface is configured to add, to the combined format, a separate object description for an audio object, where the object description comprises at least one of a direction, a distance, a diffusivity, or any other object attribute, where this object has a single direction across all frequency bands and is either already static or in motion at a rate slower than a velocity threshold. This feature is further elaborated in more detail with respect to the fourth aspect of the present invention described with respect to Fig. 4a and Fig. 4b. First Coding Alternative: Combination and processing of different audio representations through the B format or an equivalent representation. A first form of implementation of the intended encoder can be achieved by converting all input formats into a combined B format as shown in Fig. 11. Fig. 11: Overview of the DirAC-based encoder / decoder system that combines different input formats into a combined B format Since DirAC is originally designed for analyzing a B-format signal, the system converts the various audio formats into a combined B-format signal. The formats are first individually converted to a B-format signal before being combined by summing their BW, X, Y, Z components. First-Order Ambisonics (FOA) components can be normalized and reordered into a B-format. Assuming the FOA is in ACN / N3D format, 1724882 of 109 The four signals of the input format B are obtained by means of: ' iv = y0° ' I2-I y = β71 Where Y1m denotes the Ambisonics component of order 1 and the index m, -1 m +1. Since the FOA components are fully contained in the higher-order Ambisonics format, the HOA format only needs to be truncated before being converted to the B format. Since the objects and channels have determined positions in space, it is possible to project each individual object and channel in spherical harmonics (SH) onto the central position, such as the recording or reference position. The sum of these projections allows different objects and multiple channels to be combined into a single B-format, which can then be processed by DirAC analysis. The coefficients in B-format (W, X, Y, Z) are then given by: k r1=Σ F'Sit=l Y k X = (cos(^t) C0^É)) t = lk Y = ^wrÉsÉ(sin(^)cos(^É)) É = 1 k Z = ^wÉsÉ(siii(pÉ)) É = 1 where si are independent signals located in space in the positions defined by the azimuth angle 1724882 of 109 θί and the elevation angle φι, and wi are weights as a function of distance. If the distance is not available or is simply ignored, then wi = 1. For example, the independent signals correspond to audio objects located at the given position or the signal associated with a speaker channel at the specified position. In applications where an Ambisonics representation of orders higher than the first order is desired, the generation of Ambisonics coefficients presented above for the first order is extended by means of the additional consideration of higher order components. The Transport Channel Generator 160 can directly receive multi-channel signals, object waveform signals, and higher-order Ambisonics components. The Transport Channel Generator reduces the number of input channels to be transmitted by downmixing them. The channels can be mixed together, as in MPEG surround, into a mono or stereo downmix, while object waveform signals can be passively synthesized into a mono downmix. Furthermore, from higher-order Ambisonics, it is possible to extract a lower-order representation or create, via beamforming, a stereo downmix or any other spatial segmentation. If the downmixes obtained from the different input formats are compatible, they can be combined using a simple summing operation. Alternatively, transport channel generator 160 can receive the same combined B format as that transmitted to the DirAC analyzer. In this case, a subset of the components or the result of beamforming (or other processing) forms the 1,724,882 of 109 transport channels to be encoded and transmitted to the decoder. In the proposed system, conventional audio encoding is required, which may be based on, but is not limited to, the 3GPP EVS standard encoding. 3GPP EVS is the preferred codec choice due to its ability to encode speech or music signals at low bit rates with high quality while requiring relatively low latency, thus enabling real-time communication. At very low bit rates, the number of channels to be transmitted must be limited to one, and therefore only the omnidirectional microphone signal W of format B is transmitted. If the bit rate permits, the number of transport channels can be increased by selecting a subset of the format B components. Alternatively, the format B signals can be combined in a 160 beamformer directed at specific partitions of space. As an example, two cardioids can be designed to point in opposite directions, for example, to the left and right of the spatial scene: [L = V2IV + Y U = V2W - Y These two stereo channels L and R can then be efficiently encoded by a joint stereo encoding. The two signals will then be appropriately exploited by DirAC Synthesis on the decoder side for soundstage representation. Other beamforming can be envisaged; for example, a virtual cardioid microphone can be pointed in any direction from a given azimuth Θ and elevation φ. C = V2W + cqs((?) cos(^) Y + sin(í) co^) Y + sin (φ)Ζ Other ways of forming channels can be foreseen 1724882 of 109 transmissions that carry more spatial information than a single monophonic transmission channel would. Alternatively, the four coefficients of format B can be transmitted directly. In that case, the DirAC metadata can be extracted directly on the decoder side, without the need to transmit additional information for the spatial metadata. Figure 12 shows another alternative method for combining the different input formats. Figure 12 is also an overview of the DirAC-based encoder / decoder system combined in the pressure / velocity domain. Both multi-channel and Ambisonics signal components are introduced into a DirAC analysis 123, 124. For each input format, a DirAC analysis is performed consisting of a time-frequency analysis of the components in BY format and the determination of the pressure and velocity vectors: PÉ(n,k) = С^ΟύΟ UÉ(n,k) = + Pfon)®,. -F'XvIq where i is the input index and, kyn are the time and frequency indices of the time-frequency mosaic, and ex, ey, ez represent the Cartesian unit vectors. P(n,k) and U(n,k) are needed to calculate the DirAC parameters, namely DOA and diffusivity. The DirAC metadata combiner can exploit N sources playing together, resulting in a linear combination of their pressures and particle velocities measured when played individually. The combined quantities are then derived by: jV Pfn.k) = ^R^η,^ t = l N U(nrk) = ^L7É0i,l·) 1724882 of 109 The combined DirAC parameters are calculated 143 through the calculation of the combined intensity vector: where (.) denotes a complex conjugation. The diffusivity of the combined sound field is given by: , IIE(7(M»II =1where E{.} designates the time averaging operator, c the speed of sound and E(k,n) the sound field energy given by:Etican) E(k, n) = || U(k,«) ||2+ ^ | Ρ&,η) |2 The direction of arrival (DOA) is expressed by means of the unit vector, defined as ||7(fc,n)|| If an audio object is imported, the DirAC parameters can be extracted directly from the object metadata, while the pressure vector Pi(k,n) is the object's core signal (waveform). More precisely, the direction is derived directly from the object's position in space, while the diffusivity is directly given in the object metadata or (if unavailable) can be set to zero by default. From the DirAC parameters, the pressure and velocity vectors are directly given by: PÉ(k,n) = J1 - ^^k^P^k,^ ΰ'&,η) =--— PÉ(fc,n).e¿Oj4(fc,n) Poc The combination of objects or the combination of an object with different input formats below, is obtained by adding the pressure and velocity vectors as explained above. In summary, the combination of different The analysis of 1724882 of 109 input contributions (Ambisonics, channels, objects) is performed in the pressure / velocity domain, and the result is then subsequently converted into DirAC direction / diffusivity parameters. Operating in the pressure / velocity domain is theoretically equivalent to operating in B format. The main benefit of this alternative compared to the previous one is the ability to optimize the DirAC analysis according to each input format, as proposed in [3] for the 5.1 surround sound format. The main drawback of such a fusion in a combined B format or a pressure / velocity domain is that the conversion occurring at the front end of the processing chain is already a bottleneck for the entire encoding system. In effect, converting higher-order Ambisonics audio representations, objects, or channels to a B-format (first-order) signal already introduces a significant loss of spatial resolution that cannot be recovered afterward. Second Coding Alternative: combination and processing in the DirAC domain To overcome the limitations of converting all input formats into a combined B-format signal, this alternative proposes deriving the DirAC parameters directly from the original format and then combining them later in the DirAC parameter domain. An overview of such a system is shown in Fig. 13. Fig. 13 is an overview of the DirAC-based encoder / decoder system that combines different input formats in the DirAC domain with the possibility of object manipulation on the decoder side. In what follows, we can also consider channels 1724882 of 109 individual multi-channel signals as an audio object input for the encoding system. The object metadata is then static in time and represents the speaker's position and distance relative to the listener's position. The aim of this alternative solution is to avoid the systematic combination of different input formats into a combined B format or equivalent representation. The goal is to calculate the DirAC parameters before combining them. This method avoids any bias in the direction and diffusivity estimation caused by the combination. Furthermore, the characteristics of each audio representation can be optimally exploited during DirAC analysis or during the determination of the DirAC parameters. The combination of DirAC metadata occurs after determining the DirAC parameters, diffusivity, and direction for each input format, as well as the pressure contained in the transmitted transport channels. DirAC analysis can estimate the parameters of an intermediate format B, obtained by converting the input format as previously explained. Alternatively, the DirAC parameters can be advantageously estimated directly from the input format without going through format B, which could further improve the accuracy of the estimate. For example, in [7], it is proposed to estimate the direct diffusivity of higher-order Ambisonics. In the case of audio objects, a simple metadata converter in Fig. 15 can extract the object metadata direction and diffusivity for each object. The combination of the various DirAC metadata streams into a single combined DirAC metadata stream can be achieved in accordance with what is proposed in 1724882 of 109 [4] For some content, it is much better to directly estimate the DirAC parameters from the original format rather than converting to a combined B format first before performing a DirAC analysis. Indeed, the parameters—direction and diffusivity—can be skewed when converted to a B format [3] or during the combination of different sources. Furthermore, this alternative allows for... Another simpler alternative can average the parameters of the different sources by weighting them according to their energies: N=Eί&,η) ψ^^η) N For each object, it is still possible to send its own address and, optionally, its distance, diffusivity, or any other relevant attribute as part of the bitstream transmitted from the encoder to the decoder (see, e.g., Figs. 4a, 4b). This additional side information enriches the combined DirAC metadata and allows the decoder to reconstruct and / or manipulate the object separately. Since an object has only one address across all frequency bands and can be considered either static or moving slowly, this additional information needs to be updated less frequently than other DirAC parameters and only generates a very small additional bit rate. On the decoder side, directional filtering can be carried out as taught in [5] for object manipulation. Directional filtering is based on a short-term spectral attenuation technique. It is performed in the spectral domain by a function of 35 1724882 of 109 zero-phase gain, which depends on the direction of the objects. The direction can be contained in the bit stream if the object addresses are transmitted as side information. Otherwise, the address could also be shaped interactively by the user. Third alternative: decoder-side combination Alternatively, the combination can be performed on the decoder side. Figure 14 shows an overview of the DirAC-based encoder / decoder system combining different input formats on the decoder side via a DirAC metadata combiner. In Figure 14, the DirAC-based encoding scheme operates at higher bit rates than before, but allows the transmission of individual DirAC metadata. The different DirAC metadata streams are combined, for example, as proposed in [4] in the decoder before DirAC synthesis. The DirAC metadata combiner can also obtain the position of an individual object for subsequent object manipulation in DirAC analysis. Figure 15 shows an overview of the DirAC-based encoder / decoder system that combines different input formats from the decoder side into DirAC synthesis. Bitrate permitting, the system can be further enhanced, as shown in Figure 15, by sending each input component (FOA / HOA, MC, Object) its own downmix signal along with its associated DirAC metadata. Even so, the different DirAC streams share a common DirAC synthesis 220, 240 in the decoder to reduce complexity. 1724882 of 109 Figure 2a illustrates a concept for synthesizing a plurality of audio scenes according to a second additional aspect of the present invention. An apparatus illustrated in Figure 2a comprises an input interface 100 for receiving a first DirAC description of a first scene and for receiving a second DirAC description of a second scene, and one or more transport channels. In addition, a DirAC 220 synthesizer is provided for synthesizing a plurality of audio scenes in a spectral domain to obtain an audio signal in the spectral domain representing the plurality of audio scenes. Furthermore, a spectral time converter 214 is provided, which converts the audio signal in the spectral domain into a time domain in order to output a time-domain audio signal that can be output to loudspeakers, for example. In this case, the DirAC synthesizer is configured to perform the rendering of the loudspeaker output signal. Alternatively, the audio signal could be a stereo signal that can be output to headphones. Again, alternatively, the output of the audio signal by the spectral time converter 214 can be a description of the sound field in B format.All these signals, i.e., speaker signals for more than two channels, headphone signals, or descriptions of sound fields, are time-domain signals for further processing, such as output through speakers or headphones, or for transmission or storage in the case of descriptions of sound fields, such as first-order Ambisonics signals or higher-order Ambisonics signals. Furthermore, the device in Fig. 2a additionally comprises a user interface 260 for the 1724882 of 109 control of the DirAC synthesizer 220 in the spectral domain. In addition, one or more transport channels can be provided to the input interface 100 to be used together with the first and second DirAC descriptions which are, in this case, the parametric descriptions that provide, for each time / frequency tile, arrival direction information and, optionally and additionally, diffusivity information. Typically, the input of two different DirAC descriptions to interface 100 in Fig. 2a describes two different audio scenes. In this case, the DirAC synthesizer 220 is configured to perform a combination of these audio scenes. An alternative combination is illustrated in Fig. 2b. In this case, a scene combiner 221 is configured to combine the two DirAC descriptions in the parametric domain; that is, the parameters are combined to obtain a combined address of arrival (DoA) parameters and optionally diffusivity parameters at the output of block 221. This data is then fed into the DirAC renderer 222, which additionally receives one or more transport channels to obtain the audio signal in the spectral domain. The combination of the DirAC parametric data is preferably carried out as illustrated in Fig.1d and, in accordance with what has been described with respect to this figure and, in particular, with respect to the first alternative. In the event that at least one of the two description inputs in scene combiner 221 includes diffusivity values ​​of zero or no diffusivity values ​​at all, then, additionally, the second alternative may be applied, also in accordance with what was discussed in the context of Fig. 1d. Another alternative is illustrated in Fig. 2c. In this 1724882 of 109 procedure, the individual DirAC descriptions are rendered by means of a first DirAC renderer 223 for the first description and a second DirAC renderer 224 for the second description and at the output of blocks 223 and 224, a first and second spectral domain audio signal are available, and these first and second spectral domain audio signals are combined within combiner 225 to obtain, at the output of combiner 225, a spectral domain combination signal. As an example, the first DirAC renderer 223 and the second DirAC renderer 224 are configured to generate a stereo signal with a left (L) channel and a right (R) channel. Then, combiner 225 is configured to combine the left channel from block 223 and the left channel from block 224 to obtain a combined left channel. Additionally, the right channel from block 223 is combined with the right channel from block 224, resulting in a combined right channel at the output of block 225. For individual channels of a multi-channel signal, the same procedure is followed; that is, the individual channels are added one by one, such that the same channel is always added from one DirAC 223 renderer to the corresponding channel of the other DirAC renderer, and so on. The same procedure is also followed for, for example, higher-order Ambisonics signals or signals in B format. When, for example, the first DirAC 223 renderer outputs the W, X, Y, Z signals, and the second DirAC 224 renderer outputs a similar format, the combiner combines the two omnidirectional signals to obtain a combined omnidirectional W signal, and the same procedure is followed. 1724882 of 109 for the corresponding components in order to finally obtain a combined X, Y and Z component. Furthermore, as described above with respect to Fig. 2a, the input interface is configured to receive additional audio object metadata for an audio object. This audio object may already be included in the first or second DirAC description or may be separate from the first and second DirAC descriptions. In this case, the DirAC synthesizer 220 is configured to selectively manipulate the additional audio object metadata or object data related to this additional audio object metadata to, for example, perform directional filtering based on the additional audio object metadata or based on user-supplied direction information obtained from user interface 260. Alternatively or additionally, and as illustrated in Fig.2d, the DirAC 220 synthesizer is configured to perform, in the spectral domain, a zero-phase gain function. The zero-phase gain function depends on the address of an audio object, where the address is contained in a bitstream if the object addresses are transmitted as side information, or where the address is received from the user interface 260. The additional audio object metadata entered into the interface 100 as an optional feature in Fig. 2a reflects the possibility of still sending, for each individual object, its own address and optionally the distance, diffusivity, and any other relevant object attributes as part of the bitstream transmitted from the encoder to the decoder. Therefore, the additional audio object metadata can be related to an object already included in the first. 1724882 of 109 DirAC description or in the second DirAC description or is an additional object not included in the first DirAC description and already in the second DirAC description. However, it is preferable to have the metadata for additional audio objects already in a DirAC style, that is, an arrival direction and, optionally, diffusivity information. Since typical audio objects have zero diffusivity, meaning they are concentrated at their actual position, this results in a specific, concentrated arrival direction that is constant across all frequency bands and is, with respect to the frame rate, either static or moving at a slow rate. Therefore, since an object has a single direction across all frequency bands and can be considered either static or moving at a slow rate, the additional information needs to be updated less frequently than other DirAC parameters and thus incurs only a very low additional bit rate.As an example, while the first and second DirAC descriptions have DoA data and broadcast data for each spectral band and for each frame, the additional audio object metadata only requires a single DoA data for all frequency bands and this data only for every second frame or, preferably, every third, fourth, fifth or even every tenth frame in the preferred embodiment. Furthermore, with regard to the directional filtering carried out in the DirAC 220 synthesizer, which is typically included within a decoder on one side of the decoder of an encoder / decoder system, the DirAC synthesizer can, in the alternative shown in Fig. 2b, carry out the directional filtering within the parameter domain before scene combining or again carry out the directional filtering after 41 1724882 of 109 the combination of scenes. However, in this case, directional filtering is applied to the combined scene rather than the individual descriptions. Furthermore, if an audio object is not included in the first or second description but is included by its own audio object metadata, the directional filtering, as illustrated by the selective manipulator, can be selectively applied only to that additional audio object, for which additional audio object metadata exists, without affecting the first or second DirAC description or the combined DirAC description. For the audio object itself, there is also no separate transport channel representing the object's waveform signal, or the object's waveform signal is included in the downmix transport channel. Selective manipulation, as illustrated in Fig. 2b, for example, can proceed such that a certain arrival address is given by the address of the audio object introduced in Fig. 2d, included in the bitstream as side information or received from a user interface. Then, based on the address provided by the user or control information, the user can, for example, specify that the audio data from a certain address is to be enhanced or attenuated. Therefore, the object (metadata) for the object in question is amplified or attenuated. In the case of actual waveform data, such as object data fed into selective manipulator 226 from the left in Fig. 2d, the audio data would indeed be attenuated or enhanced depending on the control information. However, in the case of object data that have, in addition to the arrival direction and 1724882 of 109 optionally diffusivity or distance, additional energy information, then the energy information for the object would be reduced in the case of a required attenuation for the object or the energy information would be increased in the case of a necessary amplification of the object data. Therefore, directional filtering is based on a short-term spectral attenuation technique, and the spectral domain is filtered by a zero-phase gain function that depends on the direction of the objects. The direction can be contained within the bit stream if the object directions are transmitted as side information. Otherwise, the direction can also be interactively defined by the user.Naturally, the same procedure can not only be applied to the individual object given and reflected by the additional audio object metadata typically provided by DoA data for all frequency bands and DoA data with a low update ratio with respect to the image frequency and also proposed by the energy information for the object, but directional filtering can also be applied to the first DirAC description independent of the second DirAC description or vice versa or it can also be applied to the combined DirAC description, according to the case. Furthermore, it should be noted that the feature with respect to additional audio object data can also be applied to the first aspect of the present invention illustrated with respect to Figs. 1a to 1f. The input interface 100 of Fig. 1a further receives the additional audio object data as discussed with respect to Fig. 2a, and the format combiner can be implemented 43 1724882 of 109 as the DirAC synthesizer in the spectral domain 220 controlled by a user interface 260. Furthermore, the second aspect of the present invention, as illustrated in Fig. 2, differs from the first aspect in that the input interface already receives two DirAC descriptions, i.e., descriptions of a sound field that are in the same format, and therefore, for the second aspect, the format converter 120 of the first aspect is not necessarily required. On the other hand, when the input to the format combiner 140 of Fig. 1a consists of two DirAC descriptions, then the format combiner 140 can be implemented in accordance with what was discussed with respect to the second aspect illustrated in Fig. 2a, or, alternatively, the devices 220, 240 of Fig. 2a, can be implemented in accordance with what was discussed with respect to the format combiner 140 of Fig. 1a of the first aspect. Figure 3a illustrates an audio data converter comprising an input interface 100 for receiving an object description of an audio object containing audio object metadata. The input interface 100 is followed by a metadata converter 150, which also corresponds to the metadata converters 125 and 126 discussed in relation to the first aspect of the present invention, for converting the audio object metadata into DirAC metadata. The output of the audio converter in Figure 3a is an output interface 300 for transmitting or storing the DirAC metadata. The input interface 100 can also receive a waveform signal, as illustrated by the second arrow input on the interface 100. Furthermore, the output interface 300 can be implemented 44 1724882 of 109 to introduce, typically, a coded representation of the waveform signal at the output signal output by block 300. If the audio data converter is set to convert only a description of a single object, including metadata, output interface 300 also provides a DirAC description of this single audio object along with the waveform signal, typically coded as the DirAC transport channel. Specifically, audio object metadata contains the object's position, and DirAC metadata contains an arrival direction relative to a reference position derived from the object's position. Specifically, metadata converters 150, 125, and 126 are configured to convert DirAC parameters derived from the object data format into pressure / velocity data. The metadata converter is then configured to apply DirAC analysis to this pressure / velocity data, as illustrated in the flowchart in Fig. 3c, which consists of blocks 302, 304, and 306. For this purpose, the DirAC parameters output by block 306 are of higher quality than the DirAC parameters derived from the object metadata obtained by block 302; that is, they are enhanced DirAC parameters.3b illustrates the conversion of a position for an object in the direction of arrival with respect to a reference position for the specific object. Figure 3f illustrates a schematic diagram to explain the functionality of the metadata converter 150. The metadata converter 150 receives the position of the object indicated by the vector P in a coordinate system. Furthermore, the reference position, to which the DirAC metadata must be related, is given by the vector R in the same coordinate system. Therefore, the direction 45 1724882 of 109 of the arrival vector DoA extends from the tip of vector R to the tip of vector B. Therefore, the actual vector DoA is obtained by subtracting the reference position vector R from the object position vector P. In order to obtain normalized DoA information indicated by the DoA vector, the difference vector is divided by the magnitude or duration of the DoA vector. Furthermore, if necessary and desired, the length of the DoA vector can also be included in the metadata generated by the metadata converter 150. Additionally, the object's distance from the reference point is also included in the metadata, allowing for selective manipulation of the object based on its distance from the reference position. Specifically, the extract address block 148 in Fig. 1f can also function as discussed in relation to Fig. 3f, although other alternatives for calculating the DoA information and, optionally, the distance information can also be applied. Moreover, as discussed previously in relation to Fig.3a, blocks 125 and 126 illustrated in Fig. 1c or 1d can operate in a similar manner to that described with respect to Fig. 3f. Furthermore, the device in Fig. 3a can be configured to receive a plurality of audio object descriptions, and the metadata converter is configured to convert each metadata description directly into a DirAC description, and then the metadata converter is configured to combine the individual DirAC metadata descriptions to obtain a combined DirAC description such as the DirAC metadata illustrated in Fig. 3a. In one embodiment, 46 1724882 of 109 The combination is carried out by means of the calculation 320 of a weighting factor for a first arrival direction by the use of a first energy and the calculation 322 of a weighting factor for a second arrival direction by the use of a second energy, where the arrival direction is processed by blocks 320, 332 in relation to the same time / frequency compartment. Then, in block 324, a weighted sum is also carried out in accordance with what was discussed with respect to point 144 in Fig. 1d. Thus, the procedure illustrated in Fig. 3a represents a form of implementation of the first alternative of Fig. 1d. However, with respect to the second alternative, the procedure would be to set all diffusivity to zero or a small value, and for a given time / frequency compartment, consider all the different directions of arrival values ​​for that time / frequency compartment. The longest direction of arrival value is then selected as the combined direction of arrival value for that time / frequency compartment. In other embodiments, the second-largest direction could also be selected, provided that the energy values ​​for these two directions of arrival values ​​are not significantly different. The direction of arrival value selected is the one whose energy is either the largest among the energies of the different contributions for that time / frequency compartment or the second or third highest. Therefore, the third aspect, as described in relation to Figs. 3a to 3f, differs from the first aspect in that the third aspect is also useful for converting a single object description into a 47 1724882 of 109 DirAC metadata. Alternatively, input interface 100 can receive several object descriptions that are in the same object / metadata format. Therefore, no format converter is required, as discussed regarding the first aspect in Fig. 1a. Therefore, the implementation of the Fig. 3a can be useful in the context of receiving two different object descriptions by using different object waveform signals and different object metadata as the first scene description and the second description as input to the format combiner 140, and the output of the metadata converter 150, 125, 126 or 148 can be a DirAC rendering with DirAC metadata and, therefore, the DirAC analyzer 180 of Fig. 1 is also not required. However, the other elements with respect to the transport channel generator 160 that correspond to the downstream mixer 163 of Fig. 3a can be used in the context of the third aspect, as well as the transport channel encoder 170, the metadata encoder 190 and, in this context, the output interface 300 of Fig. 3a corresponds to the output interface 200 of Fig. 1a.Therefore, all corresponding descriptions given with respect to the first aspect also apply to the third aspect. Figures 4a and 4b illustrate a fourth aspect of the present invention in the context of an apparatus for performing audio data synthesis. In particular, the apparatus has an input interface 100 for receiving a DirAC description of an audio scene containing DirAC metadata and, furthermore, for receiving an object signal containing object metadata. This audio scene encoder illustrated in Figure 4b further comprises the generator of 48 1724882 of 109 metadata 400 for generating a combined metadata description comprising DirAC metadata on one hand and object metadata on the other. DirAC metadata comprises the arrival direction of individual time / frequency tiles, and object metadata comprises a direction or, additionally, a distance or diffusivity of an individual object. In particular, input interface 100 is configured to additionally receive a transport signal associated with the DirAC description of the audio scene, as illustrated in Fig. 4b, and the input interface is further configured to receive an object waveform signal associated with the object signal. Therefore, the scene encoder further comprises a transport signal encoder for encoding the transport signal and the object waveform signal, and the transport encoder 170 may correspond to encoder 170 in Fig. 1a. In particular, the metadata generator 400, which generates the combined metadata, can be configured as described for the first, second, or third aspect. In a preferred embodiment, the metadata generator 400 is configured to generate a single bandwidth address per time for object metadata, i.e., during a certain time frame, and the metadata generator is configured to update this single bandwidth address less frequently than the DirAC metadata. The procedure described with respect to Fig. 4b allows for combined metadata that includes metadata for a complete DirAC description and also metadata for an additional audio object, but in 49 1724882 of 109 DirAC format in such a way that a very useful DirAC rendering can be carried out by, at the same time, selective directional filtering or modification can be carried out in accordance with what was discussed previously with respect to the second aspect. Thus, the fourth aspect of the present invention, and in particular the metadata generator 400, represents a specific format converter in which the common format is the DirAC format, and the input is a DirAC description for the first scene in the first format discussed with respect to Fig. 1a, and the second scene is either a single or a combination such as a SAOC object signal. Therefore, the output of the format converter 120 represents the output of the metadata generator 400, but, in contrast to an actual specific combination of the metadata by one of the two alternatives, for example, as discussed with respect to Fig. 1d, the object metadata is included in the output signal; that is, the combined metadata is separated from the metadata for the DirAC description to allow selective modification of the object data. Therefore, the “direction / distance / diffusivity” indicated in point 2 on the right side of Fig. 4a corresponds to the additional audio object metadata input at input interface 100 in Fig. 2a, but, in the embodiment of Fig. 4a, only for a single DirAC description. Thus, in a sense, Fig. 2a could be said to represent a decoder-side implementation of the encoder illustrated in Fig. 4a, 4b, with the condition that the decoder side of the device in Fig. 2a receives only a single DirAC description and the object metadata generated by metadata generator 400 within 1724882 of 109 of the same bitstream as the additional audio object metadata.” Therefore, a completely different modification of the additional object data can be carried out when the encoded transport signal has a separate representation of the object waveform signal, separate from the DirAC transport stream. And yet, since the transport encoder 170 downmixes both data—that is, the transport channel for the DirAC description and the waveform signal from the object—the separation will be less perfect. However, by using additional object energy information, even a separation of a combined downmix channel and selective modification of the object with respect to the DirAC description are possible. Figures 5a to 5d represent a fifth additional aspect of the invention in the context of an apparatus for performing audio data synthesis. To this end, an input interface 100 is provided for receiving a DirAC description of one or more audio objects and / or a DirAC description of a multi-channel signal and / or a DirAC description of a first-order Ambisonics signal and / or a higher-order Ambisonics signal, wherein the DirAC description comprises position information for the one or more objects or side information for the first-order Ambisonics signal or the higher-order Ambisonics signal or position information for the multi-channel signal as side information or from a user interface. In particular, a 500 manipulator is configured for manipulating the DirAC description of one or more audio objects, the DirAC description of the multi-channel signal, the DirAC description of first-order Ambisonics signals, or the DirAC description of 1724882 of 109 higher-order Ambisonics signals are used to obtain a manipulated DirAC description. To synthesize this manipulated DirAC description, a DirAC 220, 240 synthesizer is configured for the synthesis of this manipulated DirAC description to obtain synthesized audio data. In a preferred embodiment, the DirAC synthesizer 220, 240 comprises a DirAC renderer 222 as illustrated in Fig. 5b and a subsequently connected spectral time converter 240 that outputs the manipulated time-domain signal. In particular, the manipulator 500 is configured to perform a position-dependent weighting operation prior to DirAC rendering. In particular, when the DirAC synthesizer is configured to output a plurality of objects from a first-order Ambisonics signal, a higher-order Ambisonics signal, or a multi-channel signal, the DirAC synthesizer is configured to use a separate spectral time converter for each object or each component of the first- or higher-order Ambisonics signals, or for each channel of the multi-channel signal, as illustrated in Fig. 5D in blocks 506 and 508. As indicated in block 510 then the output of the corresponding separate conversions are added together on the condition that all signals are in a common format, i.e., in a compatible format. Therefore, in the case of the input interface 100 of Fig. 5a, after receiving more than one, i.e., two or three representations, each representation could be manipulated separately, as illustrated in block 502 in the parameter domain according to what was discussed earlier with regard to Figs. 2b or 2c, and then a synthesis 52 could be carried out 1724882 of 109 in accordance with block 504 for each manipulated description, and the synthesis could then be added in the time domain as discussed with respect to block 510 in Fig. 5d. Alternatively, the result of the individual DirAC synthesis procedures in the spectral domain could already be summed in the spectral domain, and then a single time-domain conversion could also be used. In particular, the manipulator 500 can be implemented as the manipulator discussed with respect to Fig. 2D or discussed with respect to any other aspect above. Therefore, the fifth aspect of the present invention provides a significant feature with respect to the fact that, when individual DirAC descriptions of very different sound signals are entered, and when a certain manipulation of the individual descriptions is carried out in accordance with what was discussed with respect to block 500 of Fig. 5a, where an input in the manipulator 500 can be a DirAC description of any format, which includes only a single format, while the second aspect was concentrated on the reception of at least two different DirAC descriptions or when the fourth aspect, for example, was related to the reception of a DirAC description on the one hand and a description of the object signal on the other hand. Subsequently, reference is made to Fig. 6. Fig. 6 illustrates another implementation for performing a different synthesis of the DirAC synthesizer. When, for example, a sound field analyzer generates, for each source signal, a separate mono signal S and an original arrival direction, and when, depending on the translation information, a new arrival direction is calculated, then the Ambisonics 430 signal generator of the 53 1724882 of 109 Fig. 6, for example, would be used to generate a sound field description for the sound source signal, i.e., the mono signal S, but for the new data arrival direction (DoA) consisting of a horizontal angle θ or an elevation angle θ and an azimuth angle φ. Then, a procedure carried out by the sound field calculator 420 of Fig.6 would be to generate, for example, a first-order Ambisonics sound field representation for each sound source with the new arrival direction, and then a further modification could be carried out per sound source by using a scaling factor that depends on the distance of the sound field to the new reference location, and then all the sound fields of the individual sources could be superimposed on each other to finally obtain the modified sound field, once again, for example, in an Ambisonics representation related to a certain new reference location. When each time / frequency compartment processed by the DirAC 422 analyzer is interpreted as representing a certain sound source (bandwidth limited), then the Ambisonics 430 signal generator could be used, instead of the DirAC 425 synthesizer, to generate, for each time / frequency compartment, a complete Ambisonics representation by using the downmix signal or the pressure or omnidirectional component signal for this time / frequency compartment as the “mono signal S” in Fig. 6. An individual time-frequency conversion in the 426 time-frequency converter for each of the W, X, Y, Z components would then result in a description of the sound field different from that illustrated in Fig. 6. Later, further explanations are given about a Figure 7a illustrates a DirAC analyzer as originally described, for example, in the 2009 IWPASH reference "Directional Audio Coding." The DirAC analyzer comprises a band filter bank (1310), an energy analyzer (1320), an intensity analyzer (1330), a time averaging block (1340), a diffusivity calculator (1350), and a direction calculator (1360). In DirAC, both analysis and synthesis are performed in the frequency domain. Several methods exist for dividing sound into frequency bands, each with distinct properties. The most commonly used frequency transforms include the short-time Fourier transform (STFT), and the Quadrature Mirror Filter (QMF) bank.In addition to these, there is complete freedom to design a filter bank with arbitrary filters optimized for specific purposes. The goal of directional analysis is to estimate the direction of arrival of the sound in each frequency band, along with an estimate of whether the sound is arriving from one or multiple directions simultaneously. In principle, this can be accomplished with a number of techniques; however, sound field energy analysis has been found to be suitable, as illustrated in Fig. 7a. Energy analysis can be performed when the pressure and velocity signals in one, two, or three dimensions are captured from a single position. In first-order B-format signals, the omnidirectional signal is called the W signal, which has been reduced by the square root of two. The sound pressure can be estimated as 5 = V² * IV, expressed in the STFT domain. The X, Y, and Z channels have the directional pattern of a 1724882 of 109 dipoles directed along the Cartesian axis, which together form a vector U = [X, Y, Z]. This vector estimates the sound field velocity vector and is also expressed in the STFT domain. The sound field energy E is calculated. Signal capture in B format can be achieved either with coincident positioning of directional microphones or with a closely spaced array of omnidirectional microphones. In some applications, microphone signals can be formed in a computational, i.e., simulated, domain. The sound direction is defined as the opposite direction of the intensity vector I. The direction is denoted as corresponding angular azimuth and elevation values ​​in the transmitted metadata. The sound field diffusivity is also calculated using an expectation operator of the intensity vector and energy.The result of this equation is a real number between zero and one, which indicates whether the sound energy is arriving from a single direction (diffusivity is zero) or from all directions (diffusivity is one). This procedure is appropriate when complete 3D or lower-dimensional velocity information is available. Figure 7b illustrates a DirAC synthesis, which again has a bandfilter bank 1370, a virtual microphone block 1400, a direct / diffuse synthesizer block 1450, and a certain speaker configuration or a planned virtual speaker configuration 1460. Additionally, a diffusivity gain transformer 1380, a vector-based amplitude panning gain table (VBAP) block 1390, a microphone compensation block 1420, a speaker gain averaging block 1430, and a splitter 1440 for other channels are used. In this DirAC synthesis with speakers, the high-quality version of the synthesis 1724882 of 109 The DirAC shown in Fig. 7b receives all signals in B format, for which a virtual microphone signal is calculated for each speaker direction in the 1460 speaker configuration. The directional pattern typically used is a dipole. The virtual microphone signals are then modified nonlinearly, depending on the metadata. The low-bitrate version of DirAC is not shown in Fig. 7b; however, in this situation, only one audio channel is transmitted, as illustrated in Fig. 6. The difference in processing is that all virtual microphone signals are replaced by the single received audio channel. The virtual microphone signals are split into two streams: diffuse and non-diffuse streams, which are processed separately. Non-diffuse sound is reproduced as point sources using vector-based amplitude panning (VBAP). In VBAP, a monophonic sound signal is applied to a subset of the speakers after multiplication by speaker-specific gain factors. These gain factors are calculated using information about a speaker configuration and the specified panning direction. In the low-bitrate version, the input signal is simply panned to the directions implied by the metadata. In the high-quality version, each virtual microphone signal is multiplied by its corresponding gain factor, producing the same effect as panning but being less prone to nonlinear artifacts. In many cases, address metadata is subject to abrupt temporal changes. To avoid artifacts, the gain factors for loudspeakers calculated with VBAP are smoothed by time integration with time constants dependent on the 57 1724882 of 109 frequencies, which equates to approximately 50 cycle periods in each band. This effectively eliminates artifacts; however, changes in direction are not perceived as slower than without averaging in most cases. The goal of diffuse sound synthesis is to create the perception of sound surrounding the listener. In the low-bitrate version, diffuse stream is reproduced by decorrelating the input signal and playback from each speaker. In the high-quality version, the signals from the diffuse stream virtual microphones are already somewhat incoherent and only need to be slightly correlated. This approach provides better spatial quality of surround reverb and ambient sound than the low-bitrate version.For DirAC synthesis with headphones, DirAC is formulated with a certain number of virtual loudspeakers around the listener for the non-diffuse current and a certain number of loudspeakers for the diffuse current. The virtual loudspeakers are implemented as a convolution of the input signals with head-related transfer functions (HRTFs). Subsequently, a general relationship is given in addition with respect to the different aspects and, in particular, with respect to other implementations of the first aspect in accordance with what is discussed with respect to Fig. 1a. In general, the present invention relates to the combination of different scenes in different formats by the use of a common format, where the common format may, for example, be the B-format domain, the pressure / speed domain, or the metadata domain in accordance with what is discussed, for example, in points 120, 140 of Fig. 1a. When the combination is not carried out directly 1724882 of 109 in the common DirAC format, then a DirAC 802 analysis is carried out on one of the alternatives before transmission in the encoder according to what was discussed previously with respect to point 180 of Fig. 1a. Following the DirAC analysis, the result is encoded as previously discussed with respect to encoder 170 and metadata encoder 190, and the encoded result is transmitted via the encoded output signal generated by output interface 200. Alternatively, the result could be directly rendered by a device of Fig. 1a when the output of block 160 and the output of block 180 of Fig. 1a are forwarded to a DirAC renderer. In this case, the device of Fig. 1a would not be a dedicated encoder device, but rather a corresponding analyzer and processor. An additional alternative is illustrated in the right-hand branch of Fig. 8, where transmission from the encoder to the decoder takes place. As illustrated in block 804, DirAC analysis and DirAC synthesis are performed after transmission, i.e., on the decoder side. This procedure would be the case when using the alternative in Fig. 1a, i.e., when the encoded output signal is a B-format signal without spatial metadata. After block 808, the result could be rendered for playback or, alternatively, it could even be encoded and transmitted again. It is thus evident that the procedures of the invention, as defined and described with respect to the various aspects, are highly flexible and can be readily adapted to specific use cases. 1724882 of 109 First Appearance of the Invention: Universal DirAC-based spatial audio encoding / rendering A DirAC-based spatial audio encoder that can encode multi-channel signals, Ambisonics formats, and audio objects separately or simultaneously. Benefits and Advantages over the State of the Art • Universal DirAC-based spatial audio coding scheme for the most relevant immersive audio input formats • Universal audio rendering of different input formats to different output formats Second Aspect of the Invention: Combination of two or more DirAC descriptions in a decoder The second aspect of the invention relates to the combination and rendering of two or more DirAC descriptions in the spectral domain. Benefits and Advantages over the State of the Art • Efficient and precise combination of DirAC streams • Enables the use of DirAC that universally represents any scene and efficiently combines different streams in the parameter domain or the spectral domain • Efficient and intuitive scene manipulation of individual DirAC scenes or the combined scene in the spectral domain and subsequent conversion in the time domain of the manipulated combined scene. Third Aspect of the Invention: Conversion of objects 1724882 of 109 audio in the DirAC domain The third aspect of the invention relates to the conversion of object metadata and optionally object waveform signals directly into the DirAC domain and, in one embodiment, the combination of several objects into one object representation. Benefits and Advantages over the State of the Art • Efficient and accurate DirAC metadata estimation by means of a simple metadata transcoder of audio object metadata • Enables DirAC to encode complex audio scenes that include one or more audio objects • Efficient method for encoding audio objects through DirAC into a single parametric representation of the complete audio scene. Fourth Aspect of the Invention: Combination of Object Metadata and Regular DirAC Metadata The third aspect of the invention relates to amending the DirAC metadata with the directions and, ideally, the distance or diffusivity of the individual objects comprising the combined audio scene represented by the DirAC parameters. This additional information is easily encoded, since it consists primarily of a single broadband address per unit of time and can be updated less frequently than the other DirAC parameters because the objects can be assumed to be static or moving at a slow rate. Benefits and Advantages over the State of the Art • DirAC allows encoding a complex audio scene involving one or more audio objects 1724882 of 109 • An efficient and accurate DirAC metadata estimate by means of simple metadata transcoding of audio object metadata. • More efficient method for encoding audio objects through DirAC by efficiently combining their metadata in the DirAC domain. • Efficient method for encoding audio objects through DirAC by efficiently combining their audio representations into a single parametric representation of the audio scene. Fifth Aspect of the Invention: Manipulation of MC and FOA / HOA C Object Scenes in DirAC Synthesis The fourth aspect relates to the decoder side and exploits the known positions of audio objects. These positions can be provided by the user through an interactive interface and can also be included as additional information within the bitstream. The goal is to be able to manipulate an output audio scene comprising a number of objects by individually changing object attributes such as levels, equalization, and / or spatial positions. It may also be possible to filter the entire object or restore individual objects from the combined stream. Manipulation of the output audio scene can be achieved through joint processing of the spatial parameters of the DirAC metadata, object metadata, interactive user input if present, and audio signals carried on the transport channels. 1724882 of 109 Benefits and Advantages over the State of the Art • Allows DirAC to output audio objects from the decoder side according to what is presented at the encoder input. • Enables DirAC playback to manipulate individual audio objects by applying gain, rotation, or... • The capability requires minimal additional computational effort since it only requires a position-dependent weighting operation before rendering and a synthesis filter bank at the end of DirAC synthesis (additional object outputs will only require one additional synthesis filter bank per object output). References that are incorporated in their entirety as references in this report: [1] V. Pulkki, MV Laitinen, J. Vilkamo, J. Ahonen, T. Lokki and T. Pihlajamaki, “Directional audio coding perception-based reproduction of spatial sound”, International Workshop on the Principles and Application on Spatial Hearing, Nov. 2009, Zao; Miyagi, Japan. [2] Ville Pulkki. “Virtual source positioning using vector base amplitude panning”. J. Audio Eng. Soc., 45(6): 456 a 466, junio de 1997. [3] M. V. Laitinen and V. Pulkki, Converting 5.1 audio recordings to B-format for directional audio coding reproduction, 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)), Praga, 2011, págs. 61 a 64. [4] G. Del Galdo, F. Kuech, M. Kallinger and R. Schultz-Amling, Efficient merging of multiple audio streams for spatial sound reproduction in Directional Audio Coding, 2009 IEEE International Conference on Acoustics, 1724882 de 109 Speech and Signal Processing, Taipei, 2009 págs. 265 a 268. [5] Jürgen HERRE, CORNELIA FALCH, DIRK MAHNE, GIOVANNI DEL GALDO, MARKUS KALLINGER, AND OLIVER THIERGART, “Interactive Teleconferencing Combining Spatial Audio Object Coding and DirAC Technology, J. Audio Eng. Soc., Vol. 59, Núm. 12, diciembre de 2011. [6] R. Schultz-Amling, F. Kuech, M. Kallinger, G. Del Galdo, J. Ahonen, V. Pulkki, “Planar Microphone Array Processing for the Analysis and Reproduction of Spatial Audio using Directional Audio Coding, Audio Engineering Society Convention 124, Amsterdam, Netherlands, 2008. [7] Daniel P. Jarrett and Oliver Thiergart and Emanuel AP Habets and Patrick A. Naylor, “Coherence-Based Diffuseness Estimation in the Spherical Harmonic Domain, IEEE 27th Convention of Electrical and Electronics Engineers in Israel (IEEEI), 2012. [8] United States Patent 9,015,051. The present invention provides, in additional embodiments, and in particular with respect to the first aspect and also with respect to the other aspects, different alternatives. These alternatives are as follows: First, the combination of different formats in the format B domain and either the performance of DirAC analysis in the encoder or the transmission of the combined channels to a decoder and the performance of DirAC analysis and synthesis there. Secondly, the combination of different formats in the pressure / velocity domain and the performance of DirAC analysis in the encoder. Alternatively, the pressure / velocity data is transmitted to the decoder, and the DirAC analysis and synthesis are performed in the decoder. 1724882 of 109 Third, the combination of different formats in the metadata domain and the transmission of a single DirAC stream or the transmission of several DirAC streams to a decoder before combining them and making the combination in the decoder. Furthermore, the embodiments or aspects of the present invention are related to the following aspects: First, the combination of different audio formats according to the three alternatives above. Secondly, a reception, combination and rendering of two DirAC descriptions already in the same format is carried out. Third, a specific object is implemented in the DirAC converter with a direct conversion of object data to DirAC data. Fourth, object metadata in addition to normal DirAC metadata and a combination of both metadata; also data that exists side-by-side in the bitstream, but also audio objects are described by the DirAC metadata style. Fifth, the objects and the DirAC current are transmitted separately to a decoder and the objects are selectively manipulated within the decoder before converting the output audio signals (speaker) in the time domain. It should be mentioned here that all alternatives or aspects as discussed above and all aspects as defined by means of the independent claims in the following claims may be used individually, i.e., without any other alternative or object than the alternative, object or independent claim 1724882 of 109 contemplated. However, in other embodiments, two or more of the alternatives or aspects or independent claims may be combined with each other and, in other embodiments, all aspects, or alternatives and all independent claims may be combined with each other. An audio signal encoded according to the invention can be stored on a digital storage medium or a non-transient storage medium or can be transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium, such as the Internet. While some aspects have been described in the context of an apparatus, it is evident that these aspects also represent a description of the corresponding method, where a block or device corresponds to a step of the method or a feature of a method step. Similarly, the aspects described in the context of a method step also represent a description of a corresponding block or an element or feature of a corresponding apparatus. Depending on certain implementation requirements, the embodiments of the invention can be implemented in hardware or software. Implementation can be carried out using a digital storage medium, for example, a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, which has electronically readable control signals stored therein, and which cooperates (or is capable of cooperating) with a programmable computer system in such a way as to carry out the respective method. Some embodiments according to the invention comprise a data carrier with signals of 1724882 of 109 electronically readable controls, which are capable of cooperating with a programmable computer system, in such a way that one of the methods described in this memory is carried out. Generally, the embodiments of the present invention can be implemented as a computer program product with program code. The program code is operational for carrying out one of the methods when the computer program product is executed on a computer. The program code may be stored on a machine-readable medium, for example. Other embodiments include the computer program to carry out one of the methods described in this document, stored on a machine-readable medium or a non-transient storage medium. In other words, one embodiment of the method according to the invention is, therefore, a computer program having program code for carrying out one of the methods described herein, when the computer program is executed on a computer. A further embodiment of the methods of the invention is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded therein, the computer program for carrying out one of the methods described herein. A further embodiment of the method according to the invention is, therefore, a data stream or a sequence of signals representing the computer program for carrying out one of the methods described herein. The stream of 1724882 of 109 data or signal sequences can, for example, be configured to be transferred through a data communication connection, for example, via the Internet. A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured or adapted to carry out one of the methods described herein. One embodiment further comprises a computer having installed therein the computer program to carry out one of the methods described in this document. In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably implemented by any hardware device. The embodiments described above are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be evident to those skilled in the art. Therefore, the intention is to be limited only by the scope of the imminent patent claims and not by the specific details presented herein as a description and explanation of the embodiments.

Claims

1. An audio data converter, comprising: an input interface (100) for receiving a plurality of audio object descriptions of a plurality of audio objects, an audio object description having audio object metadata, wherein the audio object metadata having an audio object position in space; characterized in that it comprises a metadata converter (150, 125, 126, 148) for converting the audio object metadata into DirAC metadata, wherein the DirAC metadata having an arrival direction with respect to a reference position, wherein the metadata converter (150, 125, 126, 148) is configured to convert each audio object metadata of the plurality of audio object descriptions into an individual DirAC data description comprising individual DirAC metadata for each audio object description,and to combine the individual DirAC metadata for audio object descriptions to obtain combined DirAC metadata as DirAC metadata, or to convert DirAC parameters in the individual DirAC metadata derived from audio object metadata for each audio object description of the plurality of audio object descriptions into individual pressure vectors and individual velocity vectors, to sum the individual pressure vectors derived from each audio object description of the plurality of audio object descriptions to obtain a combined pressure vector, to sum the individual velocity vectors derived from each audio object description of the plurality of audio object descriptions to obtain a combined velocity vector,and applying a DirAC analysis to the combined pressure vector and the combined velocity vector to obtain the DirAC metadata; and an output interface (300) for transmitting or storing the DirAC metadata. Eleven claims follow,