Apparatus and method for object-based spatial audio mastering
By grouping audio objects into processing objects for simultaneous adjustment, the apparatus and method address the inefficiencies of traditional object-based spatial audio mastering, enabling efficient and creative 3D audio content production.
Patent Information
- Application Number
- JP2025031874
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-04-19
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-01
AI Technical Summary
The existing mastering processes for object-based spatial audio are inefficient and require manual, individual adjustments to each audio object, which is time-consuming and limited by the acoustic characteristics of the monitoring environment, restricting the quality and creativity of 3D audio content production.
An apparatus and method that groups audio objects into processing objects, allowing simultaneous adjustment of multiple audio objects using effect parameters specified by an interface, with a processor unit applying these parameters to the audio object signals or metadata, enabling centralized mastering on an encoder or decoder side.
Facilitates efficient, real-time adjustment of audio objects, reducing the effort required for mastering and enhancing the quality and creativity of 3D audio content production by allowing simultaneous processing of multiple audio objects.
Smart Images

Figure 2025098034000001_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the processing, encoding, and decoding of audio objects, and in particular to audio mastering of audio objects.
Background Art
[0002] Object-based spatial audio is an approach to interactive three-dimensional audio playback. This concept changes not only the way content creators or authors interact with audio, but also the way audio is stored and transmitted. To enable this, a new process needs to be established in the playback chain called "rendering". The rendering process generates loudspeaker signals from the description of an object-based scene. In recent years, research has been conducted on recording and mixing, but there is little concept of object-based mastering. The main difference from channel-based audio mastering is that instead of adjusting audio channels, it is necessary to modify audio objects. This requires a fundamentally new approach to mastering. In this document, a new method for mastering object-based audio is provided.
[0003] In recent years, the object-based audio approach has attracted much attention. As a result of spatial audio production, the audio scene is described by audio objects as compared to channel-based audio where loudspeaker signals are stored. An audio object can be considered as a virtual sound source consisting of an audio signal with additional metadata such as position and gain. To play an audio object, a so-called audio renderer is required. Audio rendering is a process of generating signals for loudspeakers or headphones based on additional information such as the position of loudspeakers in a virtual scene or the position of a listener.
[0004] The process of creating audio content can be divided into three main parts: recording, mixing, and mastering. Over the past few decades, all three steps have been extensively covered for channel-based audio, but object-based audio will require new workflows in future applications. Generally, even though future technologies may bring new possibilities, there is still no need to change the recording step [1], [2]. In the case of the mixing process, the situation is somewhat different because sound engineers no longer create a spatial mix by panning signals to dedicated speakers. Instead, the positions of all audio objects are generated by a spatial authoring tool, which allows the metadata part of each audio object to be defined. The complete mastering process for audio objects has not yet been established [3].
[0005] In traditional audio mixing, multiple audio tracks are routed to a specific number of output channels. For this reason, it is necessary to create individual mixes for different playback configurations, but the output channels can be efficiently processed during mastering [4]. When using an object-based audio approach, the audio renderer is responsible for creating all speaker signals in real time. By placing a large number of audio objects within the framework of an innovative mixing process, a complex audio scene is generated. However, since the renderer can play the audio scene on several different loudspeaker means, it is not possible to directly process the output channels during production. The mastering concept may, therefore, be based only on individually modifying the audio objects.
[0006] Up to now, traditional audio production has been directed towards very special auditory equipment and their channel configurations, such as for stereo or surround sound reproduction. Therefore, the decision of the playback device on which the content is to be configured needs to be made at the start of production. The production process itself consists of recording, mixing, and mastering. The mastering process optimizes the final mix to ensure that the mix is reproduced with satisfactory quality on all consumer systems with different speaker characteristics. Since the desired output format of the mix is fixed, the mastering engineer (ME) can create a master optimized for this playback configuration.
[0007] Since one can rely on the final check of those mixes during mastering, it is advisable for the creator to produce the audio in a second-best acoustic environment at the mastering stage. This reduces the barriers to entry for the production of specialized content. On the other hand, due to the fact that the MEs themselves have provided a wide range of mastering tools over the years, the ability to make corrections and extensions has been dramatically improved. Nevertheless, the final content is usually limited to the playback means on which it is configured.
[0008] This limitation is generally overcome by object-based spatial audio production (OBAP). In contrast to channel-based audio, OBAP is based on individual audio objects with metadata including their positions in an artificial environment, also called a "scene". Only at the final listening output does a dedicated rendering unit, the renderer, calculate the final speaker signals in real time based on the listener's speaker means.
[0009] OBAP provides each audio object and its metadata to the renderer individually, but cannot be directly adjusted on a channel basis during production. Therefore, existing mastering tools for conventional playback equipment cannot be used. On the other hand, OBAP needs to perform all final adjustments in the mix. Implementing overall sound adjustment by manually processing each of the individual audio objects is not only extremely inefficient, but this situation also places high demands on the monitoring equipment of each creator, and the sound quality of object-based 3D audio content is strictly limited by the acoustic characteristics of the created environment.
[0010] Ultimately, by developing tools on the creator side that can similarly enable powerful OBAP mastering processing, the barriers to production are reduced, and a new space for the aesthetics and quality of sound is opened, thereby improving the acceptance of 3D audio content production.
[0011] The first ideas regarding spatial mastering are generally publicly available [5], and this document provides a new approach on how to adapt conventional mastering tools and what kinds of new tools are considered useful for mastering object-based spatial audio. Therefore, [5] describes the basic sequence of the method of using metadata to derive object-specific parameters from global properties to objects. Furthermore, [6] describes the concept of regions of interest with surrounding transition regions in the context of OBAP applications.
Summary of the Invention
Problems to be Solved by the Invention
[0012] Therefore, it is desired to provide an improved concept of object-based audio mastering.
Means for Solving the Problems
[0013] An apparatus according to claim 1, an encoder according to claim 14, a decoder according to claim 15, a system according to claim 17, a method according to claim 18, and a computer program according to claim 19 are provided.
[0014] In one embodiment, an apparatus for generating a processed signal while using a plurality of audio objects is provided. Each audio object of the plurality of audio objects includes an audio object signal and audio object metadata. The audio object metadata includes the position of the audio object and the gain parameter of the audio object. The apparatus includes an interface for a user to specify at least one effect parameter of at least one processing object group of the audio objects. The processing object group of the audio objects includes two or more audio objects among the plurality of audio objects. The apparatus further includes a processor unit, and the apparatus is configured to generate a processed signal such that at least one effect parameter specified by the interface is applied to the audio object signal or the audio object metadata of each audio object of the processing object group of the audio objects. One or more audio objects among the plurality of audio objects do not belong to the processing object group of the audio objects.
[0015] The method further includes generating a processed signal while using a plurality of audio objects. Each audio object of the plurality of audio objects includes an audio object signal and audio object metadata. The audio object metadata includes the position of the audio object and the gain parameter of the audio object. The method - A step of specifying, by an interface (110), at least one effect parameter of a processing object group of audio objects on the user side, wherein the processing object group of audio objects comprises two or more audio objects among a plurality of audio objects; - A step of generating a processed signal by a processor unit (120) such that at least one effect parameter specified by the interface is applied to an audio object signal or audio object metadata of each audio object of the processing object group of audio objects; including.
[0016] Furthermore, a computer program including program code for executing the method described above is provided.
[0017] The provided audio mastering is based on the mastering of audio objects. In embodiments, these can be freely arranged in real time at any position in the scene. In embodiments, for example, the characteristics of general audio objects are affected. In their function as artificial containers, any number of audio objects can be included respectively. Each adjustment to the mastering object is converted into an individual adjustment to the same audio object in real time.
[0018] Such a mastering object is also called a processing object.
[0019] Therefore, instead of adjusting a number of audio objects individually, the user can use the mastering object to perform mutual adjustment on a plurality of audio objects simultaneously.
[0020] For example, according to an embodiment, a set of target audio objects of a mastering object can be defined in many ways. From a spatial perspective, the user can define a customized effective range around the position of the mastering object. Alternatively, regardless of the position, individually selected audio objects can be linked to the mastering object. The mastering object also takes into account potential changes in the position of the audio object over time.
[0021] A second characteristic of the mastering object according to an embodiment may be the ability to calculate how each audio object is individually affected, for example, based on an interaction model. Similar to a channel strip, the mastering object can inherit common mastering effects such as an equalizer or a compressor. An effect plugin usually provides the user with a number of parameters, for example, as frequency or gain controls. When a new mastering effect is added to the mastering object, all audio objects in the target set of the mastering object are automatically copied. However, not all effect parameter values are transferred without change. Depending on how the target set is calculated, some mastering effect parameters can be weighted before being applied to a specific audio object. The weights can be based on any metadata or the sound characteristics of the audio object.
[0022] Hereinafter, preferred embodiments of the present invention will be described with reference to the drawings.
Brief Description of the Drawings
[0023]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
[0024] Figure 1 shows an apparatus for generating a processed signal while using a plurality of audio objects. Each audio object of the plurality of audio objects includes an audio object signal and audio object metadata, and the audio object metadata includes the position of the audio object and the gain parameter of the audio object.
[0025] The apparatus includes an interface 110 for specifying, on the user side, at least one effect parameter of at least one processing object group of audio objects, and the processing object group of audio objects includes two or more audio objects among the plurality of audio objects.
[0026] The apparatus further includes a processor unit 120 configured to generate a processed signal such that at least one effect parameter specified by the interface 110 is applied to the audio object signal or the audio object metadata of each audio object of the processing object group of audio objects.
[0027] One or more audio objects among the plurality of audio objects do not belong to the processing object group of audio objects.
[0028] The apparatus described in FIG. 1 above implements an efficient form of audio mastering for audio objects.
[0029] In the case of audio objects, there is a problem that there are a large number of audio objects in an audio scene. When changing these, it takes a considerable amount of effort to specify each audio object individually.
[0030] According to the present invention, a group of two or more audio objects is immediately organized into a group of audio objects called a processing object group. Therefore, the processing object group is a group of audio objects organized into this special group, the processing object group.
[0031] According to the present invention, the user may immediately specify one or more (at least one) effect parameters by means of the interface 110. The processor unit 120 guarantees, by means of a single input of the effect parameter, that the effect parameter is applied to all of the two or more audio objects of the processing object group.
[0032] Such an application of the effect parameter may, for example, consist of the effect parameter modifying a specific frequency range of the audio object signals of each audio object of the processing object group.
[0033] Or, for example, the gain parameter of the audio object metadata of each audio object of the processing object group may increase or decrease depending on the effect parameter.
[0034] Or, for example, depending on the effect parameter, the position of the audio object metadata of each audio object of the processing object group may be changed. For example, it is conceivable that all audio objects of the processing object group shift by +2 along the x-axis, -3 along the y-axis, and +4 along the z-axis.
[0035] It is also conceivable that applying the effect parameters to the audio objects of the processing object group has different effects on each audio object of the processing object group. For example, the axis around which the positions of all audio objects in the processing object group are mirrored can be defined as an effect parameter. Therefore, changing the positions of the audio objects in the processing object group will have different effects on each audio object in the processing object group.
[0036] For example, in one embodiment, the processor unit 120 may be configured not to apply at least one effect parameter specified, for example, by an interface, to any audio object signal and any audio object metadata of one or more audio objects that do not belong to the processing object group of the audio object.
[0037] In the case of such an embodiment, it is specified that the effect parameter is not applied exactly to the audio objects that do not belong to the processing object group.
[0038] In principle, the mastering of audio objects may be performed centrally on the encoder side. Or, on the decoder side, the end user as the recipient of the background of the audio object can modify the audio object by himself according to the present invention.
[0039] An embodiment of implementing the mastering of the audio object according to the present invention on the encoder side is shown in FIG. 2.
[0040] An embodiment of implementing the mastering of the audio object according to the present invention on the decoder side is shown in FIG. 3.
[0041] FIG. 2 shows an apparatus according to another embodiment, and the apparatus is an encoder.
[0042] In FIG. 2, the processor unit 120 is configured to generate a downmix signal while using the audio object signals of a plurality of audio objects. In this context, the processor unit 120 is configured to generate a metadata signal while using the audio object metadata of a plurality of audio objects.
[0043] Furthermore, the processor unit 120 in FIG. 2 is configured to generate a downmix signal as a processed signal, and for each audio object in the processing object group of the audio object, at least one modified object signal is mixed into the downmix signal, and the processor unit 120 is configured to generate a modified object signal for each audio object in the processing object group of the audio object by applying at least one effect parameter specified by the interface 110 to the audio object signal of the audio object.
[0044] Alternatively, the processor unit 120 in FIG. 2 is configured to generate a metadata signal as a processed signal, the metadata signal includes at least one modified position for each audio object in the processing object group of the audio object, and the processor unit 120 is configured to generate a modified position for each audio object in the processing object group of the audio object by applying at least one effect parameter specified by the interface 110 to the position of the audio object.
[0045] Alternatively, the processor unit 120 in FIG. 2 is configured to generate a metadata signal as a processed signal, the metadata signal includes at least one correction gain parameter for each audio object of a processing object group of audio objects, and the processor unit 120 is configured to generate a correction gain parameter for each audio object of the processing object group of audio objects by applying at least one effect parameter specified by the interface 110 to the gain parameter of the audio object.
[0046] FIG. 3 shows an apparatus according to another embodiment, and the apparatus is a decoder. The apparatus in FIG. 3 is configured to receive a downmix signal in which a plurality of audio object signals of a plurality of audio objects are mixed. Further, the apparatus in FIG. 3 is configured to receive a metadata signal, and the metadata signal includes audio object metadata of each audio object of the plurality of audio objects.
[0047] The processor unit 120 in FIG. 3 is configured to reconstruct a plurality of audio object signals of a plurality of audio objects based on the downmix signal.
[0048] Furthermore, the processor unit 120 in FIG. 3 is configured to generate an audio output signal having one or more audio output channels as a processed signal.
[0049] Furthermore, the processor unit 120 in FIG. 3 is configured to apply at least one effect parameter specified by the interface 110 to the audio object signals of each audio object in the processing object group of the audio object, to generate a processed signal, or to apply at least one effect parameter specified by the interface 110 to the position or gain parameter of the audio object metadata of each audio object in the processing object group of the audio object, to generate a processed signal.
[0050] In the decoding of an audio object, rendering on the decoder side is well known to those skilled in the art, for example, from the SAOC standard (Spatial Audio Object Coding). See [8].
[0051] On the decoder side, one or more rendering parameters can be specified by user input via the interface 110.
[0052] For example, in one embodiment, the interface 110 in FIG. 3 may be configured to specify one or more rendering parameters on the user side. For example, the processor unit 120 in FIG. 3 may be configured to generate a processed signal while using one or more rendering parameters depending on the position of each audio object in the processing object group of the audio object.
[0053] FIG. 4 shows a system according to one embodiment including an encoder 200 and a decoder 300.
[0054] The encoder 200 in FIG. 4 is configured to generate a downmix signal based on the audio object signals of a plurality of audio objects and generate a metadata signal based on the audio object metadata of the plurality of audio objects, and the audio object metadata includes the position of the audio object and the gain parameter of the audio object.
[0055] The decoder 400 in FIG. 4 is configured to generate an audio output signal including one or more audio output channels based on the downmix signal and based on the metadata signal.
[0056] The encoder 200 of the system in FIG. 4 may be the device according to FIG. 2.
[0057] Alternatively, the decoder 300 of the system in FIG. 4 may be the device according to FIG. 3.
[0058] Alternatively, the encoder 200 of the system in FIG. 4 may be the device according to FIG. 2, and the decoder 300 of the system in FIG. 4 may be the device according to FIG. 3.
[0059] The following embodiments may be equally implemented in the device of FIG. 1, the device of FIG. 2, and the device of FIG. 3. Also, they may be implemented in the encoder 200 of the system in FIG. 4 and the decoder 300 of the system in FIG. 4.
[0060] According to one embodiment, the processor unit 120 may be configured to generate a processed signal such that at least one effect parameter specified by the interface 110, for example, is applied to the audio object signal of each audio object in the processing object group of audio objects. In this context, the processor unit 120 may be configured to not apply at least one effect parameter specified by the interface, for example, to the audio object signal of any one or more of the plurality of audio objects that do not belong to the processing object group of audio objects.
[0061] The application of such an effect parameter may be configured such that, for example, the application of the effect parameter to the audio object signal of each audio object in the processing object group changes a specific frequency range of the audio object signal of each audio object in the processing object group.
[0062] In one embodiment, the processor unit 120 may be configured to generate a processed signal such that at least one effect parameter specified by the interface 110, for example, is applied to the gain parameter of the metadata of each audio object in the processing object group of audio objects. In this context, the processor unit 120 may be configured to not apply at least one effect parameter specified by the interface, for example, to the gain parameter of the audio object metadata of any one or more of the plurality of audio objects that do not belong to the processing object group of audio objects.
[0063] As described above, in such an embodiment, the gain parameter of the audio object metadata of each audio object in the processing object group may be increased (e.g., +3 dB) or decreased as a function of the effect parameter.
[0064] According to one embodiment, the processor unit 120 may be configured to generate a processed signal such that at least one effect parameter specified by, for example, the interface 110 is applied to the position of the metadata of each audio object in the processing object group of the audio object. In this context, the processor unit 120 may be configured not to apply at least one effect parameter specified by the interface to any position of the audio object metadata of one or more audio objects that do not belong to the processing object group of the audio object.
[0065] As already described, in such an embodiment, the position of the audio object metadata of each audio object in the processing object group may be changed accordingly, for example, as a function of the effect parameter. This may be performed, for example, by specifying corresponding x, y, z coordinate values at which the position of each audio object is shifted. Or, for example, a shift (displacement) of rotating about a specific angle, for example, a center point defined around the position of the user may be specified. Or, for example, doubling (or halving) the distance from a specific point may be provided as an effect parameter for the position of each audio object in the processing object group.
[0066] In one embodiment, the interface 110 may be configured to specify, for example, at least one definition parameter of a processing object group of audio objects on the user side. In this context, the processor unit 120 may be configured to determine, for example, to which processing object group of audio objects a plurality of audio objects belong according to at least one definition parameter of at least one processing object group of audio objects specified by the interface 110.
[0067] For example, according to one embodiment, at least one definition parameter of a processing object group of audio objects may include at least one position of a region of interest (the position of the region of interest is, for example, the center or centroid of the region of interest). The region of interest may be associated with a processing object group of audio objects. The processor unit 120 may be configured to determine, for example, for each audio object among the plurality of audio objects, whether this audio object belongs to the processing object group of audio objects depending on the position of the audio object metadata of this audio object and depending on the position of the region of interest.
[0068] In one embodiment, at least one definition parameter of a processing object group of audio objects may further include, for example, a radius range of a region of interest associated with the processing object group of audio objects. The processor unit 120 may be configured to determine, for example, for each audio object among the plurality of audio objects, whether this audio object belongs to the processing object group of audio objects depending on the position of the audio object metadata of this audio object, depending on the position of the region of interest, and depending on the radius range of the region of interest.
[0069] For example, the user may specify the position of the processing object group and the radius range of the processing object group. The position of the processing object group may specify the center point of the space, and the radius range of the processing object group may define a circle together with the center point of the processing object group. All audio objects located within the circle or on the line of the circle may be defined as the audio objects of this processing object group. That is, all audio objects located outside the circle will not be included in the processing object group. The area within the circle and on the line of the circle can be understood as the "region of interest".
[0070] According to one embodiment, the processor unit 120 may be configured to determine a weighting factor for each audio object of the processing object group of audio objects, for example, according to the distance between the position of the audio object metadata of this audio object and the position of the region of interest. The processor unit 120 may be configured to apply the weighting factor of this audio object to the gain parameter of the audio object signal or the audio object metadata of this audio object together with at least one effect parameter specified by the interface 110 for each audio object of the processing object group of audio objects, for example.
[0071] In such an embodiment, the influence of the effect parameter on each individual audio object of the processing object group is individualized for each audio object in addition to the effect parameter, and is individualized for each audio object by determining the weighting factor applied to the audio object.
[0072] In one embodiment, at least one defined parameter of a processing object group of audio objects may include, for example, at least one angle that specifies a direction from a defined user position where there is a region of interest associated with the processing object group of audio objects. The processor unit 120 may, for example, depend on the position of the metadata of this audio object and on an angle that specifies a direction from the defined user position where the region of interest is located, for each audio object of a plurality of audio objects, determine whether the audio object belongs to the processing object group of audio objects. For each audio object, it may be configured to determine whether the audio object belongs to the processing object group of audio objects.
[0073] According to one embodiment, the processor unit 120 may be configured to determine a weighting factor depending on a difference between a first angle and another angle, for example, for each audio object of the processing object group of audio objects, where the first angle is an angle that specifies a direction from the defined user position where the region of interest is located, and the other angle may be configured to depend on the defined user position and the position of the metadata of this audio object. The processor unit 120 may be configured to apply the weighting factor of this audio object to a gain parameter of the audio object signal or the audio object metadata of this audio object together with at least one effect parameter specified by the interface 110, for example, for each audio object of the processing object group of audio objects.
[0074] In one embodiment, the processing object group of audio objects may be, for example, a first processing object group of audio objects, and, for example, one or more other processing object groups of audio objects may additionally exist.
[0075] Each processing object group of the processing object groups of one or more other audio objects may include one or more audio objects among a plurality of audio objects, and at least one audio object among the processing object groups of the processing object groups of one or more other audio objects is not an audio object of the first processing object group of audio objects.
[0076] Here, the interface 110 may be configured to specify, on the user side, at least one other effect parameter for this processing object group for each processing object group of the processing object groups of one or more other audio objects.
[0077] In this context, the processor unit 120 may be configured to generate a processed signal such that, for each processing object group of the processing object groups of one or more other audio objects, at least one other effect parameter specified by the interface 110 is applied to the respective audio object signal or audio object metadata of one or more audio objects of this processing object group, where one or more audio objects among the plurality of audio objects do not belong to this processing object group.
[0078] Here, the processor unit 120 may be configured, for example, not to apply at least one other effect parameter of this processing object group specified by the interface to any audio object signal and any audio object metadata of one or more audio objects that do not belong to this processing object group.
[0079] In such an embodiment, it means that there may be multiple processing object groups. For each processing object group, one or more individual effect parameters are determined.
[0080] According to one embodiment, the interface 110 may be configured to specify, on the user side, one or more other processing object groups of one or more audio objects in addition to the first processing object group of audio objects. The interface 110 may be configured to specify, on the user side, at least one definition parameter of this processing object group for each processing object group of one or more other processing object groups of one or more audio objects.
[0081] In this context, the processor unit 120 may be configured to determine, for each processing object group of one or more other processing object groups of one or more audio objects, which audio objects among the plurality of audio objects belong to this processing object group, depending on at least one definition parameter of this processing object group specified by the interface 110.
[0082] Hereinafter, the concept and preferred embodiments of the embodiments of the present invention will be described.
[0083] In an embodiment, any type of global adaptation in OBAP can be made possible by converting the global adaptation into individual changes of the audio objects affected (e.g., by the processor unit 120).
[0084] Spatial mastering for object-based audio production can be implemented as follows, for example, by implementing the processing objects of the present invention.
[0085] The proposed implementation of the overall adaptation is implemented by a processing object (PO). Similar to conventional audio objects, it may be freely placed anywhere in the scene in real time. The user may apply any signal processing to a processing object (processing object group), such as an equalizer (EQ) or compression. In each of these processing tools, the parameter settings of the processing object may be converted into object-specific settings. Various methods are provided for this calculation.
[0086] The region of interest is shown below.
[0087] TIFF2025098034000002.tif15170
[0088] TIFF2025098034000003.tif40170
[0089] TIFF2025098034000004.tif79169
[0090] TIFF2025098034000005.tif13169
[0091] The following describes the calculation of inverse parameters according to an embodiment.
[0092] TIFF2025098034000006.tif15170
[0093] User adjustments to the processing object transformed by Equation (1) do not always yield the desired results at a sufficient speed because the exact position of the audio object is not taken into account. For example, if the area around the processing object is very large and the included audio object is far from the position of the processing object, the effect of the calculated adjustment may not be audible even at the position of the processing object.
[0094] TIFF2025098034000007.tif75169
[0095] TIFF2025098034000008.tif20170
[0096] In the following modified embodiments, angle-based calculations are performed.
[0097] TIFF2025098034000009.tif46170
[0098] Therefore, FIG. 7 shows the relative angle of the audio object with respect to the processing object according to one embodiment.
[0099] TIFF2025098034000010.tif20170
[0100] FIG. 8 shows an equalizer object having a new radial perimeter according to one embodiment.
[0101] TIFF2025098034000011.tif39168
[0102] TIFF2025098034000012.tif23170
[0103] In one embodiment, the implemented application is equalization.
[0104] Equalization can be regarded as the most important tool in mastering because the frequency response of the mix is the most important factor for proper conversion (transformation) throughout the playback system.
[0105] The proposed implementation of equalization is realized via an EQ object. Since all other parameters are independent of distance, only the gain parameter is particularly important.
[0106] In another embodiment, the implemented application is dynamic control.
[0107] In conventional mastering, dynamic compression is used to control the dynamic variations of the mix over time. As a result, depending on the compression settings, the perceived density and the transient response of the mix change. In the case of fixed compression, the change in perceived density is also called "glue", and stronger compression settings may use the side-chain effect in pump or beat-heavy mixes.
[0108] Using OBAP, the user can easily specify the same compression settings for multiple nearby objects and obtain multi-channel compression. However, the total compression for a group of audio objects is not only beneficial for time-critical workflows, but also increases the likelihood that the psychoacoustic impression is achieved by the so-called "glued" signal.
[0109] TIFF2025098034000013.tif9170
[0110] According to another embodiment, the implemented application is a deformation of the scene.
[0111] In stereo mastering, mid / side processing is a commonly used technique to expand or stabilize the stereo image of the mix. For spatial audio mixing, the characteristics of the room or speakers may be asymmetric, and similar options may be useful if the mix is created in an acoustically important environment. New creative opportunities for ME may be provided to improve the effects of the mix.
[0112] TIFF2025098034000014.tif22169
[0113] TIFF2025098034000015.tif50169
[0114] TIFF2025098034000016.tif92170
[0115] Since the position of an audio object may change over time, the coordinate position may be interpreted as a time-dependent function.
[0116] In one embodiment, a dynamic equalizer is implemented. In other embodiments, multi-band compression is implemented.
[0117] Object-based tone adjustment is not limited to the introduced equalizer application.
[0118] The above description is supplemented again below by a more general description of the embodiments.
[0119] Object-based 3D audio production follows an approach in which an audio scene is calculated and played back in real time for almost any speaker configuration via a rendering process. The audio scene describes the placement of audio objects as a function of time. An audio object is composed of an audio signal and metadata. These metadata include, in particular, the position and volume in the room, etc. Previously, in order to edit a scene, the user had to individually change all the audio objects in the scene.
[0120] When a processing object group and a processing object are mentioned on the one hand and on the other hand, it should be noted that for each processing object, the processing object group is defined to include audio objects. The processing object group is also called a container for processing objects. Therefore, for each processing object, a group of audio objects is defined among a plurality of audio objects. The corresponding processing object group includes the specified group of audio objects. Therefore, the processing object group is a group of audio objects.
[0121] The processing object may be defined as an object that can change the characteristics of other audio objects. The processing object is an artificial container that can associate any audio object, that is, the container is used to address all of the associated audio objects. The associated audio objects are affected by the number of effects. Therefore, the processing object enables the user to process multiple audio objects simultaneously.
[0122] The processing object has, for example, a position, an association method, a container, a weighting method, an audio signal processing effect, and a metadata effect.
[0123] The position is the position of the processing object in the virtual scene.
[0124] The association method associates audio objects with the processing object (optionally while using their positions).
[0125] The container (or connection) is a set of all audio objects (or, optionally, additional other processing objects) associated with the processing object.
[0126] The weighting method is an algorithm for calculating the individual effect parameter values of the assigned audio objects.
[0127] The audio signal processing effect changes each audio component (e.g., equalizer, dynamics) of the audio object.
[0128] The metadata effect changes the metadata of the audio object and / or the processing object (e.g., distortion of the position).
[0129] Similarly, the above positions, assignment methods, containers, association methods, audio signal processing effects, and metadata effects may also be associated with a processing object group. Here, the audio object of the container of the processing object is the audio object of the processing object group.
[0130] FIG. 11 shows the connection of processing objects resulting in audio signal effects and metadata effects according to one embodiment.
[0131] Hereinafter, the characteristics of the processing object will be described by specific embodiments.
[0132] The processing object may be arbitrarily placed within a scene by the user, and the position may be set over a certain period of time or as a function of time.
[0133] The processing object may have an effect assigned by the user that modifies the audio signal and / or metadata of the audio object. Examples of effects are equalization of the audio signal, processing of the dynamics of the audio signal, or change of the position coordinates of the audio object.
[0134] The processing object may have any number of effects assigned in any order.
[0135] The effect modifies the audio signal and / or metadata of the assigned set of audio objects and is either constant over time or time-dependent.
[0136] The effect has parameters that control the signal processing and / or metadata. These parameters are divided by the user into constant or weight parameters, or are defined by their respective types.
[0137] The effects of the processing object are copied and applied to the associated audio objects. The values of the constant parameters are adopted without being changed by each audio object. The values of the weight parameters are calculated individually for each audio object by using different weighting methods. The user may select the weighting method for each effect or may activate or deactivate it for individual audio sources.
[0138] The weighting method takes into account the individual metadata and / or the signal characteristics of the individual audio objects. For example, this may correspond to the distance between the audio object and the processing object or the frequency spectrum of the audio object. The weighting method may also take into account the listener's listening position. Furthermore, the weighting method may combine the aforementioned properties of the audio object to derive the parameter values individually. For example, the sound level of the audio object may be added in the context of dynamic processing to individually derive the change in volume of each audio object.
[0139] The effect parameters may be set to be constant or time-dependent over time. The weighting method takes such time variations into account.
[0140] The weighting method can also process the information that the audio renderer analyzes from the scene.
[0141] The sequence of assignment of effects to the processing object corresponds to the sequence of the processing signals or metadata of each audio object. That is, the data modified by the previous effect is used by the next effect as the basis for its calculation. The first effect functions based on the data of the audio object that has not yet been changed.
[0142] Individual effects can be deactivated. Then, if there is calculation data for a previous effect, the calculation data for the previous effect will be transferred to the effect following the deactivated effect.
[0143] An explicitly newly developed effect is a change in the position of an audio object by homography (the "distortion effect"). The user is presented with a rectangle having corners that can be individually moved to the position of the processing object. When the user moves a corner, this distortion transformation matrix is calculated from the previous state and the newly distorted state of the rectangle. The matrix is then applied to all position coordinates of the audio objects associated with the processing object, and their positions are changed by the distortion.
[0144] Effects that only change metadata may also be applied to other processing objects (especially the "distortion effect").
[0145] An audio source may be associated with a processing object in various ways. The number of associated audio objects may also change over time depending on the type of association. All such changes are taken into account in all calculations.
[0146] The affected area may be defined around the position of the processing object.
[0147] All audio objects placed within the affected area form a set of associated audio objects to which the effect of the processing object is applied.
[0148] The affected area may be any body (3D) or any shape (2D) defined by the user.
[0149] The center of the affected area may correspond to the position of the processing object, but it does not have to. This is specified by the user.
[0150] If its position is within a three-dimensional body, the audio object is within the region affected by three-dimensional effects.
[0151] If its position projected onto the horizontal plane is within a two-dimensional shape, the audio object is within the region affected by two-dimensional effects.
[0152] Since the affected region can assume all surrounding sizes that are not specified, all audio objects in the scene are placed within the affected region.
[0153] If necessary, the affected region adapts to changes in scene properties (e.g., scene scaling).
[0154] Regardless of the affected region, the processing object may be linked to any selection of audio objects in the scene.
[0155] The coupling may be defined by the user, and all selected audio objects form a set of audio objects to which the effect of the processing object is applied.
[0156] Alternatively, the coupling may be defined by the user in such a way that the processing object adjusts its position as a function of time according to the positions of the selected audio objects. This adjustment of the position may take into account the listener's listening position. In this context, the effect of the processing object does not necessarily have to be applied to the coupled audio objects.
[0157] The relationship may be automatically performed based on user-defined criteria. In this context, all audio objects in the scene are continuously inspected for the defined criteria, and if the criteria are met, they are associated with the processing object. The period of association may be limited to the time when the criteria are met, and a transition period may be defined. The transition period determines the period during which one or more criteria are continuously met by the audio object and it is associated with the processing object, or the period during which one or more criteria are continuously ignored and the relationship to the processing object is ignored again.
[0158] The processing objects may be deactivated by the user, so their properties are retained and continue to be displayed to the user without the audio objects being affected by the processing objects.
[0159] The user can couple any number of properties of a processing object with similar properties of any number of other processing objects. These properties include effect parameters. The user can select whether the coupling is absolute or relative. In the case of constant coupling, the changed property values of the processing object are exactly adopted by all the coupled processing objects. In the case of relative coupling, the changed value is offset with respect to the property values of the coupled processing objects.
[0160] The processing object may be replicated. In so doing, a second processing object with the same properties as the original processing object is generated. The properties of the processing objects are independent of each other.
[0161] The properties of the processing object can be permanently inherited, for example when copying, so that changes made by the parent are automatically adopted by the child.
[0162] Figure 12 shows the modification of audio objects and audio signals in response to user input according to an embodiment.
[0163] Another new application of the processing object is the calculation of intelligent parameters using scene analysis. The user defines the effect parameters at a specific position via the processing object. The audio renderer performs predictive scene analysis to detect the audio sources that affect the position of the processing object. Then, considering the scene analysis, the effect is applied to the selected audio source in such a way that the user-defined effect settings are optimized at the position of the processing object.
[0164] In the following, further embodiments of the present invention visually represented by FIGS. 13-25 will be described.
[0165] For example, FIG. 13 shows a processing object PO4 having a rectangle M for the distortion of the corner portions C1, C2, C3, and C4 on the user side. FIG. 13 schematically shows the possible distortion to M' having corner portions C1', C2', C3', and C4', and the corresponding effects with sources S1, S2, S3, and S4 having new positions S1', S2', S3', and S4'.
[0166] TIFF2025098034000017.tif31170
[0167] TIFF2025098034000018.tif27170
[0168] FIG. 16 shows a possible schematic implementation of an equalizer effect applied to a processing object. Buttons such as w next to each parameter can be used to activate the weight of each parameter. m1, m2, and m3 provide options for the weighting method for the aforementioned weight parameters.
[0169] TIFF2025098034000019.tif27170
[0170] Figure 18 shows a typical implementation of a processing object to which an equalizer is applied. The cyan object with the wave symbol on the right side of the image represents the processing object of the audio scene and can be freely moved by the user with the mouse. Within the cyan transparent and uniform area around the processing object, the equalizer parameters are not changed and are applied to the audio objects Src1, Src2, and Src3 defined on the left side of the image. The shading that moves into the transparent area around the uniform circular area indicates the area where all parameters except the gain parameter are adopted without being changed by the source. In contrast, the gain parameter of the equalizer is weighted by the distance between the source and the processing object. Since only sources Src4 and Src24 are in this area, in this case, the weighting is only done for those parameters. Source Src22 is not affected by the processing object. The user controls the size of the radius of the circular area around the processing object with the "Area" slider. With the "Feather" slider, the user controls the size of the radius of the surrounding transition area.
[0171] Figure 19 shows a processing object like that in Figure 18 but in a different position and without a transition area. All parameters of the equalizer are adopted for sources Src22 and Src4 without being changed. Sources Src3, Src2, Src1, and Src24 are not affected by the processing object.
[0172] Figure 20 shows a processing object having an area defined as an area affected by its azimuth angle, and sources Src22 and Src4 are associated with the processing object. The vertex of the affected area in the center of the right side of the image corresponds to the position of the listener / user. When the processing object is moved, the area is moved by the azimuth angle. Using the "Area" slider, the user determines the angular size of the affected area. The user can change from a circular to an angle-based affected plane via a low selection range by the "Area" / "Feather" slider, which is currently displayed as "Radius".
[0173] Figure 21 shows a processing object like that in Figure 20, but having an additional transition area that can be controlled by the user via the "Feather" slider.
[0174] Figure 22 shows several processing objects in a scene having different affected areas. The gray processing objects are deactivated by the user. That is, they do not affect the audio objects within their affected areas. On the left side of the image, the equalizer parameters of the currently selected processing object are always displayed. The selection range is indicated by a thin bright cyan line around the object.
[0175] Figure 23 shows that the red square on the right side of the image indicates a processing object for the horizontal distortion of the position of an audio object. The user can distort the scene by dragging a corner in any direction with the mouse.
[0176] Figure 24 shows the scene after the user has dragged a corner of the processing object. The positions of all sources have changed due to the distortion.
[0177] Figure 25 shows a possible visualization of the association of individual audio objects having processing objects.
[0178] Although some aspects have been described in the context of apparatus, it will be apparent that such aspects also represent a description of corresponding methods, and that blocks or components of the apparatus should also be understood as corresponding method steps or as functions of method steps. Similarly, aspects described in the context of method steps represent a description of corresponding blocks, details or functions of corresponding apparatus. Some or all of the method steps may be executed by (or using) hardware devices such as, for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, some one or more of the most important method steps may be executed by such devices.
[0179] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or in software or in at least a part of the hardware or in at least a part of the software. The implementation can be carried out using a digital storage medium, such as a floppy disk, a DVD, a Blu-ray disc, a CD, a ROM, a PROM, an EPROM, an EEPROM or a flash memory, a hard disk or any other magnetic or optical memory, which has electronically readable control signals stored thereon and which can cooperate or cooperates with a programmable computer system such that respective methods are executed. Therefore, the digital storage medium can be made computer-readable.
[0180] Some embodiments according to the present invention comprise a data carrier containing electronically readable control signals which can cooperate with a programmable computer system such that any of the methods described herein are executed.
[0181] In general, embodiments of the present invention can be implemented as a computer program product with program code that is operative to perform any of the methods of the present invention when the computer program product runs on a computer.
[0182] The program code can be stored, for example, in a machine-readable carrier.
[0183] Other embodiments include a computer program for executing the methods described in the present specification, and the computer program is stored in a machine-readable carrier. In other words, one embodiment of the method of the present invention is a computer program having program code for executing any of the methods described in the present specification when the computer program operates on a computer.
[0184] A further embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) on which a computer program for executing any of the methods described in the present specification is recorded. The data carrier or digital storage medium or computer-readable medium is usually tangible and / or non-volatile.
[0185] A further embodiment of the method of the present invention is a sequence of data streams or signals representing a computer program for executing any of the methods described in the present specification. The sequence of data streams or signals can be configured to be transferred, for example, by a data communication connection, such as the Internet.
[0186] A further embodiment includes processing means, such as a computer or a programmable logic device, configured or adapted to execute any of the methods described in the present specification.
[0187] A further embodiment includes a computer on which a computer program for executing any of the methods described in the present specification is installed.
[0188] Further embodiments of the present invention include an apparatus or system configured to transfer a computer program that executes at least one of the methods described in the present specification to a receiver. The transfer is made, for example, electronically or optically. The receiver can be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system can include, for example, a file server that transfers a computer program to the receiver.
[0189] In some embodiments, a programmable logic device (e.g., a field programmable gate array, FPGA) can be used to perform some or all of the functions of the methods described in the present specification. In some embodiments, the field programmable gate array can cooperate with a microprocessor to perform any of the methods described in the present specification. Generally, the method is preferably executed by any hardware device. The hardware device may be a generally applicable hardware such as a computer processor (CPU), or may be a method-specific hardware such as an ASIC.
[0190] The above-described embodiments merely illustrate the principles of the present invention. Modifications and changes to the configurations and details described in the present specification will be apparent to those skilled in the art. Therefore, it is intended that the present invention be limited only by the scope of the impending claims and not by the specific details represented by the description and explanation of the embodiments in the present specification.
[0191] References [1]Coleman, P., Franck, A., Francombe, J., Liu, Q., Campos, T. D., Hughes, R., Men-zies, D., Galvez, M. S., Tang, Y., Woodcock, J., Jackson, P., Melchior, F., Pike, C., Fazi, F., Cox, T, and Hilton, A., "An Audio-Visual System for Object-Based Audio: From Recording to Listening," IEEE Transactions on Multimedia, PP(99), pp. 1-1, 2018, ISSN 1520- 9210, doi:10.1109 / TMM.2018.2794780. [2] Gasull Ruiz, A., Sladeczek, C., and Sporer, T., "A Description of an Object-Based Audio Workflow for Media Productions," in Audio Engineering Society Conference: 57th International Conference: The Future of Audio Entertainment Technology, Cinema, Television and the Internet, 2015. [3] Melchior, F., Michaelis, U., and Steffens, R., "Spatial Mastering - a new concept for spatial sound design in object-based audio scenes," in Proceedings of the International Computer Music Conference 2011, 2011. [4] Katz, B. and Katz, R. A., Mastering Audio: The Art and the Science, Butterworth-Heinemann, Newton, MA, USA, 2003, ISBN 0240805453, AES Conference on Spatial Reproduction, Tokyo, Japan, 2018 August 6 - 9, page 2 [5] Melchior, F., Michaelis, U., and Steffens, R., "Spatial Mastering - A New Concept for Spatial Sound Design in Object-based Audio Scenes," Proceedings of the International Computer Music Conference 2011, University of Huddersfield, UK, 2011. [6] Sladeczek, C., Neidhardt, A., Boehme, M., Seeber, M., and Ruiz, A. G., "An Approach for Fast and Intuitive Monitoring of Microphone Signals Using a Virtual Listener," Proceedings, International Conference on Spatial Audio (ICSA), 21.2. - 23.2.2014, Erlangen, 2014 [7] Dubrofsky, E., Homography Estimation, Master's thesis, University of British Columbia, 2009. [8] ISO / IEC 23003-2:2010 Information technology - MPEG audio technologies - Part 2: Spatial Audio Object Coding (SAOC); 2010
Claims
1. 1. An apparatus for generating a processed signal using a plurality of audio objects, each audio object of the plurality of audio objects comprising an audio object signal and audio object metadata, the audio object metadata comprising a position of the audio object and a gain parameter of the audio object, the apparatus comprising: an interface (110) for a user to specify at least one effect parameter of a processing object group of audio objects, the processing object group of audio objects including two or more audio objects of the plurality of audio objects; a processor unit (120) configured to generate the processed signal such that the at least one effect parameter specified by the interface (110) is applied to the audio object signal or to the audio object metadata of each audio object of a processed object group of the audio objects; An apparatus comprising:
2. one or more of the plurality of audio objects do not belong to a processing object group of the audio objects, and the processor unit (120) is configured not to apply the at least one effect parameter specified by the interface to any audio object signal and any audio object metadata of the one or more audio objects that do not belong to a processing object group of the audio objects.
2. The apparatus of claim 1.
3. the processor unit (120) is configured to generate the processed signal such that the at least one effect parameter specified by the interface (110) is applied to the audio object signal of each of the audio objects of a processing object group of the audio objects; the processor unit (120) is configured to not apply the at least one effect parameter specified by the interface to any audio object signal of the one or more audio objects of the plurality of audio objects that do not belong to a processing object group of the audio object.
3. The apparatus of claim 2.
4. the processor unit (120) is configured to generate the processed signal such that the at least one effect parameter specified by the interface (110) is applied to the gain parameter of the metadata of each audio object of a processing object group of the audio objects; the processor unit (120) is configured to not apply the at least one effect parameter specified by the interface to any gain parameter of the audio object metadata of the one or more audio objects of the plurality of audio objects that do not belong to a processing object group of the audio object.
4. Apparatus according to claim 2 or claim 3.
5. the processor unit (120) is configured to generate the processed signal such that the at least one effect parameter specified by the interface (110) is applied to the position of the metadata of each audio object of a processing object group of the audio objects; the processor unit (120) is configured to not apply the at least one effect parameter specified by the interface to any position in the audio object metadata of the one or more audio objects of the plurality of audio objects that do not belong to a processing object group of the audio object. Apparatus according to any one of claims 2 to 4.
6. the interface (110) is configured for the user to specify at least one definition parameter of a processing object group of the audio object; the processor unit (120) is configured to determine, depending on the definition parameter of the at least one of the processing object groups of the audio objects specified by the interface (110), which audio objects of the plurality of audio objects belong to the processing object group of the audio objects.
6. Apparatus according to any one of claims 1 to 5.
7. said at least one definition parameter of said processing object group of said audio objects comprises at least one location of a region of interest associated with said processing object group of said audio objects; the processor unit (120) is configured to determine, for each audio object of the plurality of audio objects, whether the audio object belongs to a processing object group of said audio objects depending on the location of the audio object metadata of said audio object and depending on the location of the region of interest.
7. The apparatus of claim 6.
8. and wherein the at least one definition parameter of the processing object group of the audio object further comprises a radial extent of the region of interest associated with the processing object group of the audio object; the processor unit (120) is configured to determine, for each audio object of the plurality of audio objects, whether the audio object belongs to a processing object group of the audio object in dependence on the position of the audio object metadata of the audio object, in dependence on the position of the region of interest, and in dependence on the radial range of the region of interest.
8. The apparatus of claim 7.
9. the processor unit (120) is adapted to determine, for each audio object of the processing object group of audio objects, a weighting factor depending on the distance between the location of the audio object metadata of said audio object and the location of the region of interest; the processor unit (120) is configured for applying, for each audio object of the processing object group of audio objects, the weighting factor of that audio object together with the at least one effect parameter specified by the interface (110) to the gain parameter of the audio object signal or of the audio object metadata of that audio object.
9. Apparatus according to claim 7 or claim 8.
10. said at least one definition parameter of said processing object group of audio objects comprises at least one angle specifying a direction from a defined user position in which a region of interest associated with said processing object group of audio objects lies; the processor unit (120) is configured to determine, for each audio object of the plurality of audio objects, whether the audio object belongs to a processing object group of the audio object depending on the position of the metadata of the audio object and depending on the angle specifying the direction from the defined user position in which the region of interest is located.
7. The apparatus of claim 6.
11. the processor unit (120) is configured to determine, for each audio object of a processing object group of audio objects, a weighting factor that depends on a difference between a first angle and another angle, the first angle being the angle specifying the direction from the defined user position in which the region of interest is located, and the other angle being dependent on the defined user position and the position of the metadata of the audio object; the processor unit (120) is configured for each audio object of the processing object group of audio objects to apply the weighting factor of said audio object together with the at least one effect parameter specified by the interface (110) to the gain parameter of the audio object signal or of the audio object metadata of said audio object.
11. The apparatus of claim 10.
12. the processing object group of the audio object is a first processing object group of audio objects, and there are also processing object groups of one or more further audio objects, each processing object group of the one or more further audio objects comprising one or more audio objects of the plurality of audio objects, and at least one audio object of the processing object group of the processing object group of the one or more further audio objects is not an audio object of the first processing object group of the audio objects, the interface (110) is configured for a user to specify, for each processing object group of the one or more further audio objects, at least one further effect parameter for the processing object group of the audio object; the processor unit (120) is configured to generate the processed signal such that, for each processing object group of the one or more further audio objects, the at least one further effect parameter of the processing object group specified by the interface (110) is applied to the audio object signal or the audio object metadata of each of the one or more audio objects of the processing object group, wherein one or more audio objects of the plurality of audio objects do not belong to the processing object group, and the processor unit (120) is configured not to apply the at least one further effect parameter of the processing object group specified by the interface to any audio object signal and any audio object metadata of the one or more audio objects that do not belong to the processing object group.
12. Apparatus according to any one of claims 1 to 11.
13. the interface (110) is configured for the user to specify, in addition to the first processing object group of the audio objects, the one or more further processing object groups of one or more audio objects, the interface (110) being configured for the user to specify, for each processing object group of the one or more further processing object groups of one or more audio objects, at least one definition parameter of the processing object group; the processor unit (120) is configured to determine, for each processing object group of the one or more further processing object groups of one or more audio objects, which audio objects belong to the plurality of audio objects of the processing object group in dependence on the at least one definition parameter of the processing object group specified by the interface (110).
13. The apparatus of claim 12.
14. the apparatus is an encoder, the processor unit (120) is configured to generate a downmix signal using the audio object signals of the plurality of audio objects, and the processor unit (120) is configured to generate a metadata signal using the audio object metadata of the plurality of audio objects, the processor unit (120) is configured to generate the downmix signal as the processed signal, and for each audio object of a processing object group of the audio objects, at least one modified object signal is mixed into the downmix signal, and the processor unit (120) is configured to generate, for each audio object of a processing object group of the audio objects, the modified object signal of the audio object by applying the at least one effect parameter specified by the interface (110) to the audio object signal of the audio object, or the processor unit (120) is configured to generate the metadata signal as the processed signal, the metadata signal including at least one modified position for each audio object of a processing object group of the audio objects, and the processor unit (120) is configured to generate, for each audio object of a processing object group of the audio objects, the modified position of the audio object by applying the at least one effect parameter specified by the interface (110) to the position of the audio object, or the processor unit (120) is configured to generate the metadata signal as the processed signal, the metadata signal including at least one modified gain parameter for each audio object of a processing object group of the audio objects, and the processor unit (120) is configured to generate, for each audio object of a processing object group of the audio objects, the modified gain parameter of the audio object by applying the at least one effect parameter specified by the interface (110) to the gain parameter of the audio object.
14. Apparatus according to any one of claims 1 to 13.
15. the apparatus is a decoder, the apparatus is configured to receive a downmix signal in which the audio object signals of the plurality of audio objects are mixed, the apparatus is further configured to receive a metadata signal, the metadata signal comprising, for each audio object of the plurality of audio objects, the audio object metadata of that audio object, the processor unit (120) is configured to reconstruct the audio object signals of the audio objects based on a downmix signal; the processor unit (120) is configured to generate, as the processed signal, an audio output signal comprising one or more audio output channels; the processor unit (120) is configured to apply the at least one effect parameter specified by the interface (110) to the audio object signal of each of the audio objects of a processing object group of the audio objects to generate the processed signal, or to apply the at least one effect parameter specified by the interface (110) to the position or the gain parameter of the audio object metadata of each of the audio objects of a processing object group of the audio objects to generate the processed signal.
14. Apparatus according to any one of claims 1 to 13.
16. The interface (110) is further adapted for the user to specify one or more rendering parameters; the processor unit (120) is configured to generate the processed signal using the one or more rendering parameters as a function of the position of each audio object of the processing object group of the audio objects.
16. The apparatus of claim 15.
17. an encoder (200) for generating a downmix signal based on audio object signals of a plurality of audio objects and for generating a metadata signal based on audio object metadata of said plurality of audio objects, said audio object metadata including positions of said audio objects and gain parameters of said audio objects; a decoder (300) for generating an audio output signal comprising one or more audio output channels based on the downmix signal and based on the metadata signal; Including, The encoder (200) is an apparatus according to claim 14, or The decoder (300) A device according to claim 15 or claim 16, or The encoder (200) is a device according to claim 14, and the decoder (300) comprises 17. The device according to claim 15 or claim 16, system.
18. 1. A method for generating a processed signal using a plurality of audio objects, each audio object of the plurality of audio objects comprising an audio object signal and audio object metadata, the audio object metadata comprising a position of the audio object and a gain parameter of the audio object, the method comprising: a user using an interface (110) specifying at least one effect parameter of a processing object group of audio objects, the processing object group of audio objects comprising at least two audio objects of the plurality of audio objects; generating, by a processor unit (120), the processed signal such that the at least one effect parameter specified by a user using the interface is applied to the audio object signal or to the audio object metadata of each of the two or more audio objects of a processed object group of audio objects; A method comprising:
19. A computer program comprising a program code for performing the method according to claim 18.