Apparatus and method for object-based spatial audio-mastering

The device and method for audio mastering organize audio objects into processing groups, enabling simultaneous adjustments through an interface, addressing inefficiencies in conventional tools and enhancing sound quality in object-based spatial audio production.

EP3756363B1Active Publication Date: 2025-10-15FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
EP2019710283
Authority / Receiving Office
EP · EP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-04-19
Filing Date
2019-02-18
Publication Date
2025-10-15
Estimated Expiration
2039-02-18

AI Technical Summary

Technical Problem

Conventional audio mastering tools are inadequate for object-based spatial audio, as they cannot efficiently adjust individual audio objects within complex scenes, leading to inefficient production processes and limited sound quality due to environmental constraints.

Method used

A device and method for audio mastering that organizes audio objects into processing object groups, allowing users to apply effect parameters to multiple objects simultaneously through an interface, with a processor unit determining which objects belong to the group based on metadata and user-defined criteria, enabling real-time adjustments.

Benefits of technology

Facilitates efficient and flexible audio mastering by allowing simultaneous adjustments to multiple audio objects, improving sound quality and reducing production barriers for object-based spatial audio content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF0001
    Figure IMGF0001
  • Figure IMGF0002
    Figure IMGF0002
  • Figure IMGF0003
    Figure IMGF0003
Patent Text Reader

Abstract

The invention relates to an apparatus for generating a processed signal using a plurality of audio objects according to an embodiment, each audio object of the plurality of audio objects comprising an audio object signal and audio object metadata, and the audio object metadata comprising a position of the audio object and a gain parameter of the audio object. The apparatus comprises an interface (110) for the user to specify at least one effect parameter of a processing object group of audio objects, the processing object group of audio objects comprising two or more audio objects of the plurality of audio objects. The apparatus also comprises a processor unit (120) which is designed to generate the processed signal such that the at least one effect parameter specified by means of the interface (110) is applied to the audio object signal or to the audio object metadata of each of the audio objects of the processing object group of audio objects. One or more audio objects of the plurality of audio objects do not belong to the processing object group of audio objects.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The application relates to audio object processing, audio object encoding and audio object decoding and, in particular, audio mastering for audio objects.

[0002] Object-based spatial audio is an approach to interactive three-dimensional audio reproduction. This concept not only changes how content creators or authors can interact with the audio, but also how it is stored and transmitted. To enable this, a new process in the reproduction chain, called "rendering," must be established. The rendering process generates loudspeaker signals from an object-based scene description. Although recording and mixing have been explored in recent years, concepts for object-based mastering are almost nonexistent. The main difference compared to channel-based audio mastering is that instead of adjusting the audio channels, the audio objects must be modified. This requires a fundamentally new approach to mastering. This paper presents a new method for mastering object-based audio.

[0003] In recent years, the object-based audio approach has generated considerable interest. Compared to channel-based audio, where speaker signals are stored as the result of spatial audio production, the audio scene is described using audio objects. An audio object can be considered a virtual sound source consisting of an audio signal with additional metadata, such as position and gain. To reproduce audio objects, an audio renderer is required. Audio rendering is the process of generating speaker or headphone signals based on additional information, such as the position of speakers or the position of the listener in the virtual scene.

[0004] The process of audio content creation can be divided into three main parts: recording, mixing, and mastering. While all three steps have been extensively addressed for channel-based audio over the past decades, object-based audio requires new workflows for future applications. So far, the recording step generally does not need to be changed, even though future technologies might bring new possibilities [1], [2]. The mixing process is somewhat different, as the sound engineer no longer creates a spatial mix by panning signals to dedicated speakers. Instead, all positions of audio objects are created by a spatial authoring tool that allows defining the metadata part of each audio object. A complete mastering process for audio objects has not yet been established [3].

[0005] Conventional audio mixes route multiple audio tracks to a specific number of output channels. This requires creating individual mixes for different playback configurations, but allows for efficient handling of the output channels during mastering [4]. When using the object-based audio approach, the audio renderer is responsible for creating all loudspeaker signals in real time. Arranging a large number of audio objects as part of a creative mixing process results in complex audio scenes. However, since the renderer can reproduce the audio scene in several different loudspeaker setups, it is not possible to directly address the output channels during production. The mastering concept can therefore only be based on individually modifying audio objects.

[0006] To date, conventional audio production is geared toward highly specific listening devices and their channel configurations, such as stereo or surround playback. The decision regarding which playback device(s) the content is intended for must therefore be made at the beginning of its production. The production process itself then consists of recording, mixing, and mastering. The mastering process optimizes the final mix to ensure that it plays back with satisfactory quality on all consumer systems with different speaker characteristics. Since the desired output format of a mix is ​​fixed, the mastering engineer (ME) can create an optimized master for this playback configuration.

[0007] The mastering phase makes it more beneficial for creators to produce audio in suboptimal acoustic environments, as they can rely on a final check of their mix during mastering. This lowers the barriers to entry for producing professional content. On the other hand, over the years, MEs themselves have been offered a wide range of mastering tools, drastically improving their options for corrections and enhancements. Nevertheless, the final content is usually limited to the playback setup for which it was designed.

[0008] This limitation is fundamentally overcome by Object-Based Spatial Audio Production (OBAP). Unlike channel-based audio, OBAP is based on individual audio objects with metadata that encompasses their position in an artificial environment, also referred to as a "scene." Only at the final listening output does a dedicated rendering unit, the renderer, calculate the final loudspeaker signals in real time based on the listener's loudspeaker setup.

[0009] Although OBAP provides each audio object and its metadata individually to the renderer, direct channel-based adjustments are not possible during production, and thus, existing mastering tools for conventional playback setups cannot be used. Meanwhile, OBAP requires that all final adjustments be made in the mix. While the requirement to implement overall sound adjustments by manually handling each individual audio object is not only highly inefficient, it also places significant demands on each creator's monitoring setup and strictly limits the sound quality of object-based 3D audio content to the acoustic properties of the environment in which it was created.

[0010] Ultimately, developing tools to enable a similarly powerful mastering process for OBAP on the creator side could improve the acceptance of producing 3D audio content by lowering production barriers and opening up new space for sound aesthetics and sound quality.

[0011] While initial ideas about spatial mastering have been made public [5], this paper introduces new approaches for adapting conventional mastering tools and which types of new tools might be considered helpful in mastering for object-based spatial audio. [5] describes a basic sequence for using metadata to derive object-specific parameters from global properties. Furthermore, [6] describes the concept of a region of interest with a surrounding transition region in the context of OBAP applications.

[0012] THIBAUT CARPENTIER: "Panoramix: 3D mixing and postproduction workstation", 42ND INTERNATIONAL COMPUTER MUSIC CONFERENCE (ICMC) 2016, 1 September 2016 (2016-09-01), pages 122-127, XP055586741, presents Panoramix, a device for post-processing 3D audio content by mixing, reverberation, and surround sound generation.

[0013] Anonymous: "Spatial Audio Workstation User Manual, Version 2.4.0" Barco® Audio Technologies Spatial Audio Workstation User Manual", according to authors last edited on July 31, 2017 (2017-07-31), pages 1-50, XP055586950, URL: http: / / www.iosono-sound.com / uploads / downloads / SAW_User_Manual_24.pdf shows a device for processing numerous audio objects, whereby the position of the audio objects can be changed during processing.

[0014] SCHEIRER ED ET AL: "AUDIOBIFS: DESCRIBING AUDIO SCENES WITH THE MPEG-4 MUL TI MEDIA STANDARD", IEEE TRANSACTIONS ON MUL TI MEDIA, IEEE SERVICE CENTER, PISCATAWAY, NJ, US, Vol. 1, No. 3, 1 September 1999 (1999-09-01), pages 237-250, XP001011325, ISSN: 1520-9210, DOI: 10.1109 / 6046.784463 shows a system for the flexible generation of complex audio scenes using compressed audio information, synthesized audio and 3D audio content.

[0015] DE 10 2010 030534 A1 (IOSONO GMBH [DE]) December 29, 2011 (2011-12-29) shows a device for modifying an audio scene, comprising a direction determiner and an audio scene processing device, wherein the audio scene has at least one audio object comprising an audio signal and associated metadata.

[0016] WO 2013 / 006338 A2 (DOLBY LAB LICENSING CORP [US]; ROBINSON CHARLES Q [US] ET AL.) January 10, 2013 (2013-01-10) shows an adaptive audio system for processing a number of independent audio data streams.

[0017] US 2010 / 223552 A1 (VERAX TECHNOLOGIES INC [US]; METCALF RANDALL B [US]) September 2, 2010 (2010-09-02) shows a system for further processing a large number of audio objects, in which the audio objects can be individually controlled.

[0018] Anonymous: "Spatial Audio Workstation User Manual Version 2.3.0", April 5, 2017 (2017-04-05), pages 1-49, XP055926479, Found on the Internet: URL:https: / / web.archive.org / web / 20170405125152 / http: / / www.iosonosound.com / uploads / downloads / SAW-User-Manual-230.pdf also shows a device for processing numerous audio objects, whereby the position of the audio objects can be changed during processing.

[0019] It is therefore desirable to provide improved object-based audio mastering concepts.

[0020] The object of the invention is achieved by the subject matter of the independent patent claims. Particular embodiments are shown in the dependent claims.

[0021] A device for generating a processed signal using a plurality of audio objects according to one embodiment is provided, wherein each audio object of the plurality of audio objects comprises an audio object signal and audio object metadata, wherein the audio object metadata comprises a position of the audio object and a gain parameter of the audio object. The device comprises: an interface for specifying at least one effect parameter of a processing object group of audio objects by a user, wherein the processing object group of audio objects comprises two or more audio objects of the plurality of audio objects. Furthermore, the device comprises a processor unit configured to generate the processed signal such that the at least one effect parameter specified by means of the interfaceis applied to the audio object signal or to the audio object metadata of each of the audio objects of the processing object group of audio objects. One or more audio objects of the plurality of audio objects do not belong to the processing object group of audio objects. The interface is designed for the user to specify at least one definition parameter of the processing object group of audio objects, wherein the processor unit is designed to determine, depending on the at least one definition parameter of the processing object group of audio objects specified by means of the interface, which audio objects of the plurality of audio objects belong to the processing object group of audio objects, wherein the at least one definition parameter of the processing object group of audio objects comprises at least one angle that specifies a direction from a defined user position,in which there is an area of ​​interest associated with the processing object group of audio objects, and wherein the processor unit is configured to determine for each audio object of the plurality of audio objects, depending on the position of the metadata of this audio object and depending on the angle specifying the direction from the defined user position in which the area of ​​interest is located, whether this audio object belongs to the processing object group of audio objects.

[0022] Furthermore, a method for generating a processed signal using a plurality of audio objects is provided, wherein each audio object of the plurality of audio objects comprises an audio object signal and audio object metadata, wherein the audio object metadata comprises a position of the audio object and a gain parameter of the audio object. The method comprises: specifying at least one effect parameter of a processing object group of audio objects by a user via an interface, wherein the processing object group of audio objects comprises two or more audio objects of the plurality of audio objects. And: generating the processed signal by a processor unit such that the at least one effect parameter specified via the interfaceapplied to the audio object signal or to the audio object metadata of each of the audio objects of the processing object group of audio objects. The method further comprises: specifying at least one definition parameter of the processing object group of audio objects by the user, determining which audio objects of the plurality of audio objects belong to the processing object group of audio objects depending on the at least one definition parameter of the processing object group of audio objects specified by means of the interface, wherein the at least one definition parameter of the processing object group of audio objects comprises at least one angle specifying a direction from a defined user position in which a region of interest associated with the processing object group of audio objects is located,and wherein for each audio object of the plurality of audio objects, it is determined whether this audio object belongs to the processing object group of audio objects depending on the position of the metadata of this audio object and depending on the angle that specifies the direction from the defined user position in which the area of ​​interest is located.

[0023] Furthermore, a computer program with a program code for carrying out the method described above is provided.

[0024] The provided audio mastering is based on the mastering of audio objects. In some embodiments, these can be positioned anywhere in a scene and freely in real time. In some embodiments, for example, the properties of general audio objects are influenced. In their function as artificial containers, they can each contain an arbitrary number of audio objects. Any adjustment to a mastering object is converted in real time into individual adjustments to audio objects of the same.

[0025] Such mastering objects are also called processing objects.

[0026] Thus, instead of adjusting numerous audio objects separately, the user can use a mastering object to perform mutual adjustments on multiple audio objects simultaneously.

[0027] For example, according to embodiments, the set of target audio objects for a mastering object can be defined in numerous ways. From a spatial perspective, the user can define a user-defined scope around the position of the mastering object. Alternatively, it is possible to link individually selected audio objects to the mastering object regardless of their position. The mastering object also takes into account potential changes in the position of audio objects over time.

[0028] A second property of mastering objects according to embodiments may, for example, be their ability to calculate how each audio object is individually affected based on interaction models. Similar to a channel strip, a mastering object can, for example, incorporate any general mastering effect, such as equalizers and compressors. Effect plug-ins typically provide the user with numerous parameters, e.g., for frequency or gain control. When a new mastering effect is added to a mastering object, it is automatically copied to all audio objects in the target set. However, not all effect parameter values ​​are transferred unchanged. Depending on the calculation method for the target set, some parameters of the mastering effect may be weighted before being applied to a specific audio object.The weighting can be based on any metadata or a sound characteristic of the audio object.

[0029] Preferred embodiments of the invention are described below with reference to the drawings.

[0030] The drawings show: Fig. 1 shows a device for generating a processed signal using a plurality of audio objects according to one embodiment. Fig. 2 shows a device according to another embodiment, wherein the device is an encoder. Fig. 3 shows a device according to another embodiment, wherein the device is a decoder. Fig. 4 shows a system according to one embodiment. Fig. 5 shows a processing object with the region A and the fading region. A f according to an embodiment. Fig. 6 shows a processing object with the area A and object radii according to an embodiment. Fig. 7 shows a relative angle of audio objects to the processing object according to an embodiment. Fig. 8 shows an equalizer object with a new radial circumference according to an embodiment. Fig. 9 shows a signal flow of a compression of the signals from n sources according to an embodiment. Fig. 10 shows a scene transformation using a control panel M according to an embodiment. Fig. 11 shows the relationship of a processing object with which audio signal effects and metadata effects are caused, according to an embodiment. Fig. 12 shows the modification of audio objects and audio signals in response to a user input according to an embodiment. Fig. 13 shows a processing object PO 4 with rectangle M for distorting the corners C 1 , C 2 , C 3 and C 4 by the user according to an embodiment.14 shows processing objects PO 1 and PO 2 with their respective overlapping two-dimensional catchment areas A and B according to one embodiment. Fig. 15 shows processing object PO 3 with rectangular, two-dimensional catchment area C and the angles between PO 3 and the associated sources S 1 , S 2 and S 3 according to one embodiment. Fig. 16 shows a possible schematic implementation of an equalizer effect applied to a processing object according to one embodiment. Fig. 17 shows processing object PO 5 with a three-dimensional catchment area D and the respective distances d S1 , d S2 and d S3 to the sources S 1 , S 2 and S 3 associated across the catchment area according to one embodiment. Fig. 18 shows a prototypical implementation of a processing object to which an equalizer has been applied according to one embodiment. Fig. 19 shows a processing object as in . Fig. 18 , only at a different position and without a transition area according to one embodiment. Fig. 20 shows a processing object with an area defined by its azimuth as a catchment area, so that the sources Src22 and Src4 are assigned to the processing object according to one embodiment. Fig. 21 shows a processing object as in Fig. 20 , but with an additional transition area that can be controlled by the user via the "Feather" slider according to one embodiment. Fig. 22 shows several processing objects in the scene, with different catchment areas according to one embodiment. Fig. 23 shows the red square on the right side of the image shows a processing object for horizontally distorting the position of audio objects according to one embodiment. Fig. 24 shows the scene after the user has distorted the corners of the processing object. The position of all sources has changed according to the distortion according to one embodiment. Fig. 25 shows a possible visualization of the assignment of individual audio objects to a processing object according to one embodiment.

[0031] Fig. 1 shows an apparatus for generating a processed signal using a plurality of audio objects according to one embodiment, wherein each audio object of the plurality of audio objects comprises an audio object signal and audio object metadata, wherein the audio object metadata comprises a position of the audio object and a gain parameter of the audio object.

[0032] The apparatus comprises: an interface 110 for specifying at least one effect parameter of a processing object group of audio objects by a user, wherein the processing object group of audio objects comprises two or more audio objects of the plurality of audio objects.

[0033] Furthermore, the device comprises a processor unit 120 which is designed to generate the processed signal such that the at least one effect parameter which was specified by means of the interface 110 is applied to the audio object signal or to the audio object metadata of each of the audio objects of the processing object group of audio objects.

[0034] One or more audio objects of the majority of audio objects do not belong to the processing object group of audio objects.

[0035] The device described above of the Fig. 1 realizes an efficient form of audio mastering for audio objects.

[0036] The problem with audio objects is that an audio scene often contains a large number of audio objects. If these are to be modified, it would be a considerable effort to specify each audio object individually.

[0037] According to the invention, a group of two or more audio objects is organized into a group of audio objects, referred to as a processing object group. A processing object group is therefore a group of audio objects that are organized into this specific group, the processing object group.

[0038] According to the invention, a user now has the option of specifying one or more (at least one) effect parameters via interface 110. Processor unit 120 then ensures that the effect parameter is applied to all two or more audio objects of the processing object group by a single input of the effect parameter.

[0039] Such an application of the effect parameter can now consist, for example, in the effect parameter modifying a specific frequency range of the audio object signal of each of the audio objects of the processing object group.

[0040] Or, the gain parameter of the audio object metadata of each of the audio objects of the processing object group can be increased or decreased accordingly depending on the effect parameter, for example.

[0041] Or, the position of the audio object metadata of each of the audio objects in the processing object group can be changed accordingly, for example, depending on the effect parameter. For example, it is conceivable that all audio objects in the processing object group are shifted by +2 along the x-coordinate axis, -3 along the y-coordinate axis, and +4 along the z-coordinate axis.

[0042] It is also conceivable that applying an effect parameter to the audio objects in the processing object group could have a different effect on each audio object in the processing object group. For example, an axis could be defined as an effect parameter, along which the position of all audio objects in the processing object group is mirrored. Changing the position of the audio objects in the processing object group would then have a different effect on each audio object in the processing object group.

[0043] In one embodiment, the processor unit 120 may, for example, be configured not to apply the at least one effect parameter specified by means of the interface to any audio object signal and any audio object metadata of the one or more audio objects that do not belong to the processing object group of audio objects.

[0044] For such an embodiment, it is specified that the effect parameter is not applied to audio objects that do not belong to the processing object group.

[0045] In principle, audio object mastering can either be performed centrally on the encoder side. Or, on the decoder side, the end user, as the receiver of the audio object scenery, can modify the audio objects themselves according to the invention.

[0046] An embodiment that implements audio object mastering on the encoder side according to the invention is described in Fig. 2 shown.

[0047] An embodiment that implements audio object mastering on the decoder side according to the invention is described in Fig. 3 shown.

[0048] Fig. 2 shows a device according to another embodiment, wherein the device is an encoder.

[0049] In Fig. 2 The processor unit 120 is configured to generate a downmix signal using the audio object signals of the plurality of audio objects. The processor unit 120 is configured to generate a metadata signal using the audio object metadata of the plurality of audio objects.

[0050] Furthermore, the processor unit 120 is in Fig. 2 designed to generate the downmix signal as the processed signal, wherein at least one modified object signal for each audio object of the processing object group of audio objects is mixed into the downmix signal, wherein the processor unit 120 is designed to generate the modified object signal of this audio object for each audio object of the processing object group of audio objects by applying the at least one effect parameter, which was specified by means of the interface 110, to the audio object signal of this audio object.

[0051] Or, the processor unit 120 of the Fig. 2 is designed to generate the metadata signal as the processed signal, wherein the metadata signal comprises at least one modified position for each audio object of the processing object group of audio objects, wherein the processor unit 120 is designed to generate, for each audio object of the processing object group of audio objects, the modified position of this audio object by applying the at least one effect parameter, which was specified by means of the interface 110, to the position of this audio object.

[0052] Or, the processor unit 120 of the Fig. 2 is designed to generate the metadata signal as the processed signal, wherein the metadata signal comprises at least one modified gain parameter for each audio object of the processing object group of audio objects, wherein the processor unit 120 is designed to generate, for each audio object of the processing object group of audio objects, the modified gain parameter of this audio object by applying the at least one effect parameter, which was specified by means of the interface 110, to the gain parameter of this audio object.

[0053] Fig. 3 shows a device according to a further embodiment, wherein the device is a decoder. The device of Fig. 3 is designed to receive a downmix signal in which the plurality of audio object signals of the plurality of audio objects are mixed. Furthermore, the device of Fig. 3 designed to receive a metadata signal, wherein the metadata signal for each audio object of the plurality of audio objects comprises the audio object metadata of that audio object.

[0054] The processor unit 120 of the Fig. 3 is configured to reconstruct the plurality of audio object signals of the plurality of audio objects based on a downmix signal.

[0055] Furthermore, the processor unit 120 of the Fig. 3 designed to generate an audio output signal comprising one or more audio output channels as the processed signal.

[0056] Furthermore, the processor unit 120 of the Fig. 3 designed to apply the at least one effect parameter specified by means of the interface 110 to the audio object signal of each of the audio objects of the processing object group of audio objects in order to generate the processed signal, or to apply the at least one effect parameter specified by means of the interface 110 to the position or to the gain parameter of the audio object metadata of each of the audio objects of the processing object group of audio objects in order to generate the processed signal.

[0057] In audio object decoding, rendering on the decoder side is well known to those skilled in the art, for example from the SAOC standard (Spatial Audio Object Coding), see [8].

[0058] On the decoder side, one or more rendering parameters can be specified, for example, by user input via interface 110.

[0059] Thus, in one embodiment, the interface 110 of the Fig. 3 For example, it may further be configured for the user to specify one or more rendering parameters. The processor unit 120 of the Fig. 3 For example, it may be configured to generate the processed signal using the one or more rendering parameters depending on the position of each audio object of the processing object group of audio objects.

[0060] Fig. 4 shows a system according to one embodiment comprising an encoder 200 and a decoder 300.

[0061] The encoder 200 of the Fig. 4 is designed to generate a downmix signal based on audio object signals of a plurality of audio objects and to generate a metadata signal based on audio object metadata of the plurality of audio objects, wherein the audio object metadata comprises a position of the audio object and a gain parameter of the audio object.

[0062] The decoder 400 of the Fig. 4 is designed to generate an audio output signal comprising one or more audio output channels based on the downmix signal and based on the metadata signal.

[0063] The encoder 200 of the system of Fig. 4 a device according to Fig. 2 be.

[0064] Or, the decoder 300 of the system of Fig. 4 is a device according to Fig. 3 be.

[0065] Or, the encoder 200 of the system of Fig. 4 a device according to Fig. 2 and the decoder 300 of the system of Fig. 4 can be a device of Fig. 3 be.

[0066] The following embodiments are equally applicable in a device of Fig. 1 and in a device of Fig. 2 and in a device of Fig. 3 They can also be implemented in an encoder 200 of the system of Fig. 4 realizable, as well as in a decoder 300 of the system of Fig. 4 .

[0067] According to one embodiment, the processor unit 120 can, for example, be configured to generate the processed signal such that the at least one effect parameter specified by means of the interface 110 is applied to the audio object signal of each of the audio objects in the processing object group of audio objects. The processor unit 120 can, for example, be configured not to apply the at least one effect parameter specified by means of the interface to any audio object signal of the one or more audio objects of the plurality of audio objects that do not belong to the processing object group of audio objects.

[0068] Such an application of the effect parameter can now, for example, consist in the application of the effect parameter to the audio object signal of each audio object of the processing object group, e.g. modifying a specific frequency range of the audio object signal of each of the audio objects of the processing object group.

[0069] In one embodiment, the processor unit 120 can, for example, be configured to generate the processed signal such that the at least one effect parameter specified via the interface 110 is applied to the gain parameter of the metadata of each of the audio objects in the processing object group of audio objects. In this case, the processor unit 120 can, for example, be configured not to apply the at least one effect parameter specified via the interface to any gain parameter of the audio object metadata of the one or more audio objects of the plurality of audio objects that do not belong to the processing object group of audio objects.

[0070] As already described, in such an embodiment, the gain parameter of the audio object metadata of each of the audio objects of the processing object group can be increased (e.g., increased by +3 dB) or decreased accordingly depending on the effect parameter, for example.

[0071] According to one embodiment, the processor unit 120 can, for example, be configured to generate the processed signal such that the at least one effect parameter specified by means of the interface 110 is applied to the position of the metadata of each of the audio objects in the processing object group of audio objects. The processor unit 120 can, for example, be configured not to apply the at least one effect parameter specified by means of the interface to any position of the audio object metadata of the one or more audio objects of the plurality of audio objects that do not belong to the processing object group of audio objects.

[0072] As already described, in such an embodiment, the position of the audio object metadata of each of the audio objects in the processing object group can be changed accordingly, for example, depending on the effect parameter. This can be done, for example, by specifying the corresponding x, y, and z coordinate values ​​by which the position of each of the audio objects is to be shifted. Or, for example, a shift by a certain angle, rotated around a defined center point, for example, around a user position, can be specified. Or, for example, a doubling (or, for example, halving) of the distance to a certain point can be provided as an effect parameter for the position of each audio object in the processing object group.

[0073] The interface 110 is configured for the user to specify at least one definition parameter of the processing object group of audio objects. The processor unit 120 is configured to determine, depending on the at least one definition parameter of the processing object group of audio objects specified via the interface 110, which audio objects of the plurality of audio objects belong to the processing object group of audio objects.

[0074] Thus, according to one embodiment, the at least one definition parameter of the processing object group of audio objects can, for example, comprise at least one position of a region of interest (where the position of the region of interest is, for example, the center or center of gravity of the region of interest). The region of interest can be assigned to the processing object group of audio objects. The processor unit 120 can, for example, be configured to determine for each audio object of the plurality of audio objects, depending on the position of the audio object metadata of this audio object and depending on the position of the region of interest, whether this audio object belongs to the processing object group of audio objects.

[0075] In one embodiment, the at least one definition parameter of the processing object group of audio objects can, for example, further comprise a radius of the region of interest assigned to the processing object group of audio objects. Processor unit 120 can, for example, be configured to decide for each audio object of the plurality of audio objects whether this audio object belongs to the processing object group of audio objects, depending on the position of the audio object metadata of this audio object and depending on the position of the region of interest and depending on the radius of the region of interest.

[0076] For example, a user can specify a position of the processing object group and a radius of the processing object group. The position of the processing object group can specify a spatial center point, and the radius of the processing object group, together with the center point of the processing object group, then defines a circle. All audio objects with a position within the circle or on the circle line can then be defined as audio objects of this processing object group; all audio objects with a position outside the circle are then not included in the processing object group. The area within the circle line and on the circle line can then be understood as an "area of ​​interest."

[0077] According to one embodiment, the processor unit 120 can, for example, be configured to determine a weighting factor for each of the audio objects in the processing object group of audio objects as a function of a distance between the position of the audio object metadata of this audio object and the position of the region of interest. For example, the processor unit 120 can be configured to apply the weighting factor of this audio object, together with the at least one effect parameter specified via the interface 110, to the audio object signal or to the gain parameter of the audio object metadata of this audio object for each of the audio objects in the processing object group of audio objects.

[0078] In such an embodiment, the influence of the effect parameter on the individual audio objects of the processing object group is individualized for each audio object by determining, in addition to the effect parameter, an individual weighting factor for each audio object, which is applied to the audio object.

[0079] The at least one definition parameter of the processing object group of audio objects is configured to include at least one angle that specifies a direction from a defined user position in which a region of interest is located that is assigned to the processing object group of audio objects. The processor unit 120 is configured to determine for each audio object of the plurality of audio objects, depending on the position of the metadata of this audio object and depending on the angle that specifies the direction from the defined user position in which the region of interest is located, whether this audio object belongs to the processing object group of audio objects.

[0080] According to one embodiment, the processor unit 120 can, for example, be configured to determine a weighting factor for each of the audio objects in the processing object group of audio objects, which weighting factor depends on a difference between a first angle and a further angle, wherein the first angle is the angle that specifies the direction from the defined user position in which the region of interest is located, and wherein the further angle depends on the defined user position and on the position of the metadata of this audio object. For example, the processor unit 120 can be configured to apply the weighting factor of this audio object, together with the at least one effect parameter specified by means of the interface 110, to the audio object signal or to the gain parameter of the audio object metadata of this audio object for each of the audio objects in the processing object group of audio objects.

[0081] In one embodiment, the processing object group of audio objects may, for example, be a first processing object group of audio objects, wherein, for example, one or more further processing object groups of audio objects may also exist.

[0082] In this case, each processing object group of the one or more further processing object groups of audio objects can comprise one or more audio objects of the plurality of audio objects, wherein at least one audio object of a processing object group of the one or more further processing object groups of audio objects is not an audio object of the first processing object group of audio objects.

[0083] In this case, the interface 110 can be designed for each processing object group of the one or more further processing object groups of audio objects to allow the user to specify at least one further effect parameter for this processing object group of audio objects.

[0084] In this case, the processor unit 120 can be designed to generate the processed signal such that, for each processing object group of the one or more further processing object groups of audio objects, the at least one further effect parameter of this processing object group, which was specified by means of the interface 110, is applied to the audio object signal or to the audio object metadata of each of the one or more audio objects of this processing object group, wherein one or more audio objects of the plurality of audio objects do not belong to this processing object group.

[0085] In this case, the processor unit 120 can, for example, be designed not to apply the at least one further effect parameter of this processing object group, which was specified by means of the interface, to any audio object signal and any audio object metadata of the one or more audio objects that do not belong to this processing object group.

[0086] In such embodiments, more than one processing object group may exist. One or more separate effect parameters are defined for each processing object group.

[0087] According to one embodiment, the interface 110 can be configured, in addition to the first processing object group of audio objects, for example, for the user to specify the one or more further processing object groups of one or more audio objects, in that the interface 110 is configured for each processing object group of the one or more further processing object groups of one or more audio objects to specify at least one definition parameter of this processing object group by the user.

[0088] In this case, the processor unit 120 can, for example, be designed to determine for each processing object group of the one or more further processing object groups of one or more audio objects, depending on the at least one definition parameter of this processing object group, which was specified by means of the interface 110, which audio objects of the plurality of audio objects belong to this processing object group.

[0089] In the following, concepts of embodiments of the invention and preferred embodiments are presented.

[0090] In embodiments, any kind of global adjustments in OBAP are made possible by converting global adjustments into individual changes of the affected audio objects (e.g., by the processor unit 120).

[0091] Spatial mastering for object-based audio production can be realized, for example, as follows by implementing processing objects according to the invention.

[0092] The proposed implementation of overall adjustments is realized via processing objects (POs). These can be positioned anywhere in a scene, just like ordinary audio objects, and freely in real time. The user can apply any signal processing to the processing object (to the processing object group), for example, equalizer (EQ) or Kompression. For each of these processing tools, the parameter settings of the processing object can be converted into object-specific settings. Various methods are presented for this calculation.

[0093] An area of ​​interest is considered below.

[0094] Fig. 5 shows a processing object with the range A and the fading area A f according to one embodiment.

[0095] As in Fig. 5 As shown, the user defines an area A and a blanking area A f around the processing object. The processing parameters of the processing object are divided into constant parameters and weighted parameters. Values ​​of constant parameters are passed on unchanged by all audio objects within A and A f inherited. Weighted parameter values ​​are only inherited by audio objects within A. Audio objects within A f are weighted with a distance factor. The decision as to which parameters are weighted and which are not depends on the parameter type.

[0096] The custom value p M such a weighted parameter for the processing object, for each audio object S i , the parameter function p i defined as follows: p i t = p M t , p M t ∗ f i t , 0 , for S i ∈ A for S i ∈ A f else . , where the factor f i as follows: f i t = r A f − r S i r A f − r A .

[0097] Consequently, if the user r A = 0 specifies that there is no range of validity within which weighted parameters are kept constant.

[0098] Im Folgenden a calculation of inverse parameters according to one embodiment is described.

[0099] Fig. 6 shows a processing object with area A and object radii according to one embodiment.

[0100] User adjustments to the processing object converted via equation (1) may not always produce the desired results quickly enough, because the precise position of audio objects is not taken into account. For example, if the area around the processing object is very large and the audio objects it contains are far away from the processing object position, the effect of calculated adjustments may not even be audible at the processing object position.

[0101] For gain parameters, a different calculation method is conceivable based on the decay rate of each object. Again, within a user-defined range of interest, which is Fig. 6 is shown, the individual parameter p i for each audio object is then calculated as follows. p i t = h i t , 0 , for S i ∈ A else . , where h i could be defined as follows h i t = sgng e t ∗ g e t + 10 ∗ log 10 a i d i t 2 . a i is a constant for the closest possible distance to an audio object, and d i ( t ) is the distance from the audio object to the EQ object. Derived from the inverse distance law, the function has been modified to correctly handle possible positive or negative EQ gain changes.

[0102] In the following modified embodiment, an angle-based calculation is carried out.

[0103] The previous calculations are based on the distance between audio objects and the processing object. However, from a user perspective, the angle between the processing object and the surrounding audio objects can sometimes represent their auditory impression more accurately. [5] proposes the global control of any audio plug-in parameter via the azimuth of audio objects. This approach can be adopted by using the difference in angle α i between the processing object with offset angle α eq and audio objects S i in whose vicinity is calculated, as in Fig. 7 is shown.

[0104] This shows Fig. 7 a relative angle of audio objects to the processing object according to one embodiment.

[0105] The user-defined area of ​​interest mentioned above could be defined accordingly using the angles α A and α Af be changed, which in Fig. 8 is shown.

[0106] This shows Fig. 8 an equalizer object with a new radial circumference according to one embodiment.

[0107] Regarding the blanking area, A f , f i be redefined as follows: f i t = α A f − α S i α A f − α A .

[0108] Although for the modified approach presented above, the distance d i In this context, if the angle between the audio object and the EQ object could simply be interpreted as the angle between the audio object and the EQ object, this would no longer justify applying the inverse ratio law. Therefore, only the user-defined range is changed, while the gain calculation remains as before.

[0109] In one embodiment, equalization is implemented as an application.

[0110] Equalization can be considered the most important tool in mastering, as the frequency response of a mix is ​​the most critical factor for good translation across playback systems.

[0111] The proposed equalization implementation is realized using EQ objects. Since all other parameters are not distance-dependent, only the gain parameter is of particular interest.

[0112] In a further embodiment, dynamic control is implemented as an application.

[0113] In traditional mastering, dynamic compression is used to control dynamic variations in a mix over time. Depending on the compression settings, this changes the perceived density and transient response of a mix. In the case of fixed compression, the perceived change in density is referred to as 'glue,' while stronger compression settings can be used for pumping or side-chain effects on beat-heavy mixes.

[0114] With OBAP, the user could easily specify identical compression settings for multiple adjacent objects to achieve multi-channel compression. However, summed compression on groups of audio objects would not only be advantageous for time-critical workflows, but would also be more likely to fulfill the psychoacoustic impression of so-called "glued" signals.

[0115] Fig. 9 shows a signal flow of a compression of the signals from n sources according to one embodiment.

[0116] According to a further embodiment, scene transformation is implemented as an application.

[0117] In stereo mastering, mid / side processing is a commonly used technique for expanding or stabilizing the stereo image of a mix. For spatial audio mixes, a similar option can be useful if the mix was created in an acoustically critical environment with potentially asymmetric room or speaker characteristics. It could also provide new creative possibilities for the ME to enhance the effects of a mix.

[0118] Fig. 10 shows a scene transformation using a control panel M according to one embodiment. Specifically, Fig. 10 a schematic implementation using a distortion area with user-draggable edges C 1 toC 4 .

[0119] A two-dimensional transformation of a scene in the horizontal plane can be realized using a homography transformation matrix H that represents each audio object at position p i to a new position p' i depicts, see also [7]: H : = h 1 h 2 h 3 h 4 h 5 h 6 h 7 h 8 h 9 , p i ′ = Hp i .

[0120] When the user moves a control field M to M' using the four draggable corners c 1-4 distorted (see Figur 6 ), their 2D coordinates can x 1 − 4 y 1 − 4 for a linear system of equations (7) to find the coefficients of H to obtain [7]. x 1 y 1 1 0 0 0 − x 1 ′ x 1 − x 1 ′ y 1 0 0 0 x 1 y 1 1 − y 1 ′ x 1 − y 1 ′ y 1 x 2 y 2 1 0 0 0 − x 2 ′ x 2 − x 2 ′ y 2 0 0 0 x 2 y 2 1 − y 2 ′ x 2 − y 2 ′ y 2 x 3 y 3 1 0 0 0 − x 3 ′ x 3 − x 3 ′ y 3 0 0 0 x 3 y 3 1 − y 3 ′ x 3 − y 3 ′ y 3 x 4 y 4 1 0 0 0 − x 4 ′ x 4 − x 4 ′ y 4 0 0 0 x 4 y 4 1 − y 4 ′ x 4 − y 4 ′ y 4 ∗ h 1 h 2 h 3 h 4 h 5 h 6 h 7 h 8 = x 1 ′ y 1 ′ x 2 ′ y 2 ′ x 3 ′ y 3 ′ x 4 ′ y 4 ′

[0121] Since audio object positions can vary over time, the coordinate positions can be interpreted as time-dependent functions.

[0122] Some embodiments implement dynamic equalizers. Other embodiments implement multi-band compression.

[0123] Object-based sound adjustments are not limited to the introduced equalizer applications.

[0124] The above description is supplemented below by a more general description of embodiments.

[0125] Object-based three-dimensional audio production follows the approach of calculating and playing audio scenes in real time for virtually any speaker configuration using a rendering process. Audio scenes describe the arrangement of audio objects over time. Audio objects consist of audio signals and metadata. This metadata includes, among other things, position in space and volume. To edit a scene, the user previously had to change all audio objects in a scene individually.

[0126] When we refer to a processing object group and a processing object in the following, it should be noted that for each processing object, a processing object group is always defined that contains audio objects. The processing object group is also referred to, for example, as the container of the processing object. For each processing object, a group of audio objects from the plurality of audio objects is defined. The corresponding processing object group contains the specified group of audio objects. A processing object group is therefore a group of audio objects.

[0127] Processing objects can be defined as objects that can modify the properties of other audio objects. Processing objects are artificial containers to which any audio objects can be assigned, meaning that all of its assigned audio objects are addressed via the container. The assigned audio objects can be influenced by any number of effects. Thus, processing objects offer the user the ability to edit multiple audio objects simultaneously.

[0128] For example, a processing object has position, mapping methods, containers, weighting methods, audio signal processing effects, and metadata effects.

[0129] The position is a position of the processing object in a virtual scene.

[0130] The mapping procedure assigns audio objects to the processing object (possibly using their position).

[0131] The container (or connections) is the set of all audio objects assigned to the processing object (or any additional other processing objects).

[0132] Weighting methods are the algorithms for calculating the individual effect parameter values ​​for the assigned audio objects.

[0133] Audio signal processing effects change the audio component of audio objects (e.g. equalizer, dynamics).

[0134] Metadata effects change the metadata of audio objects and / or processing objects (e.g. position distortion).

[0135] Likewise, the processing object group can be assigned the above-described position, allocation method, container, weighting method, audio signal processing effects, and metadata effects. The audio objects of the processing object container are the audio objects of the processing object group.

[0136] Fig. 11 shows the context of a processing object that causes audio signal effects and metadata effects, according to one embodiment.

[0137] In the following, properties of processing objects are described according to specific embodiments: Processing objects can be placed arbitrarily in a scene by the user, the position can be set constant over time or time-dependent.

[0138] Processing objects can be assigned effects by the user that modify the audio signal and / or the metadata of audio objects. Examples of effects include equalizing the audio signal, editing the dynamics of the audio signal, or changing the position coordinates of audio objects.

[0139] Processing objects can be assigned any number of effects in any order.

[0140] Effects change the audio signal and / or the metadata of the assigned set of audio objects, either constantly over time or time-dependently.

[0141] Effects have parameters for controlling signal and / or metadata processing. These parameters are divided into constant and weighted parameters, either user-defined or fixed depending on the type.

[0142] The effects of a processing object are copied and applied to its associated audio objects. The values ​​of constant parameters are retained unchanged by each audio object. The values ​​of weighted parameters are calculated individually for each audio object using various weighting methods. The user can select a weighting method for each effect or enable or disable it for individual audio sources.

[0143] The weighting methods take into account individual metadata and / or signal characteristics of individual audio objects. This corresponds, for example, to the distance of an audio object from the processing object or the frequency spectrum of an audio object. The weighting methods can also consider the listener's listening position. Furthermore, the aforementioned properties of audio objects can be combined for the weighting methods to derive individual parameter values. For example, the sound levels of audio objects can be added together during dynamic processing to derive a volume change for each audio object individually.

[0144] Effect parameters can be set constant over time or time-dependent. The weighting methods take such temporal changes into account.

[0145] Weighting methods can also process information that the audio renderer analyzes from the scene.

[0146] The order in which effects are applied to the processing object corresponds to the sequence in which signals and / or metadata are processed for each audio object. This means that the data modified by a previous effect is used as the basis for the next effect's calculations. The first effect operates on the still-unchanged data of an audio object.

[0147] Individual effects can be deactivated. The calculated data from the previous effect, if one exists, is then passed to the effect following the deactivated effect.

[0148] A newly developed effect is the change in the position of audio objects using homography ("distortion effect"). This involves displaying a rectangle with individually movable corners at the position of the processing object. If the user moves a corner, a transformation matrix for this distortion is calculated from the previous state of the rectangle and the newly distorted state. The matrix is ​​then applied to all position coordinates of the audio objects assigned to the processing object, so that their position changes according to the distortion.

[0149] Effects that only change metadata can also be applied to other processing objects (including the "distortion effect").

[0150] Audio sources can be assigned to processing objects in various ways. Depending on the type of assignment, the number of assigned audio objects can also change over time. This change is taken into account in all calculations.

[0151] A catchment area can be defined around the position of processing objects.

[0152] All audio objects positioned within the catchment area form the assigned set of audio objects to which the effects of the processing object are applied.

[0153] The catchment area can be any body (three-dimensional) or any shape (two-dimensional) defined by the user.

[0154] The center of the catchment area can, but does not have to, correspond to the position of the processing object. This is determined by the user.

[0155] An audio object lies within a three-dimensional catchment area if its position lies within the three-dimensional body.

[0156] An audio object lies within a two-dimensional catchment area if its position projected onto the horizontal plane lies within the two-dimensional shape.

[0157] The catchment area can take on an unspecified all-encompassing size so that all audio objects of a scene are within the catchment area.

[0158] The catchment areas adapt to changes in the scene properties (e.g. scene scaling) if necessary.

[0159] Regardless of the catchment area, processing objects can be linked to any selection of audio objects in a scene.

[0160] The coupling can be defined by the user so that all selected audio objects form a set of audio objects to which the effects of the processing object are applied.

[0161] Alternatively, the user can define the coupling so that the processing object adjusts its position over time based on the position of the selected audio objects. This position adjustment can take the listener's listening position into account. The effects of the processing object do not necessarily have to be applied to the coupled audio objects.

[0162] The assignment can be performed automatically based on user-defined criteria. All audio objects in a scene are continuously examined for the defined criteria(s). If the criteria(s) are met, they are assigned to the processing object. The duration of the assignment can be limited to the time the criteria(s) are met, or transition periods can be defined. The transition periods determine how long one or more criteria must be continuously met by the audio object for it to be assigned to the processing object, or how long one or more criteria must be continuously violated for the assignment to the processing object to be canceled.

[0163] Processing objects can be deactivated by the user so that their properties are retained and continue to be displayed to the user, but audio objects are not influenced by the processing object.

[0164] Any number of properties of a processing object can be linked by the user with similar properties of any number of other processing objects. These properties include effect parameters. The user can select absolute or relative coupling. With constant coupling, the changed property value of a processing object is exactly adopted by all linked processing objects. With relative coupling, the value of the change is offset against the property values ​​of linked processing objects.

[0165] Processing objects can be duplicated. This creates a second processing object with identical properties to the original. The properties of the processing objects are then independent of each other.

[0166] Properties of processing objects can be permanently inherited, for example when copying, so that changes in the parent are automatically adopted by the children.

[0167] Fig. 12 shows the modification of audio objects and audio signals in response to user input according to one embodiment.

[0168] Another new application of processing objects is intelligent parameter calculation using scene analysis. The user defines effect parameters at a specific position using the processing object. The audio renderer performs a predictive scene analysis to detect which audio sources influence the position of the processing object. Effects are then applied to the selected audio sources, taking the scene analysis into account, in such a way that the user-defined effect settings are best achieved at the position of the processing object.

[0169] In the following, further embodiments of the invention, which are implemented by means of the Fig. 13 - Fig. 25 are visually represented.

[0170] This shows Fig. 13 Processing object PO 4 with rectangle M for the distortion of the corners C 1 , C 2 , C 3 and C 4 by the user. Fig. 13 schematically a possible distortion towards M' with the corners C 1 ', C 2 ', C 3 ' and C 4 ', as well as the corresponding effect on the sources S 1 , S 2 , S 3 and S 4 with their new positions S 1 ', S 2 ', S 3 ' and S 4 '.

[0171] Fig. 14 shows processing objects PO 1 and PO 2 with their respective overlapping two-dimensional catchment areas A and B, as well as the distances a S1, a S2 and a S3 or b S3 , b S4 and b S6 from the respective processing object to the sources S 1 , S 2 , S 3 , S 4 and S 6 assigned by the catchment areas.

[0172] Fig. 15 shows processing object PO 3 with rectangular, two-dimensional catchment area C and the angles between PO 3 and the associated sources S 1 , S 2 and S 3 for a possible weighting of parameters that takes the listening position of the listener into account.

[0173] The angles can be determined by the difference between the azimuth of the individual sources and the azimuth α po of PO 3.

[0174] Fig. 16 shows a possible schematic implementation of an equalizer effect applied to a processing object. Buttons like w next to each parameter can be used to activate the weighting for that parameter. m 1 , m 2 , and m 3 provide options for the weighting method for the weighted parameters mentioned.

[0175] Fig. 17 shows the processing object PO 5 with a three-dimensional catchment area D and the respective distances d S1, d S2 and d S3 to the sources S 1 , S 2 and S 3 assigned via the catchment area.

[0176] Fig. 18 shows a prototypical implementation of a processing object to which an equalizer has been applied. The turquoise object with the wave symbol on the right side of the image shows the processing object in the audio scene, which the user can move freely with the mouse. Within the turquoise, transparent, homogeneous area around the processing object, the equalizer parameters are applied unchanged to the audio objects Src1, Src2, and Src3, as defined on the left side of the image. Around the homogeneous circular area, the shading that fades into transparency shows the area in which all parameters except the gain parameters are adopted unchanged from the sources. The gain parameters of the equalizer, on the other hand, are weighted depending on the distance of the sources from the processing object. Since only source Src4 and source Src24 are located in this area, only their parameters are weighted in this case. Source Src22 is not affected by the processing object.The "Area" slider controls the size of the radius of the circular area surrounding the processing object. The "Feather" slider controls the size of the radius of the surrounding transition area.

[0177] Fig. 19 shows a processing object as in Fig. 18 , just at a different position and without a transition area. All equalizer parameters are applied unchanged to sources Src22 and Src4. Sources Src3, Src2, Src1, and Src24 are not affected by the processing object.

[0178] Fig. 20 shows a processing object with an area defined by its azimuth as the catchment area, so that the sources Src22 and Src4 are assigned to the processing object. The tip of the catchment area in the center of the right side of the image corresponds to the position of the listener / user. When the processing object is moved, the area moves according to the azimuth. The "Area" slider allows the user to determine the size of the angle of the catchment area. The user can change from a circular to an angle-based catchment area using the lower selection field above the "Area" / "Feather" sliders, which now displays "radius."

[0179] Fig. 21 shows a processing object as in Fig. 20 , but with an additional transition area that can be controlled by the user via the "Feather" slider.

[0180] Fig. 22 Shows multiple processing objects in the scene, each with a different range. The gray processing objects have been deactivated by the user, meaning they do not affect the audio objects within their range. The left side of the image always displays the equalizer parameters of the currently selected processing object. The selection is indicated by a thin, light turquoise line around the object.

[0181] Fig. 23 The red square on the right side of the image shows a processing object for horizontally distorting the position of audio objects. The user can drag the corners in any direction with the mouse to distort the scene.

[0182] Fig. 24 shows the scene after the user has distorted the corners of the processing object. The position of all sources has changed according to the distortion.

[0183] Fig. 25 shows a possible visualization of the assignment of individual audio objects to a processing object.

[0184] Although some aspects have been described in the context of a device, it should be understood that these aspects also represent a description of the corresponding method, so that a block or component of a device can also be understood as a corresponding method step or as a feature of a method step. Analogously, aspects described in the context of or as a method step also represent a description of a corresponding block, detail, or feature of a corresponding device. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, some or more of the key method steps may be performed by such an apparatus.

[0185] Depending on specific implementation requirements, embodiments of the invention may be implemented in hardware or in software, or at least partially in hardware or at least partially in software. The implementation may be carried out using a digital storage medium, for example a floppy disk, a DVD, a Blu-ray disc, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a FLASH memory, a hard disk, or other magnetic or optical storage device on which electronically readable control signals are stored that can interact or interact with a programmable computer system such that the respective method is carried out. Therefore, the digital storage medium may be computer-readable.

[0186] Some embodiments according to the invention thus comprise a data carrier having electronically readable control signals capable of interacting with a programmable computer system such that one of the methods described herein is carried out.

[0187] In general, embodiments of the present invention may be implemented as a computer program product having a program code, wherein the program code is effective to perform one of the methods when the computer program product is run on a computer.

[0188] The program code can, for example, also be stored on a machine-readable medium.

[0189] Other embodiments include the computer program for performing one of the methods described herein, wherein the computer program is stored on a machine-readable medium. In other words, one embodiment of the method according to the invention is thus a computer program that has program code for performing one of the methods described herein when the computer program is executed on a computer.

[0190] A further embodiment of the method according to the invention is thus a data carrier (or a digital storage medium or a computer-readable medium) on which the computer program for performing one of the methods described herein is recorded. The data carrier or the digital storage medium or the computer-readable medium is typically tangible and / or non-transitory.

[0191] A further embodiment of the method according to the invention is thus a data stream or a sequence of signals that represents the computer program for carrying out one of the methods described herein. The data stream or the sequence of signals can be configured, for example, to be transferred via a data communication connection, for example, via the Internet.

[0192] A further embodiment comprises a processing device, for example a computer or a programmable logic device, which is configured or adapted to carry out one of the methods described herein.

[0193] A further embodiment comprises a computer on which the computer program for performing one of the methods described herein is installed.

[0194] A further embodiment according to the invention comprises a device or system designed to transmit a computer program for performing at least one of the methods described herein to a recipient. The transmission can be electronic or optical, for example. The recipient can be, for example, a computer, a mobile device, a storage device, or a similar device. The device or system can, for example, comprise a file server for transmitting the computer program to the recipient.

[0195] In some embodiments, a programmable logic device (e.g., a field-programmable gate array, an FPGA) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field-programmable gate array may interact with a microprocessor to perform any of the methods described herein. In general, in some embodiments, the methods are performed by any hardware device. This may be general-purpose hardware such as a computer processor (CPU) or method-specific hardware such as an ASIC.

[0196] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. Therefore, it is intended that the invention be limited only by the scope of the following claims and not by the specific details presented in the description and explanation of the embodiments herein. Referenzen

[0197] [1] Coleman, P., Franck, A., Francombe, J., Liu, Q., Campos, T. D., Hughes, R., Menzies, D., Galvez, M. S., Tang, Y., Woodcock, J., Jackson, P., Melchior, F., Pike, C., Fazi, F., Cox, T., and Hilton, A., "An Audio-Visual System for Object-Based Audio: From Recording to Listening," IEEE Transactions on Multimedia, PP(99), pp. 1-1, 2018, ISSN 1520- 9210, doi:10.1109 / TMM.2018.2794780. [2] Gasull Ruiz, A., Sladeczek, C., and Sporer, T., "A Description of an Object-Based Audio Workflow for Media Productions," in Audio Engineering Society Conference: 57th International Conference: The Future of Audio Entertainment Technology, Cinema, Television and the Internet, 2015. [3] Melchior, F., Michaelis, U., and Steffens, R., "Spatial Mastering - a new concept for spatial sound design in object-based audio scenes," in Proceedings of the International Computer Music Conference 2011, 2011. [4] Katz, B. and Katz, R. A., Mastering Audio: The Art and the Science, Butterworth-Heinemann, Newton, MA, USA, 2003, ISBN 0240805453. AES Conference on Spatial Reproduction, Tokyo, Japan, 2018 August 6 - 9, Page 2 [5] Melchior, F., Michaelis, U., and Steffens, R., "Spatial Mastering - A New Concept for Spatial Sound Design in Object-based Audio Scenes," Proceedings of the International Computer Music Conference 2011, University of Huddersfield, UK, 2011. [6] Sladeczek, C., Neidhardt, A., Böhme, M., Seeber, M., and Ruiz, A. G., "An Approach for Fast and Intuitive Monitoring of Microphone Signals Using a Virtual Listener," Proceedings, International Conference on Spatial Audio (ICSA), 21.2. - 23.2.2014, Erlangen, 2014 [7] Dubrofsky, E., Homography Estimation, Master's thesis, University of British Columbia, 2009. [8] ISO / IEC 23003-2:2010 Information technology - MPEG audio technologies - Part 2: Spatial Audio Object Coding (SAOC); 2010.

Claims

1. A device for generating a processed signal using a plurality of audio objects, each audio object of the plurality of audio objects including an audio object signal and audio object metadata, the audio object metadata including a position of the audio object and a gain parameter of the audio object, the device comprising: an interface (110) for specification of at least one effect parameter of a processing-object group of audio objects by a user, the processing-object group of audio objects including two or more audio objects of the plurality of audio objects, and a processor unit (120) configured to generate the processed signal such that the at least one effect parameter specified by means of the interface (110) is applied to the audio object signal or to the audio object metadata of each of the audio objects of the processing-object group of audio objects, wherein the interface (110) is configured to specify at least one definition parameter of the processing-object group of audio objects by the user, wherein the processor unit (120) is configured to determine, in dependence on the at least one definition parameter of the processing-object group of audio objects specified by means of the interface (110), which audio objects of the plurality of audio objects belong to the processing-object group of audio objects, wherein the at least one definition parameter of the processing-object group of audio objects includes at least one angle specifying a direction from a defined user position in which an area of interest associated with the processing-object group of audio objects is located, and wherein the processor unit (120) is configured to determine, for each audio object of the plurality of audio objects, depending on the position of the metadata of this audio object and depending on the angle specifying the direction from the defined user position in which the area of interest is located, whether this audio object belongs to the processing-object group of audio objects.

2. The device as claimed in claim 1, wherein one or more audio objects of the plurality of audio objects do not belong to the processing-object group of audio objects, and wherein the processor unit (120) is configured not to apply the at least one effect parameter specified by means of the interface (110) to any audio object signal and any audio object metadata of the one or more audio objects which do not belong to the processing-object group of audio objects.

3. The device as claimed in claim 2, wherein the processor unit (120) is configured to generate the processed signal such that the at least one effect parameter specified by means of the interface (110) is applied to the audio object signal of each of the audio objects of the processing-object group of audio objects; and / or wherein the processor unit (120) is configured to generate the processed signal such that the at least one effect parameter specified by means of the interface (110) is applied to the gain parameter of the metadata of each of the audio objects of the processing-object group of audio objects; wherein the processor unit (120) is configured not to apply the at least one effect parameter specified by means of the interface (110) to any gain parameter of the audio object metadata of the one or more audio objects of the plurality of audio objects which do not belong to the processing-object group of audio objects.

4. The device as claimed in claim 2 or 3, wherein the processor unit (120) is configured to generate the processed signal such that the at least one effect parameter specified by means of the interface (110) is applied to the position of the metadata of each of the audio objects of the processing-object group of audio objects, wherein the processor unit (120) is configured not to apply the at least one effect parameter specified by means of the interface (110) to any position of the audio object metadata of the one or more audio objects of the plurality of audio objects which do not belong to the processing-object group of audio objects.

5. The device as claimed in any of claims 1 to 4, wherein the at least one definition parameter of the processing-object group of audio objects includes the at least one angle, wherein the processor unit (120) is configured to determine, for each of the audio objects of the processing-object group of audio objects, a weighting factor which depends on a difference of a first angle and a further angle, wherein the first angle is the angle specifying the direction from the defined user position in which the area of interest is located, and wherein the further angle depends on the defined user position and on the position of the metadata of this audio object, wherein the processor unit (120) is configured to apply, for each of the audio objects of the processing-object group of audio objects, the weighting factor of this audio object together with the at least one effect parameter specified by means of the interface (110) to the audio object signal or to the gain parameter of the audio object metadata of this audio object.

6. The device as claimed in any of the previous claims, wherein the processing-object group of audio objects is a first processing-object group of audio objects, wherein there also exist one or more further processing-object groups of audio objects, wherein each processing-object group of the one or more further processing-object groups of audio objects includes one or more audio objects of the plurality of audio objects, wherein at least one audio object of a processing-object group of the one or more further processing-object groups of audio objects is not an audio object of the first processing-object group of audio objects; wherein the interface (110) is configured for specification, for each processing-object group of the one or more further processing-object groups of audio objects, of at least one further effect parameter for that processing-object group of audio objects by the user; and wherein the processor unit (120) is configured to generate the processed signal such that for each processing-object group of the one or more further processing-object groups of audio objects, the at least one further effect parameter of this processing-object group specified by means of the interface (110) is applied to the audio object signal or to the audio object metadata of each of the one or more audio objects of this processing-object group, wherein one or more audio objects of the plurality of audio objects do not belong to this processing-object group, and wherein the processor unit (120) is configured not to apply the at least one further effect parameter of this processing-object group specified by means of the interface (110) to any audio object signal and any audio object metadata of the one or more audio objects which do not belong to this processing-object group; and / or wherein the interface (110) is configured, in addition to the first processing-object group of audio objects, for specification of the one or more further processing-object groups of one or more audio objects by the user, in that the interface (110) is configured, for each processing-object group of the one or more further processing-object groups of one or more audio objects, for specification of at least one definition parameter of this processing-object group by the user, and wherein the processor unit (120) is configured to determine, for each processing-object group of the one or more further processing-object groups of one or more audio objects, in dependence on the at least one definition parameter of this processing-object group specified by means of the interface (110), which audio objects of the plurality of audio objects belong to this processing-object group.

7. The device as claimed in any of the previous claims, the device being an encoder, wherein the processor unit (120) is configured to generate a downmix signal while using the audio object signals of the plurality of audio objects, and wherein the processor unit (120) is configured to generate a metadata signal signal while using the audio object metadata of the plurality of audio objects, wherein the processor unit (120) is configured to generate the downmix signal as the processed signal, at least one modified object signal being mixed, in the downmix signal, for each audio object of the processing-object group of audio objects, the processor unit (120) being configured to generate, for each audio object of the processing-object group of audio objects, the modified object signal of this audio object by means of applying the at least one effect parameter specified by means of the interface (110) to the audio object signal of this audio object, or wherein the processor unit (120) is configured to generate the metadata signal as the processed signal, the metadata signal including at least one modified position for each audio object of the processing-object group of audio objects, wherein the processor unit (120) is configured to generate, for each audio object of the processing-object group of audio objects, the modified position of this audio object by means of applying the at least one effect parameter specified by means of the interface (110) to the position of this audio object, or the processor unit (120) is configured to generate the metadata signal as the processed signal, the metadata signal including at least one modified gain parameter for each audio object of the processing-object group of audio objects, the processor unit (120) being configured to generate, for each audio object of the processing-object group of audio objects, the modified gain parameter of this audio object by applying the at least one effect parameter specified by means of the interface (110) to the gain parameter of this audio object.

8. The device as claimed in any of claims 1 to 6, the device being a decoder, the device being configured to receive a downmix signal in which the plurality of audio object signals of the plurality of audio objects are mixed, the device further being configured to receive a metadata signal, the metadata signal including, for each audio object of the plurality of audio objects, the audio object metadata of this audio object, wherein the processor unit (120) is configured to reconstruct the plurality of audio object signals of the plurality of audio objects on the basis of a downmix signal, wherein the processor unit (120) is configured to generate, as the processed signal, an audio output signal comprising one or more audio output channels, wherein the processor unit (120) is configured to apply the at least one effect parameter specified by means of the interface (110) to the audio object signal of each of the audio objects of the processing-object group of audio objects to generate the processed signal, or to apply the at least one effect parameter specified by means of the interface (110) to the position or to the gain parameter of the audio object metadata of each of the audio objects of the processing-object group of audio objects to generate the processed signal.

9. The device as claimed in claim 8, wherein the interface (110) is further configured for specification of one or more rendering parameters by the user, and wherein the processor unit (120) is configured to generate the processed signal while using the one or more rendering parameters in dependence on the position of each audio object of the processing-object group of audio objects.

10. A system comprising: an encoder (200) for generating a downmix signal on the basis of audio object signals of a plurality of audio objects and for generating a metadata signal on the basis of audio object metadata of the plurality of audio objects, wherein the audio object metadata includes a position of the audio object and a gain parameter of the audio object, and a decoder (300) for generating an audio output signal including one or more audio output channels on the basis of the downmix signal and on the basis of the metadata signal, wherein the encoder (200) is a device as claimed in claim 7, or wherein the decoder (300) is a device as claimed in claim 8 or 9, or wherein the encoder (200) is a device as claimed in claim 7 and the decoder (300) is a device as claimed in claim 8 or 9.

11. A method of generating a processed signal while using a plurality of audio objects, each audio object of the plurality of audio objects including an audio object signal and audio object metadata, the audio object metadata including a position of the audio object and a gain parameter of the audio object, the method including: specifying at least one effect parameter of a processing-object group of audio objects on the part of / by a user by means of an interface (110), wherein the processing-object group of audio objects comprises two or more audio objects of the plurality of audio objects, and generating the processed signal by a processor unit (120) such that the at least one effect parameter specified by means of the interface (110) is applied to the audio object signal or to the audio object metadata of each of the audio objects of the processing-object group of audio objects, wherein the method further comprises: specifying at least one definition parameter of the processing-object group of audio objects by the user, determining, in dependence on the at least one definition parameter of the processing-object group of audio objects specified by means of the interface (110), which audio objects of the plurality of audio objects belong to the processing-object group of audio objects, wherein the at least one definition parameter of the processing-object group of audio objects comprises at least one angle specifying a direction from a defined user position in which an area of interest associated with the processing-object group of audio objects is located, and wherein it is determined, for each audio object of the plurality of audio objects, depending on the position of the metadata of this audio object and depending on the angle specifying the direction from the defined user position in which the area of interest is located, whether this audio object belongs to the processing-object group of audio objects.

12. A computer program comprising a program code for performing the method as claimed in claim 11.

Citation Information

Patent Citations

  • Device for changing an audio scene and device for generating a directional function

    DE102010030534A1

  • Playback Device For Generating Sound Events

    US20100223552A1

  • System and method for adaptive audio signal generation, coding and rendering

    WO2013006338A2