Apparatus and method for screen-related audio object remapping - Patents.com

The proposed audio object remapping solution addresses the challenge of maintaining audio-visual consistency across different screen sizes by using an object metadata processor to calculate new audio object positions based on screen size changes, effectively enhancing multimedia playback experiences.

JP7675506B2Active Publication Date: 2025-05-13FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2020118271
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2014-12-08
Filing Date
2020-07-09
Publication Date
2025-05-13
Estimated Expiration
2035-03-25

AI Technical Summary

Technical Problem

Existing audio signal processing technologies struggle to effectively remap audio objects in relation to changes in screen size, particularly in multimedia playback settings, which can lead to inconsistencies in audio-visual synchronization.

Method used

An apparatus and method for audio object remapping, which includes an object metadata processor and an object renderer. The processor receives metadata indicating whether an audio object is screen-related and calculates its new position based on the original position and screen size, allowing for accurate remapping of audio objects in response to changes in screen size.

Benefits of technology

The solution enables improved integration of audio and visual multimedia content by ensuring that audio objects are optimally rendered across varying screen sizes, maintaining audio-visual consistency and enhancing the overall multimedia experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007675506000057
    Figure 0007675506000057
  • Figure 0007675506000058
    Figure 0007675506000058
  • Figure 0007675506000059
    Figure 0007675506000059
Patent Text Reader

Abstract

To provide a device for generating speaker signals for screen-related audio object mapping.SOLUTION: A device includes an object metadata processor 110 and an object renderer 120. The object renderer 120 receives an audio object. The object metadata processor 110 receives metadata that includes instructions related to whether the audio object is screen-related and further includes a first position of the audio object. The object metadata processor 110 calculates a second position of the audio object according to the first position of the audio object and the size of a screen when the audio object is instructed as being screen-related in the metadata. The object renderer 120 generates a speaker signal according to the audio object and positional information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to audio signal processing, in particular to an apparatus and method for audio object remapping, and more particularly to an apparatus and method for screen-related audio object remapping. [Background technology]

[0002] With the increasing consumption of multimedia content in daily life, the demand for sophisticated multimedia solutions is constantly growing. In this aspect, the integration of visual and audio content plays a key role. It is desirable to optimally adjust visual and audio multimedia content to the available visual and audio playback settings.

[0003] In the current state of the art, audio objects are already known. An audio object can for example be thought of as a sound track with associated metadata. The metadata can for example describe properties of the raw audio data, for example the desired playback position or volume level. The advantage of object-based audio is that a specific rendering process on the playback side can play back a given movement in the best possible way for all playback speaker layouts.

[0004] The geometric metadata can be used to define where an audio object should be rendered, e.g., azimuth or elevation angle or absolute position relative to a reference point, e.g., the listener. The metadata is stored or transmitted together with the object audio signal.

[0005] In the context of MPEG-H at the 105th MPEG meeting, the audio group considered the requirements and timelines of several different application standards. According to their considerations, it is essential to align specific points in time with specific requirements for next generation broadcast systems. According to their considerations, the system should be able to accept audio objects at the encoder input. Moreover, the system should support the signaling, delivery and rendering of audio objects and allow the user to control the objects, e.g. for dialogue enhancement, alternative language tracks and audio description languages.

[0006] In the current state of the art, different concepts are provided. According to a first prior art presented in "Method and apparatus for playback of a higher-order ambisonics audio signal" (see [1]), the playback of a spatial sound field oriented audio for its linked visual objects is adapted by applying a spatial warping process. In that prior art, the decoder warps the sound field in such a way that all audio objects in the direction of the screen are compressed or stretched according to the size ratio between the target screen and the reference screen. The possibility of encoding and transmitting as metadata together with the content the reference size of the screen used for content playback (or the viewing angle from the reference listening position) is included. Alternatively, a fixed reference screen size is assumed in the encoding and for the decoding, and the decoder knows the actual size of the target screen. In this prior art, the decoder warps the sound field in such a way that all audio objects in the direction of the screen are compressed or stretched according to the ratio between the size of the target screen and the size of the reference screen. A so-called "two-piece piecewise linear" warping function is used. The stretching is limited to the angular position of the audio items. In that prior art, for a centered screen, the definition of the warping function is similar to the definition of the mapping function of screen-related remapping. The first and third pieces of the three-piece piecewise linear mapping function can be defined as two-piece piecewise linear functions. However, in that prior art, the application is limited to HOA (HOA=Higher Order Ambisonics) (sound field directing) signals in the spatial domain. Moreover, the warping function only depends on the ratio between the reference screen and the reproduction screen, and no definition is given for a non-centered screen.

[0007] Another prior art "Vorrichtung und Verfahren zum Bestimmen einer Wiedergabe-position" (see [2]) describes a method for adapting the position of a sound source to a video playback. The playback position of the sound source is determined individually for each audio object depending on the direction and distance relative to a reference position as well as the camera parameters. The prior art also describes that a screen with a fixed reference size is assumed. In order to adapt the scene to a playback screen that is larger or smaller than the reference screen, a linear scaling of all position parameters (in Cartesian coordinates) is performed. However, according to the prior art, incorporating physical camera and projection parameters is complicated and such parameters are not always available. Moreover, the prior art method works in Cartesian coordinates (x, y, z), so that scene scaling changes not only the position of the object but also its distance. Furthermore, the prior art is not applicable to adapting the object position to changes in the relative screen size in angular coordinates (opening angle, viewing angle).

[0008] In further prior art "Verfahren zur Audiocodierung" (see [3]), a method is described which involves transmitting the current (time-varying) horizontal and vertical viewing angles in a data stream (a reference viewing angle with respect to the listener's position in the original scene). On the playback side, the size and position of the playback are analyzed and the playback of the audio objects is individually optimized to match the reference screen.

[0009] In another prior art work "Acoustical Zooming Based on a parametric Sound Field Representation" (see [4]), a method is described that allows audio rendering to follow the movement of the visual scene ("acoustic zoom"). The acoustic zoom operation is defined as a shift of the virtual recording position. The scene model of the zoom algorithm places all sound sources on a circle with an arbitrary but fixed radius. However, that prior art method works in the DirAC parameter domain, where distances and angles (directions of arrival) are varied, the mapping function is non-linear and depends on the zoom factor / parameters, and non-centered screens are not supported. [Prior art documents] [Patent documents]

[0010] [Patent Document 1] European Patent Application Publication No. 20120305271 [Patent Document 2] International Publication No. 2004073352 Brochure [Patent Document 3] European Patent Application Publication No. 20020024643 [Non-patent literature]

[0011] [Non-Patent Document 1] "Acoustical Zooming Based on a Parametric Sound Field Representation" http: / / www.aes.org / tmpFiles / elib / 20140814 / 15417.pdf Summary of the Invention [Problem to be solved by the invention]

[0012] The object of the present invention is to provide an improved concept for audio and visual multimedia content integration utilizing existing multimedia playback setups. This object is solved by an apparatus according to claim 1, a decoder device according to claim 13, a method according to claim 14 and a computer program according to claim 15. [Means for solving the problem]

[0013] An apparatus for audio object remapping is provided. The apparatus includes an object metadata processor and an object renderer. The object renderer is configured to receive an audio object. The object metadata processor is configured to receive metadata including an indication as to whether the audio object is screen-related or not and further including a first position of the audio object. Moreover, the object metadata processor is configured to calculate a second position of the audio object in response to the first position of the audio object and in response to a size of the screen if the audio object is indicated as screen-related in the metadata. The object renderer is configured to generate a speaker signal in response to the audio object and the position information. The object metadata processor is configured to provide the first position of the audio object as position information to the object renderer if the audio object is indicated as not screen-related in the metadata. Furthermore, the object metadata processor is configured to provide the second position of the audio object as position information to the object renderer if the audio object is indicated as screen-related in the metadata.

[0014] According to one embodiment, the object metadata processor may be configured to not calculate the second position of an audio object, for example if the audio object is indicated in the metadata as not being screen related.

[0015] In one embodiment, the object renderer may be configured, for example, to not determine whether the position information is a first position of an audio object or a second position of an audio object.

[0016] According to one embodiment, the object renderer may for example be configured to generate loudspeaker signals further depending on the number of loudspeakers in the playback environment.

[0017] In one embodiment, the object renderer may be configured to generate the speaker signals further depending on the speaker positions of each of the speakers in the playback environment, for example.

[0018] According to one embodiment, the object metadata processor is configured to calculate a second position of the audio object depending on a first position of the audio object and a size of the screen if the audio object is indicated in the metadata as being screen related, the first position indicating a first position in three dimensional space and the second position indicating a second position in three dimensional space.

[0019] In one embodiment, the object metadata processor may be configured to calculate a second position of the audio object depending on the first position of the audio object and the size of the screen, for example if the audio object is indicated in the metadata as being screen related, the first position indicating a first azimuth angle, a first elevation angle and a first distance, and the second position indicating a second azimuth angle, a second elevation angle and a second distance.

[0020] According to an embodiment, the object metadata processor may be configured to receive metadata including, for example, as a first instruction, an instruction as to whether the audio object is screen related or not, and if the audio object is screen related, further including a second instruction, said second instruction indicating whether the audio object is an on-screen object or not. The object metadata processor may be configured to calculate a second position of the audio object in response to the first position of the audio object and in response to the size of the screen, for example if the second instruction indicates that the audio object is an on-screen object, such that the second position takes a first value on a screen area of ​​the screen.

[0021] In one embodiment, the object metadata processor may be configured to calculate a second position of the audio object depending on the first position of the audio object and the size of the screen, for example if the second instruction indicates that the audio object is not an on-screen object, such that the second position has a second value (either in the screen area or not).

[0022] According to an embodiment, the object metadata processor may be configured to receive metadata including, for example, as a first instruction, an instruction as to whether the audio object is screen related or not, and if the audio object is screen related, further including a second instruction, said second instruction indicating whether the audio object is an on-screen object or not. The object metadata processor may be configured to calculate, for example, if the second instruction indicates that the audio object is an on-screen object, a second position of the audio object depending on the first position of the audio object, the size of the screen and a first mapping curve as a mapping curve, the first mapping curve defining a mapping of original object positions in a first value interval to remapped object positions in a second value interval. Furthermore, the object metadata processor may be configured to calculate a second position of the audio object depending on the first position of the audio object, the size of the screen and a second mapping curve as a mapping curve, for example when the second instruction indicates that the audio object is not an on-screen object, the second mapping curve defining a mapping of an original object position in a first value interval to a remapped object position in a third value interval, the second value interval being included by the third value interval and the second value interval being smaller than the third value interval.

[0023] In one embodiment, each of the first value interval, the second value interval, and the third value interval may be, for example, a value interval of an azimuth angle, and each of the first value interval, the second value interval, and the third value interval may be, for example, a value interval of an elevation angle.

[0024] According to an embodiment, the object metadata processor may be configured to calculate a second position of the audio object in response to at least one of a first linear mapping function and a second linear mapping function, the first linear mapping function being defined to map the first azimuth angle values ​​to the second azimuth angle values ​​and the second linear mapping function being defined to map the first elevation angle values ​​to the second elevation angle values; TIFF0007675506000001.tif11166 indicates the azimuth screen left edge reference, TIFF0007675506000002.tif11168 shows the azimuth screen right edge reference, TIFF0007675506000003.tif11168 shows the elevation screen top reference, TIFF0007675506000004.tif11162 shows the elevation screen bottom reference, TIFF0007675506000005.tif9151 indicates the azimuth screen left edge of the screen, TIFF0007675506000006.tif8155 shows the right azimuth screen edge of the screen, TIFF0007675506000007.tif8162 shows the elevation screen top edge of the screen, TIFF0007675506000008.tif8159 shows the elevation angle of the bottom edge of the screen. TIFF0007675506000009.tif9162 indicates the first azimuth angle value, TIFF0007675506000010.tif9156 indicates a second azimuth angle value, θ indicates a first elevation angle value, θ' indicates a second elevation angle value, and TIFF0007675506000011.tif9156 is, for example, a first azimuth value obtained by a first linear mapping function according to the following formula: From the first mapping of TIFF0007675506000012.tif9162, The second elevation angle value θ′ may result from a second mapping of the first elevation angle value θ with a second linear mapping function, for example, according to the following equation: TIFF0007675506000014.tif49161

[0025] Moreover, a decoder device is provided. The decoder device comprises a USAC decoder for obtaining one or more audio input channels, obtaining one or more input audio objects, obtaining compressed object metadata, and decoding the bitstream to obtain one or more SAOC transport channels. Furthermore, the decoder device comprises a SAOC decoder for decoding the one or more SAOC transport channels to obtain a first group of one or more rendered audio objects. Moreover, the decoder device comprises an apparatus according to the above-mentioned embodiment. The apparatus comprises an object metadata decoder which is an object metadata processor of the apparatus according to the above-mentioned embodiment and is implemented for decoding the compressed object metadata to obtain uncompressed metadata, the apparatus further comprising an object renderer of the apparatus according to the above-mentioned embodiment for rendering the one or more input audio objects in response to the uncompressed metadata to obtain a second group of one or more rendered audio objects. Furthermore, the decoder device comprises a format converter for converting the one or more audio input channels to obtain one or more transformed channels. Moreover, the decoder device includes a mixer for mixing one or more audio objects of a first group consisting of the one or more rendered audio objects, one or more audio objects of a second group consisting of the one or more rendered audio objects, and the one or more transformed channels to obtain one or more decoded audio channels.

[0026] Further, a method for generating a speaker signal is provided, the method comprising the following steps. - receiving an audio object. - receiving metadata including an indication as to whether the audio object is screen-related or not and further including a first position of the audio object; - if the audio object is indicated in the metadata as being screen-related, calculating a second position of the audio object as a function of the first position of the audio object and the size of the screen. - generating loudspeaker signals as a function of the audio object and position information;

[0027] If the audio object is indicated in the metadata as not screen-related, the position information is a first position of the audio object. If the audio object is indicated in the metadata as being screen-related, the position information is a second position of the audio object.

[0028] Moreover, a computer program is provided which, when run on a computer or signal processor, is configured to carry out the above-mentioned method. [Brief description of the drawings]

[0029] [Figure 1] FIG. 2 illustrates an apparatus for generating speaker signals according to one embodiment. [Diagram 2] FIG. 2 illustrates an object renderer according to one embodiment. [Diagram 3] FIG. 2 illustrates an object metadata processor according to one embodiment. [Figure 4] FIG. 1 illustrates azimuth remapping according to an embodiment. [Diagram 5] FIG. 1 illustrates elevation remapping according to an embodiment. [Figure 6] FIG. 1 illustrates azimuth remapping according to an embodiment. [Figure 7] FIG. 13 illustrates elevation remapping according to another embodiment. [Figure 8] FIG. 1 is a schematic diagram of a 3D audio decoder. [Figure 9] FIG. 2 is a schematic diagram of a 3D audio decoder according to one embodiment; [Figure 10] FIG. 2 illustrates the structure of a format converter. [Figure 11] FIG. 2 illustrates object-based audio rendering according to one embodiment. [Figure 12] FIG. 2 illustrates an object metadata preprocessor according to one embodiment. [Figure 13] FIG. 1 illustrates azimuth remapping according to one embodiment. [Figure 14] FIG. 13 illustrates elevation angle remapping according to one embodiment. [Figure 15] FIG. 1 illustrates azimuth angle remapping according to one embodiment. [Figure 16] FIG. 13 illustrates elevation remapping according to another embodiment. [Figure 17] FIG. 13 illustrates elevation remapping according to a further embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0030] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0031] 1 shows an apparatus for audio object remapping according to one embodiment. The apparatus comprises an object metadata processor 110 and an object renderer 120.

[0032] The object renderer 120 is configured to receive an audio object.

[0033] The object metadata processor 110 is configured to receive metadata including an indication as to whether an audio object is screen-related or not and further including a first position of the audio object. Moreover, the object metadata processor 110 is configured to calculate a second position of the audio object depending on the first position of the audio object and the size of the screen if the audio object is indicated in the metadata as being screen-related.

[0034] The object renderer 120 is configured to generate speaker signals according to the audio object and position information.

[0035] The object metadata processor 110 is configured to provide the first position of the audio object as position information to the object renderer 120 if the audio object is indicated in the metadata as not screen related.

[0036] Furthermore, the object metadata processor 110 is configured to provide the second position of the audio object as position information to the object renderer 120 if the audio object is indicated in the metadata as being screen-related.

[0037] According to one embodiment, the object metadata processor 110 may be configured to not calculate the second position of an audio object, for example, if the audio object is indicated in the metadata as not being screen related.

[0038] In one embodiment, the object renderer 120 may be configured, for example, to not determine whether the position information is a first position of an audio object or a second position of an audio object.

[0039] According to one embodiment, the object renderer 120 may be configured to generate speaker signals further depending on, for example, the number of speakers in the playback environment.

[0040] In one embodiment, the object renderer 120 may be configured to generate the speaker signals further depending on the speaker positions of each of the speakers in the playback environment, for example.

[0041] According to one embodiment, the object metadata processor 110 is configured to calculate a second position of the audio object depending on a first position of the audio object and the size of the screen if the audio object is indicated in the metadata as being screen related, the first position indicating a first position in three-dimensional space and the second position indicating a second position in three-dimensional space.

[0042] In one embodiment, the object metadata processor 110 may be configured to calculate a second position of the audio object depending on the first position of the audio object and the size of the screen, for example if the audio object is indicated in the metadata as being screen related, where the first position indicates a first azimuth angle, a first elevation angle and a first distance, and the second position indicates a second azimuth angle, a second elevation angle and a second distance.

[0043] According to one embodiment, the object metadata processor 110 may be configured to receive metadata including, for example, as a first instruction, an instruction as to whether the audio object is screen related or not, and if the audio object is screen related, further including a second instruction, said second instruction indicating whether the audio object is an on-screen object or not. The object metadata processor 110 may be configured to calculate a second position of the audio object depending on the first position of the audio object and depending on the size of the screen, for example if the second instruction indicates that the audio object is an on-screen object, such that the second position takes a first value on the screen area of ​​the screen.

[0044] In one embodiment, the object metadata processor 110 may, for example, determine if the second location is in the screen area or not in the on-screen area if the second indication indicates that the audio object is not an on-screen object. The audio object may be configured to calculate a second position of the audio object depending on the first position of the audio object and depending on the size of the screen, such that:

[0045] According to an embodiment, the object metadata processor 110 may be configured to receive metadata including, for example, as a first instruction, an instruction as to whether the audio object is screen related or not, and if the audio object is screen related, further including a second instruction, said second instruction indicating whether the audio object is an on-screen object or not. The object metadata processor 110 may be configured to calculate, for example, if the second instruction indicates that the audio object is an on-screen object, a second position of the audio object depending on the first position of the audio object, the size of the screen and a first mapping curve as a mapping curve, the first mapping curve defining a mapping of original object positions in a first value interval to remapped object positions in a second value interval. Furthermore, the object metadata processor 110 may be configured to calculate a second position of the audio object depending on the first position of the audio object, the size of the screen, and a second mapping curve as a mapping curve, for example if the second instruction indicates that the audio object is not an on-screen object, the second mapping curve defining a mapping of an original object position in a first value interval to a remapped object position in a third value interval, the second value interval being included by the third value interval and the second value interval being smaller than the third value interval.

[0046] In one embodiment, each of the first value interval, the second value interval, and the third value interval may be, for example, a value interval of an azimuth angle, and each of the first value interval, the second value interval, and the third value interval may be, for example, a value interval of an elevation angle.

[0047] Specific embodiments of the invention and optional features of some embodiments of the invention are described below.

[0048] There may be audio objects (audio signals associated with a position in 3D space, e.g., a given azimuth, elevation and distance) that are not intended for a fixed position, but whose position will change with the size of the screen in the playback setup.

[0049] If an object is signaled as screen-related (eg, by a flag in the metadata), its position is remapped / recalculated relative to the screen size according to certain rules.

[0050] FIG. 2 illustrates an object renderer according to one embodiment.

[0051] By way of introduction the following is noted:

[0052] In object-based audio formats, metadata is stored or transmitted together with the object signal. The audio objects are rendered at the playback side using the metadata and information about the playback environment, such as the number of speakers or the size of the screen. Table 1: Exemplary Metadata TIFF0007675506000015.tif235160

[0053] For an object, geometric metadata can be used to specify how the audio object should be rendered, for example, the azimuth or elevation angle relative to a reference point, for example, the listener, or the absolute position relative to a reference point, for example, the listener. The renderer calculates the speaker signals based on the geometric data and the available speakers and their positions.

[0054] The embodiments according to the present invention will become apparent from the above as follows.

[0055] To control screen-related rendering, additional metadata fields control how the geometric metadata is interpreted.

[0056] If the field is set to OFF, the geometric metadata is interpreted by the renderer to calculate the speaker signals.

[0057] If the field is set to ON, the geometric metadata is mapped by the renderer from nominal data to other values. This remapping is done on the geometric metadata so that renderers that follow the Object Metadata Processor are agnostic about the preprocessing of the object metadata and operate unchanged. Examples of such metadata fields are given in the table below. Table 2: Metadata for controlling screen-related rendering and their meanings: TIFF0007675506000016.tif50166

[0058] Additionally, the nominal screen size or the screen size used during playback of the audio content may be transmitted as metadata information. TIFF0007675506000017.tif20165

[0059] The following table provides an example of how the metadata can be efficiently coded. Table 3 - ObjectMetadataConfig() syntax according to one embodiment TIFF0007675506000018.tif106165hasOnScreenObjects This flag specifies whether there are any screen-related objects. isScreenRelatedObject This flag specifies whether object positions are screen relative or not (positions should be rendered differently so that they can be remapped but still include all valid angle values). isOnScreenObject This flag specifies that the corresponding object is "onscreen". Objects with this flag equal to 1 should be rendered differently, so that their positions only take values ​​on the screen area. According to an alternative, the flag is not used and a reference screen angle is specified. If ScreenRelativeObject=1, all angles are relative to this reference angle. There may be other use cases where an audio object needs to be known to be on the screen.

[0060] With regard to isScreenRelativeObject, it is noted that according to one embodiment there are two possibilities: remap the position but it can still take on all values ​​(screen relative), and remap it so that it can only contain values ​​that are on the screen area (on-screen).

[0061] The remapping is performed in the object metadata processor, which takes into account the local screen size and performs the mapping of the geometric metadata.

[0062] FIG. 3 illustrates an object metadata processor according to one embodiment.

[0063] Regarding screen-related geometric metadata modifications, the following can be said:

[0064] Depending on the information isScreenRelativeObject and isOnScreenObject there are two possibilities for signaling screen-related audio elements: a) Screen-relative audio elements b) On-screen audio elements

[0065] In both cases, the position data of the audio element is remapped by the object metadata processor: a curve is applied that maps the original azimuth and elevation of the position to a remapped azimuth and remapped elevation.

[0066] The reference is the nominal screen size in the metadata or the default screen size assumed.

[0067] For example, the viewing angles defined in ITU-R REC-BT.2022 (General viewing conditions for subjective assessment of the quality of SDTV and HDTV television images on flat panel displays) can be used.

[0068] The difference between these two types of screen associations is the definition of the remapping curve.

[0069] In case a), the remapped azimuth angle can take values ​​between -180° and 180°, and the remapped elevation angle can take values ​​between -90° and 90°. The curve is defined such that azimuth values ​​between the default leftmost azimuth and the default rightmost azimuth are mapped (compressed or stretched) to the interval between a given screen left edge and a given screen right edge (and thus the elevation angles as well). The other azimuth and elevation values ​​are compressed or stretched accordingly so that the full range of values ​​is covered.

[0070] FIG. 4 illustrates azimuth remapping according to an embodiment.

[0071] In case b), the remapped azimuth and elevation angles can only take values ​​that describe a position on the screen area (Azimuth(LeftScreen) Azimuth(Remapped) Azimuth(RightScreen) and Elevation(BottomScreen) Elevation(Remapped) Elevation(TopScreen)).

[0072] There are several different possibilities to handle values ​​outside these ranges. They are all mapped to the edges of the screen, such that all objects between -180° azimuth and the left edge of the screen end up on the left edge of the screen, and all objects between the right edge of the screen and 180° azimuth end up on the right hemisphere. Another possibility is to map the values ​​of the rear hemisphere to the front hemisphere. Then, in the left hemisphere, positions between -180° + azimuth (left edge of the screen) and azimuth (left edge of the screen) are mapped to the left edge of the screen. Values ​​between -180° and -180° + azimuth (left edge of the screen) are mapped to values ​​between 0° and azimuth (left edge of the screen). The right hemisphere and elevation are handled in the same way.

[0073] FIG. 5 illustrates elevation remapping according to an embodiment.

[0074] The points -x1 and +x2 (which may be different or equal to +x1) of the curve where the slope changes are default values ​​(default assumed standard screen size + position) or they may be present in the metadata (e.g. by the player who may then provide the playback screen size there).

[0075] There are also possible mapping functions that are not composed of line segments, but are instead curved.

[0076] Additional metadata may control the method of remapping, for example defining offsets or non-linear coefficients to account for panning behavior or listening resolution.

[0077] It may also signal how the mapping is performed, for example by "projecting" all objects intended for the rear onto the screen.

[0078] Such alternative mapping methods are listed in the figures below.

[0079] Thus, FIG. 6 illustrates azimuth remapping according to an embodiment.

[0080] FIG. 7 illustrates elevation remapping according to an embodiment. Regarding unknown screen size behavior, If the playback screen size is not given, - The default screen size is assumed, or - No mapping is applied even if the object is marked as screen-related or on-screen.

[0081] Returning to FIG. 4, in another embodiment, in case b), the remapped azimuth and elevation angles can only take values ​​that describe a position on the screen area (Azimuth (left edge of screen)≦Azimuth (remapped)≦Azimuth (right edge of screen) and Elevation (bottom edge of screen)≦Elevation (remapped)≦Elevation (top edge of screen)). There are several different possibilities to handle values ​​outside these ranges: In some embodiments, they are all mapped relative to the edges of the screen, such that all objects between +180° azimuth and the left edge of the screen end up at the left edge of the screen, and all objects between the right edge of the screen and −180° azimuth end up at the right edge of the screen. Another possibility is to map the values ​​in the rear half sphere to the front half sphere.

[0082] Then, in the left hemisphere, locations between +180°-Azimuth(left edge of screen) and Azimuth(left edge of screen) are mapped to the left edge of the screen. Values ​​between +180° and +180°-Azimuth(left edge of screen) are mapped to values ​​between 0° and Azimuth(left edge of screen). The right hemisphere and elevation angles are treated similarly.

[0083] Figure 16 is a diagram similar to Figure 5. In the embodiment illustrated by Figure 16, a value interval on the abscissa axis from -90° to +90° and a value interval on the ordinate axis from -90° to +90° are shown in both figures.

[0084] Figure 17 is a diagram similar to Figure 7. In the embodiment illustrated by Figure 17, a value interval on the abscissa axis from -90° to +90° and a value interval on the ordinate axis from -90° to +90° are shown in both figures.

[0085] Further embodiments of the invention and optional features of the further embodiments will now be described with reference to Figures 8 to 15.

[0086] According to some embodiments, screen-related element remapping can be processed only if, for example, the bitstream contains screen-related elements accompanied by OAM data (OAM data = related object metadata) (isScreenRelativeObject flag == 1 for at least one audio element) and the local screen size is signaled to the decoder via the LocalScreenSize() interface.

[0087] The geometric position data (OAM data before any positional modifications due to user interaction) can be mapped to different ranges of values, for example, by defining and utilizing a mapping function. The remapping can, for example, modify the geometric position data as a pre-processing step to rendering, such that the renderer is agnostic about the remapping and operates unchanged.

[0088] The screen size of a nominal reference screen (used in the mixing and monitoring process) and / or local screen size information in the reproduction room may for example be taken into account for the remapping.

[0089] If no nominal reference screen size is given, default reference values ​​may be used, for example assuming a 4k display and optimal viewing distance as an example.

[0090] If no local screen size information is given, for example, no remapping should be applied.

[0091] For example, two linear mapping functions can be defined for mapping the elevation and azimuth values.

[0092] The screen edges of the nominal screen size can for example be given by: TIFF0007675506000019.tif11159

[0093] The regenerative screen ends can be omitted, for example, by: TIFF0007675506000020.tif11163

[0094] The remapping of the azimuth and elevation position data may be defined, for example, by the following linear mapping function: TIFF0007675506000021.tif50161TIFF0007675506000022.tif49161

[0095] Figure 13 shows a position data remapping function according to one embodiment. In particular, in Figure 13, a mapping function for mapping azimuth angles is shown. In Figure 13, a curve is defined such that azimuth angle values ​​between the nominal reference left edge azimuth angle and the nominal reference right edge azimuth angle are mapped (compressed or stretched) to the interval between a given local screen left edge and a given local screen right edge. Other azimuth angle values ​​are compressed or stretched accordingly so that the full range of values ​​is covered.

[0096] The remapped azimuth angle may, for example, take on values ​​between -180° and 180°, and the remapped elevation angle may take on values ​​between -90° and 90°.

[0097] For example, according to one embodiment, if the isScreenRelativeObject flag is set to 0, no screen-relative element remapping is applied for the corresponding element and the geometric position data (OAM data + position changes due to user interaction) is directly used by the renderer to calculate the playback signal.

[0098] According to some embodiments, the positions of all screen-associated elements may be remapped according to the playback screen size, for example as an adaptation to the playback room, e.g., if no playback screen size information is given or if no screen-associated elements are present, no remapping is applied.

[0099] The remapping may be defined by a linear mapping function that takes into account eg reproduction screen size information in the reproduction room and eg screen size information of a reference screen used in the mixing and monitoring process.

[0100] An azimuth mapping function according to one embodiment is shown in Figure 13. In Figure 13 above, a mapping function for azimuth angles is shown. As in Figure 13, this function may be defined such that, for example, azimuth values ​​between the left and right edges of the reference screen are mapped (compressed or stretched) to the interval between the left and right edges of the reproduction screen. Other azimuth values ​​are compressed or stretched so that the full range of values ​​is covered.

[0101] An elevation mapping function may, for example, be defined accordingly (see FIG. 14). Screen-related processing may also take into account zoom regions, for example for zooming into high-definition video content. Screen-related processing may, for example, be defined only for elements that are labeled as screen-related, with dynamic position data.

[0102] In the following, a system overview of a 3D audio codec system is provided, in which embodiments of the present invention can be utilized, which may for example be based on the MPEG-D USAC codec for coding of channel and object signals.

[0103] According to embodiments, MPEG SAOC technology is adapted to increase the coding efficiency of large amounts of objects (SAOC=Spatial Audio Object Coding). For example, according to some embodiments, three types of renderers can perform the tasks of, for example, rendering objects for channels, rendering channels for headphones, or rendering channels for different speaker setups.

[0104] When an object signal is either explicitly transmitted or parametrically encoded using SAOC, the corresponding object metadata information is compressed and multiplexed into the 3D audio bitstream.

[0105] Figures 8 and 9 show different algorithmic blocks of a 3D audio system. In particular, Figure 8 shows a schematic of a 3D audio encoder, and Figure 9 shows a schematic of a 3D audio decoder according to one embodiment.

[0106] A possible embodiment of the module of Figures 8 and 9 will now be described.

[0107] In Fig. 8, a pre-renderer 810 (also called mixer) is shown. In the configuration of Fig. 8, the pre-renderer 810 (mixer) is optional. The pre-renderer 810 can be optionally used to convert the channel+object input scene into a channel scene before encoding. Functionally, the pre-renderer 810 on the encoder side can be related to the functionality of, for example, the object renderer / mixer 920 on the decoder side, which will be described later. Pre-rendering of objects ensures a deterministic signal entropy at the encoder input that is essentially independent of the number of simultaneously active object signals. Pre-rendering of objects does not require object metadata transmission. Discrete object signals are rendered into the channel layout that the encoder is configured to use. The weights of the objects per channel are obtained from the associated object metadata (OAM).

[0108] The core codecs for the speaker-channel signals, the discrete object signals, the object downmix signals and the pre-rendered signals are based on the MPEG-D USAC technology (USAC Core Codec). The USAC encoder 820 (e.g., shown in FIG. 8) handles the coding of multiple signals by creating channel and object mapping information based on the geometric and semantic information of the input channel and object assignments. This mapping information describes how the input channels and objects are mapped to the USAC channel elements (CPE, SCE, LFE) and the corresponding information is sent to the decoder.

[0109] Any additional payload, such as SAOC data or object metadata, is passed through the extension element and can be taken into account, for example, in the rate control of the USAC encoder.

[0110] The objects can be coded in different ways depending on the rate / distortion and interactivity requirements of the renderer. The following object coding variants are possible: Pre-rendered objects: Object signals are pre-rendered and mixed into the 22.2 channel signal before encoding. The subsequent coding chain takes into account the 22.2 channel signal. Discrete object waveform: The object is fed to the USAC Encoder 820 as a mono waveform. The USAC Encoder 820 transmits the object in addition to the channel signal using a single channel element SCE. The decoded object is rendered and mixed at the receiver side. Compressed object metadata information is sent along to the receiver / renderer. Parametric object waveform: object properties and their relation to each other are described by SAOC parameters. The downmix of the object signals is coded with USAC by the USAC encoder 820. The parametric information is transmitted together. The number of downmix channels is selected depending on the number of objects and the overall data rate. The compressed object metadata information is transmitted to the SAOC renderer.

[0111] On the decoder side, USAC decoder 910 performs USAC decoding.

[0112] Moreover, according to an embodiment, a decoder device is provided, see Fig. 9. The decoder device comprises a USAC decoder 910 for obtaining one or more audio input channels, obtaining one or more input audio objects, obtaining compressed object metadata, and decoding a bitstream to obtain one or more SAOC transport channels.

[0113] Further, the decoder device includes a SAOC decoder 915 for decoding the one or more SAOC transport channels to obtain a first group of one or more rendered audio objects.

[0114] Moreover, the decoder device comprises an apparatus 917 according to the embodiments described above in relation to figures 1 to 7 or as described below in relation to figures 11 to 15. The apparatus 917 is the object metadata processor 110 of the apparatus of figure 1 and comprises an object metadata decoder 918 implemented for decoding the compressed object metadata to obtain uncompressed metadata.

[0115] Furthermore, the apparatus 917 according to the above-mentioned embodiment comprises an object renderer 920, e.g. the object renderer 120 of the apparatus of FIG. 1, for rendering one or more input audio objects in response to the uncompressed metadata to obtain a second group of one or more rendered audio objects.

[0116] Furthermore, the decoder device comprises a format converter 922 for converting the one or more audio input channels to obtain one or more converted channels.

[0117] Moreover, the decoder device includes a mixer 930 for mixing one or more audio objects of a first group consisting of one or more rendered audio objects, one or more audio objects of a second group consisting of the one or more rendered audio objects, and one or more transformed channels to obtain one or more decoded audio channels.

[0118] In Fig. 9 a particular embodiment of a decoder device is shown. The SAOC encoder 815 for object signals (SAOC encoder 815 is optional, see Fig. 8) and the SAOC decoder 915 (see Fig. 9) are based on the MPEG SAOC technology. The system is able to regenerate, modify and render some audio objects based on fewer transmitted channels and additional parametric data (OLD, IOC, DMG) (OLD = object level difference, IOC = inter-object correlation, DMG = downmix gain). The additional parametric data represents a significantly lower data rate than would be required to transmit all objects individually, making the coding very efficient.

[0119] The SAOC Encoder 815 takes as input the object / channel signal as a mono waveform and outputs the parametric information (packed into a 3D audio bitstream) and the SAOC transport channel (encoded and transmitted using single channel elements).

[0120] The SAOC decoder 915 reconstructs the object / channel signals from the decoded SAOC transport channel and parametric information and generates an output audio scene based on the playback layout, the decompressed object metadata information and, optionally, user interaction information.

[0121] For the object metadata codec, for each object, the associated metadata specifying the geometric location and distribution of the object in 3D space is efficiently coded by quantization of object properties in time and space, e.g. by the metadata encoder 818 of Fig. 8. The compressed object metadata cOAM (cOAM = compressed audio object metadata) is transmitted as side information to the receiver. At the receiver side, the cOAM is decoded by the metadata decoder 918.

[0122] For example, in FIG. 9, the metadata decoder 918 may implement an object metadata processor according to, for example, one of the embodiments described above.

[0123] An object renderer, e.g., object renderer 920 in Fig. 9, utilizes the compressed object metadata to generate object waveforms in a given playback format. Each object is rendered to a specific output channel according to its metadata. The output of this block comes from the sum of the partial results.

[0124] For example, in FIG. 9, object renderer 920 may be implemented, for example, according to one of the embodiments described above.

[0125] In Fig. 9, the metadata decoder 918 may be implemented as an object metadata processor, for example as described according to one of the above-mentioned or later-described embodiments described with reference to Figs. 1-7 and 11-15, and the object renderer 920 may be implemented as an object renderer, for example as described according to one of the above-mentioned or later-described embodiments described with reference to Figs. 1-7 and 11-15. The metadata decoder 918 and the object renderer 920 may, for example, together implement an apparatus for generating speaker signals 917, for example as described above or later with reference to Figs. 1-7 and 11-15.

[0126] If both channel-based content and discrete / parametric objects are decoded, the channel-based waveforms and the rendered object waveforms are mixed before the resulting waveform is output, for example by the mixer 930 of FIG. 9 (or before feeding them to a post-processing device module such as a binaural renderer or speaker renderer module).

[0127] The binaural renderer module 940 can, for example, generate a binaural downmix of multi-channel audio material, where each input channel is represented by a virtual source. This processing is performed frame-by-frame in the QMF domain. The binauralization may, for example, be based on measured binaural room impulse responses.

[0128] The Speaker Renderer 922 can, for example, convert between a transmission channel configuration and a desired playback format. It is therefore referred to below as the Format Converter 922. The Format Converter 922 performs the conversion to a smaller number of output channels, for example generating a downmix. The system automatically generates downmix matrices optimized for a given combination of input and output formats and applies these matrices in the downmix process. The Format Converter 922 allows for standard speaker configurations and random configurations with non-standard speaker positions.

[0129] The structure of the format converter is shown in Fig. 10. Fig. 10 shows a downmix composer 1010 and a downmix processor for processing the downmix in the QMF domain (QMF domain = quadrature mirror filter domain).

[0130] According to some embodiments, the object renderer 920 may be configured to provide screen-related audio object remapping, such as described in connection with one of the embodiments described above with reference to Figures 1-7, or as described in connection with one of the embodiments described below with reference to Figures 11-15.

[0131] Further embodiments of the invention and concepts of embodiments of the invention are described below.

[0132] According to some embodiments, user control of an object can utilize, for example, descriptive metadata, e.g., information regarding whether the object is present within the bitstream and high-level characteristics of the object, and can utilize, for example, restrictive metadata, e.g., information regarding how interaction is possible or enabled by the content generator.

[0133] According to some embodiments, the signaling, delivery and rendering of audio objects can utilize, for example, positional metadata, structural metadata, e.g., object groupings and hierarchies, capabilities for rendering to specific speakers and signaling channel content as objects, and means for adapting object scenes to screen sizes.

[0134] The embodiment provides new metadata fields that are being developed in addition to the geometric positions and levels of objects already defined in 3D space.

[0135] When an object-based audio scene is played back in different playback settings, according to some embodiments, the positions of the rendered sound sources can be scaled to the dimensions of the playback, for example automatically. When audio-visual content is presented, the placement of the sound sources and the position of the visual source of the sound may, for example, no longer match, so that standard rendering of the audio objects for playback can result in, for example, disruption of positional audio-visual consistency.

[0136] To avoid this effect, the possibility may be exploited to signal that e.g. audio objects are not intended for a fixed position in 3D space, but that their position should change with the size of the screen in the playback setup. According to some embodiments, special handling of these audio objects, as well as the definition of scene scaling algorithms, may enable e.g. a more immersive experience, since the playback can be optimized for e.g. the local characteristics of the playback environment.

[0137] In some embodiments, the renderer or pre-processing module can take into account, for example, the local screen size in the playback room, and thus preserve the relationship between audio and video, for example, in movie or game aspects. In such embodiments, the audio scene can then be scaled according to the playback settings, for example, so that the positions of visual elements match the positions of corresponding sound sources. For example, positional consistency between audio and visual for screens of varying sizes can be maintained.

[0138] For example, according to embodiments, dialogues and speech can then be perceived from the direction of the speaker on the screen, regardless of the playback screen size, and this is possible for stationary sound sources and for moving sound sources, for which the trajectory of the sound and the movement of the visual elements must correspond.

[0139] To control screen-related rendering, an additional metadata field is introduced that allows marking an object as screen-related. If an object is marked as screen-related, its geometric position data is remapped to other values ​​before rendering. For example, Figure 13 shows an exemplary (re)mapping for azimuth angles.

[0140] In particular, some embodiments may provide that a simple mapping function is defined that operates, for example, in the angular domain (azimuth, elevation).

[0141] Moreover, some embodiments may provide that, for example, the distance of the object does not change, no "zooming" or virtual translation towards or away from the screen is performed, but rather a scaling of the object's position only.

[0142] Furthermore, some embodiments are advantageous in that the mapping function is not only based on the screen ratio, but also takes into account the azimuth and elevation angles of the screen edges, so that, for example, a non-centered reproduction screen TIFF0007675506000023.tif11161 can be processed.

[0143] Moreover, some embodiments may, for example, define special mapping functions for on-screen objects. According to some embodiments, the azimuth and elevation mapping functions may, for example, be independent, so that one may choose to remap only the azimuth or elevation values.

[0144] Further embodiments are provided below.

[0145] Fig. 11 illustrates object-based audio rendering according to one embodiment. Audio objects can be rendered on the playback side using, for example, metadata and information about the playback environment. Such information is, for example, the number of speakers or the size of the screen. The renderer 1110 can calculate speaker signals based on the geometric data as well as the available speakers and their positions.

[0146] An object metadata (pre)processor 1210 according to one embodiment will now be described with reference to FIG.

[0147] In FIG. 12, the object metadata processor 1210 is configured to perform a remapping that takes into account the local screen size and performs a mapping of the geometric metadata.

[0148] Position data of a screen-related object is remapped by the object metadata processor 1210. For example, a curve can be applied that maps the original azimuth and elevation angles of the position to remapped azimuth and remapped elevation angles.

[0149] For example, the screen size of a nominal reference screen utilized in the mixing and monitoring process as well as local screen size information in the reproduction room may for example be taken into account for remapping.

[0150] For example, a reference screen size, which may be referred to as the playback screen size, may be transmitted, for example, in the metadata.

[0151] In some embodiments, if a nominal screen size is not given, for example, a default screen size may be assumed.

[0152] For example, the viewing angles defined in ITU-R REC-BT.2022 (Reference: General Viewing Conditions for the Subjective Assessment of the Quality of SDTV and HDTV Television Images on Flat Panel Displays) can be used as an example.

[0153] In some embodiments, for example, two linear mapping functions can be defined for remapping the elevation and azimuth values.

[0154] In the following, modification of screen-related geometric metadata according to some embodiments will be explained with reference to Figs.

[0155] The remapped azimuth angles can take values ​​between -180° and 180°, and the remapped elevation angles can take values ​​between -90° and 90°. The mapping curve is typically defined such that azimuth values ​​between the default leftmost azimuth and the default rightmost azimuth are mapped (compressed or stretched) into the interval between a given screen left edge and a given screen right edge (and thus the elevation angles as well). Other azimuth and elevation values ​​are compressed or stretched accordingly so that the full range of values ​​is covered.

[0156] As already mentioned above, the screen edges of the nominal screen size can for example be given by: TIFF0007675506000024.tif11159

[0157] The regenerative screen ends can be omitted, for example, by: TIFF0007675506000025.tif11163

[0158] The remapping of the azimuth and elevation position data may be defined, for example, by the following linear mapping function: TIFF0007675506000026.tif50161TIFF0007675506000027.tif49161

[0159] The mapping function for the azimuth angle is shown in FIG. 13, and the mapping function for the elevation angle is shown in FIG.

[0160] A point on a curve where the slope can change TIFF0007675506000028.tif11159 may be set as default values ​​(default assumed standard screen size and default assumed standard screen position) or they may be present in the metadata (e.g. by the player which may then provide the playback / monitoring screen size there).

[0161] Regarding the definition of object metadata for screen-relative remapping, an additional metadata flag named "isScreenRelativeObject" is defined to control screen-relative rendering. This flag can define, for example, whether an audio object should be processed / rendered relative to the local playback screen size or not.

[0162] If screen-related elements are present in the audio scene, the possibility is provided to provide screen size information of the nominal reference screen used for mixing and monitoring (screen size in use during playback of the audio content). Table 4 - ObjectMetadataConfig() syntax according to one embodiment TIFF0007675506000029.tif167167hasScreenRelativeObjects This flag specifies whether or not screen-related objects exist. hasScreenSize This flag specifies whether the nominal screen size is defined. The definition is done in terms of the viewing angles that correspond to the screen edges. If hasScreenSize is zero, the following values ​​are used as defaults: TIFF0007675506000030.tif42157bsScreenSizeAz This field defines the azimuth angle corresponding to the left and right edges of the screen. TIFF0007675506000031.tif42156bsScreenSizeTopEl This field defines the elevation angle corresponding to the top edge of the screen. TIFF0007675506000032.tif21159bsScreenSizeBottomEl This field defines the elevation angle corresponding to the bottom edge of the screen. TIFF0007675506000033.tif21161isScreenRelativeObject This flag specifies whether object positions are screen-relative or not (positions should be rendered differently so that they can be remapped but still include all valid angle values).

[0163] According to one embodiment, if no playback screen size is given, then a default playback screen size and a default playback screen position are assumed or no mapping is applied even if the object is marked as screen-related.

[0164] Some of the embodiments realize possible variations.

[0165] In some embodiments, non-linear mapping functions are utilized. These possible mapping functions are not composed of lines, but are instead curved. In some embodiments, additional metadata controls the method of remapping, for example defining offsets or non-linear coefficients to account for panning behavior or listening resolution.

[0166] Some embodiments provide for independent processing of azimuth and elevation angles. Azimuth and elevation angles can be marked and processed as screen-related independently. Table 5 shows the syntax of ObjectMetadataConfig() according to such an embodiment. Table 5: Syntax of ObjectMetadataConfig() according to one embodiment TIFF0007675506000034.tif130166

[0167] Some embodiments utilize the definition of an on-screen object, which can be differentiated between screen-related objects and on-screen objects, where the possible syntax can be that of Table 6 below. Table 6 - ObjectMetadataConfig() syntax according to one embodiment TIFF0007675506000035.tif147165hasOnScreenObjects This flag specifies whether there are any screen-related objects. isScreenRelatedObject This flag specifies whether object positions are screen relative or not (positions should be rendered differently so that they can be remapped but still include all valid angle values). isOnScreenObject This flag specifies whether the corresponding object is "on-screen" or not. Objects with this flag equal to 1 should be rendered differently, such that their positions only take on values ​​on the screen area.

[0168] For on-screen objects, the remapped azimuth and elevation angles can only take values ​​that describe a position on the screen area. TIFF0007675506000036.tif10163.

[0169] As implemented by some embodiments, there are several different possibilities for handling these out-of-range values: They can be mapped to the edge of the screen, then in the left hemisphere, 180° and The position between TIFF0007675506000037.tif11158 is the left edge of the screen. TIFF0007675506000038.tif11166. The right hemisphere and elevation are treated similarly (non-dashed mapping function 1510 in FIG. 15).

[0170] Another possibility realized by some of the embodiments is to map the values ​​of the rear half sphere to the front half sphere. TIFF0007675506000039.tif11158 values ​​between 0° and TIFF0007675506000040.tif9151. The right hemisphere and elevation are treated similarly (dashed mapping function 1520 in FIG. 15).

[0171] FIG. 15 illustrates azimuth remapping (on-screen object) according to these embodiments.

[0172] The selection of the desired behavior can be determined by additional metadata (e.g., backward ([180° and TIFF0007675506000041.tif11158] and [-180° and The on-screen object may be signaled by a 32-bit LSB executable (a flag to "project" all on-screen objects intended for the 32-bit LSB executable [TIFF0007675506000042.tif11161]) or a 32-bit LSB executable (a flag to "project" all on-screen objects intended for the 32-bit LSB executable [TIFF0007675506000042.tif11161]) on the screen.

[0173] Although some aspects are described in terms of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, with a block or device corresponding to a method step or feature of a method step, and similarly, aspects described in terms of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.

[0174] The decomposed signals of the present invention can be stored on a digital storage medium or transmitted over a transmission medium, such as a wireless or wired transmission medium, such as the Internet.

[0175] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. Implementations can be implemented using digital storage media, such as floppy disks, DVDs, CDs, ROMs, PROMs, EPROMs, EEPROMs or flash memories, on which electronically readable control signals are stored, which cooperate (or can cooperate) with a programmable computer system such that the respective methods are implemented.

[0176] Some embodiments according to the invention include a non-transitory data carrier having electronically readable control signals capable of cooperating with a programmable computer system such that one of the methods described herein is performed.

[0177] Generally, embodiments of the present invention can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier.

[0178] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0179] In other words, therefore, an embodiment of the inventive method is a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0180] A further embodiment of the inventive method is, therefore, a data carrier (or a digital storage medium, or a computer readable medium) having recorded thereon a computer program for performing one of the methods described herein.

[0181] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, the data stream or the sequence of signals may for example be arranged to be transferred via a data communication connection, for example the Internet.

[0182] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0183] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0184] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0185] The above-described embodiments merely illustrate the principles of the embodiments of the present invention. Modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented herein for purposes of description and illustration of the embodiments.

Claims

1. 1. An apparatus for generating a speaker signal, comprising: an object metadata processor (110); an object renderer (120); The object renderer (120) is configured to receive an audio object; the object metadata processor (110) is configured to receive metadata including an indication as to whether the audio object is screen related or not and further including a first position of the audio object; the object metadata processor (110) is configured to calculate a second position of the audio object in response to the first position of the audio object and in response to a size of a screen if the audio object is indicated in the metadata as being screen related; the object renderer (120) is configured to generate the speaker signals in response to the audio objects; if the audio object is indicated in the metadata as being screen-related, the object renderer (120) is configured to generate the speaker signal in response to a second position of the audio object, the second position of the audio object being dependent on the first position of the audio object and a size of a screen; If the metadata indicates that the audio object is screen-related and an on-screen object, the device is configured to generate the speaker signal by calculating the second position according to the first position of the audio object and a size of the screen, such that the second position takes a first value on a screen area of ​​the screen, and generating the speaker signal according to the second position.

2. 2. The apparatus of claim 1, wherein the object metadata processor (110) is configured to not calculate the second position of the audio object if the audio object is indicated in the metadata as not being screen-related.

3. The apparatus of claim 1 or 2, wherein the object renderer (120) is further configured to generate the speaker signals depending on the number of speakers in a playback environment.

4. The apparatus of claim 3 , wherein the object renderer (120) is configured to generate the speaker signals further depending on a speaker position of each of the speakers of the playback environment.

5. 5. The apparatus of claim 1, wherein the object metadata processor (110) is configured to calculate, if the audio object is indicated in the metadata as being screen-related, the second position of the audio object depending on the first position of the audio object and the size of the screen, the first position indicating the first position in a three-dimensional space and the second position indicating the second position in the three-dimensional space.

6. 6. The apparatus of claim 5, wherein the object metadata processor (110) is configured to calculate, if the audio object is indicated in the metadata as being screen-related, the second position of the audio object in response to the first position of the audio object and in response to the size of the screen, the first position indicating a first azimuth angle, a first elevation angle and a first distance, and the second position indicating a second azimuth angle, a second elevation angle and a second distance.

7. the object metadata processor (110) is configured to receive the metadata including as a first indication an indication as to whether the audio object is screen associated, and if the audio object is screen associated, further including a second indication, the second indication indicating whether the audio object is an on-screen object; 7. The apparatus of claim 1, wherein the object metadata processor (110) is configured to calculate the second position of the audio object depending on the first position of the audio object and the size of the screen, such that when the second indication indicates that the audio object is an on-screen object, the second position takes the first value on the screen area of ​​the screen.

8. 8. The apparatus of claim 7, wherein the object metadata processor (110) is configured to calculate the second position of the audio object in response to the first position of the audio object and in response to the size of the screen, if the second indication indicates that the audio object is not the on-screen object, such that the second position takes a second value that is either in a screen area or not.

9. the object metadata processor (110) is configured to receive the metadata including as a first indication an indication as to whether the audio object is screen associated, and if the audio object is screen associated, further including a second indication, the second indication indicating whether the audio object is an on-screen object; the object metadata processor (110) is configured to calculate, if the second indication indicates that the audio object is an on-screen object, the second position of the audio object as a function of the first position of the audio object, the size of the screen, and a first mapping curve as a mapping curve, the first mapping curve defining a mapping of original object positions in a first value interval to remapped object positions in a second value interval; The object metadata processor (110) is configured to calculate, when the second indication indicates that the audio object is not an on-screen object, the second position of the audio object depending on the first position of the audio object, the size of the screen, and a second mapping curve as the mapping curve, the second mapping curve defining a mapping of an original object position in the first value interval to a remapped object position in a third value interval, the second value interval being included by the third value interval and the second value interval being smaller than the third value interval.

10. the first value interval, the second value interval and the third value interval are azimuth value intervals; or 10. The apparatus of claim 9, wherein the first value interval, the second value interval and the third value interval are elevation value intervals.

11. the object metadata processor (110) is configured to calculate the second position of the audio object in response to at least one of a first linear mapping function and a second linear mapping function; the first linear mapping function is defined to map first azimuth angle values ​​to second azimuth angle values; the second linear mapping function is defined to map first elevation angle values ​​to second elevation angle values; indicates the azimuth screen left edge reference, indicates the azimuth screen right edge reference, indicates the elevation screen top reference, indicates the elevation screen bottom reference, indicates the azimuth screen left edge of the screen, indicates the azimuth screen right edge of the screen, indicates the elevation angle of the screen top edge, indicates the elevation angle of the screen at the bottom edge, denotes the first azimuth angle value, denotes the second azimuth angle value, θ denotes the first elevation angle value, θ′ denotes the second elevation angle value, The second azimuth angle value is the first azimuth angle value by the first linear mapping function according to the following formula: from the first mapping of 11. The apparatus of claim 1, wherein the second elevation angle value θ′ results from a second mapping of the first elevation angle value θ with the second linear mapping function according to the following equation:

12. 1. A decoder device comprising: a USAC decoder (910) for decoding the bitstream to obtain one or more audio input channels, to obtain one or more input audio objects, to obtain compressed object metadata, and to obtain one or more SAOC transport channels; a SAOC decoder (915) for decoding the one or more SAOC transport channels to obtain a first group of one or more rendered audio objects; 12. The apparatus (917) according to any one of claims 1 to 11, An object metadata decoder (918) implemented for decoding the compressed object metadata to obtain uncompressed metadata, the object metadata processor (110) of the apparatus of any one of claims 1 to 11, an apparatus (917) comprising an object renderer (920; 120) of the apparatus of any one of claims 1 to 12 for rendering the one or more input audio objects according to the uncompressed metadata to obtain a second group of one or more rendered audio objects; a format converter (922) for converting the one or more audio input channels to obtain one or more converted channels; a mixer (930) for mixing the one or more audio objects of a first group of the one or more rendered audio objects, the one or more audio objects of a second group of the one or more rendered audio objects, and one or more transformed audio channels to obtain one or more decoded audio channels; 23. A decoder device comprising:

13. 1. A method for generating a speaker signal, comprising: receiving an audio object; receiving metadata including an indication as to whether the audio object is screen-related or not and further including a first location of the audio object; generating the speaker signals in response to the audio objects; A method for generating a speaker signal by calculating a second position according to the first position of the audio object and a size of the screen, such that the second position has a first value on a screen area of ​​the screen when the metadata indicates that the audio object is screen-related and an on-screen object, and generating the speaker signal according to the second position.

14. A computer program for carrying out the method according to claim 13 when the computer program is run on a computer or signal processor.

Citation Information

Patent Citations

  • EP20020024643

  • EP20120305271

  • Control of coke-receiving traverse of coke quenching vehicle

    JP1989022995A

  • Acoustic signal multiplex transmission system, manufacturing device, and reproduction device added with sound image localization acoustic meta-information

    JP2009278381A

  • Method and apparatus for playback of high-order ambisonics audio signal

    JP2013187908A