Device and method for audio encoding

The audio encoding apparatus uses input metadata to constrain rendering parameters, optimizing bit rate compression and quality, addressing the trade-offs in existing audio coding technologies for dynamic applications.

JP2025179172APending Publication Date: 2025-12-09KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025146640
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-14
Filing Date
2025-09-04
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing audio coding technologies face challenges in balancing flexibility, audio quality, and complexity, particularly in dynamic applications like virtual reality, where there is a trade-off between interactivity and data rate, leading to suboptimal performance and increased costs.

Method used

An audio encoding apparatus that incorporates input presentation metadata to constrain rendering parameters, generating encoded audio data with associated metadata to allow flexible and controlled rendering, while optimizing bit rate compression and quality.

Benefits of technology

This approach enhances audio quality-to-bitrate ratios, providing flexible and controlled rendering with reduced complexity and data rates, ensuring efficient and customizable audio experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025179172000001_ABST
    Figure 2025179172000001_ABST
Patent Text Reader

Abstract

To provide audio encoding devices and methods for dynamic applications such as virtual reality applications.SOLUTION: An audio encoding device includes: an audio receiver 201 for receiving an audio item indicating an audio scene; a metadata receiver 203 for receiving input presentation metadata for the audio item which describes a presentation constraint for rendering of the audio item; an audio encoder 205 which generates encoded audio data for the audio scene by encoding the plurality of audio items; a metadata circuit 207 for generating output presentation metadata from input presentation metadata; and an output 209 for generating an encoded audio data stream having encoded audio data and output presentation metadata.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an apparatus and method for audio coding, particularly, but not exclusively, for audio coding for dynamic applications such as virtual reality applications. [Background technology]

[0002] The variety and range of audio and video applications has increased significantly in recent years with the continuous development and introduction of new services and modalities for using and consuming audio, images and video.

[0003] For example, one increasingly popular service is the presentation of audio and images in a way that allows the viewer to actively and dynamically interact with the system to change the parameters of the rendering. A highly appealing feature in many applications is the ability to change the effective viewing / listening position. Such a feature, in particular, allows virtual reality experiences to be provided to users.

[0004] The trend is toward providing increased flexibility to allow for rendering-side adaptation of scenes. To provide increased rendering-side flexibility for rendering audio scenes, several audio coding and distribution approaches have been proposed, in which an audio scene is represented by a composition of different audio items. For example, an audio item may represent a separate sound source, such as a specific speaker. In some approaches, all audio items are of the same type, but there is an increasing development of systems that allow multiple different audio types to be used and supported simultaneously. For example, some audio items may be audio channels, others may be separate audio objects, and still others may be scene-based, such as ambisonic audio items. In many systems, metadata is provided along with the audio data representing the audio items. Such metadata may, for example, indicate the nominal location in the scene for the audio source of an audio item.

[0005] Such an approach allows for a high degree of customization and adaptation on the client / rendering side: for example, the audio scene can be adapted locally to changes in the listener's virtual position in the audio scene, or to the specific preferences of individual listeners.

[0006] As a particular example, the 3GPP® consortium is currently developing the so-called Immersive Voice and Audio Services (IVAS) codec, which is capable of coding audio content in various configurations, such as channel-, object-, or scene- (particularly Ambisonics-) based configurations. The goal of the coding is to carry the audio information using the minimum amount of data.

[0007] The IVAS codec will also have a renderer that converts the various audio streams into a format suitable for playback at the receiving end, for example it can map the audio to a known loudspeaker setup or it can render the audio into a binaural format for playback over headphones.

[0008] In the 3GPP® IVAS codec scope, work is underway to collect potential use cases. For these, it is considered that the codec should provide interactivity to modulate the rendering. For example, headphone audio must be rendered independent of head position and translation, which means that it must be compensated for head movement. As another example, a user may be prompted to spatially position an audio item, such as (re)positioning an object that carries the audio of a participant in a virtual meeting.

[0009] The renderer is considered part of the 3GPP® IVAS codec work item and is considered to be internal to the IVAS codec. However, it has been proposed that the codec also include a pass-through mode. This mode allows audio items to be represented at the decoder output in the same configuration as they were input at the encoder input (i.e., as 1:1 corresponding channel, object, and scene-based audio items). An external renderer has access to these items via a dedicated external rendering interface, providing an alternative rendering to the internal IVAS renderer.

[0010] Such an approach provides additional flexibility and increases the scope for customization and adaptation at the receiving end. However, this approach can also have drawbacks. For example, there is a trade-off between flexibility and audio quality and complexity. It is generally useful to allow content providers to retain some control over the rendering on the client side by constraining the degrees of freedom. This not only aids the rendering and results in more realistic rendered audio scenes, but also allows content providers to retain some control over the experience provided to the user. For example, it prevents the renderer from generating audio scenes that are unrealistic and may have adverse effects on the content and the content provider.

[0011] It is envisaged that encoded audio items can be supplemented with metadata that constrain how renderers are allowed to render the audio items. This allows for a better trade-off between different requirements in many situations. However, it may not necessarily be optimal in all situations, and may for example require an increased data rate, resulting in a reduction in flexibility and / or quality for the rendered audio scene. Summary of the Invention [Problem to be solved by the invention]

[0012] Thus, improved approaches are desirable, and in particular, approaches that allow for improved usability, improved flexibility, easier implementation, easier usability, reduced cost, reduced complexity, lower data rates, improved perceived audio quality, improved rendering control, improved trade-offs, and / or improved performance would be advantageous.

[0013] SUMMARY OF THE INVENTION Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]

[0014] According to an aspect of the present invention, there is provided an audio encoding apparatus comprising: an audio receiver for receiving a plurality of audio items representing an audio scene; a metadata receiver for receiving input presentation metadata for the plurality of audio items, the input presentation metadata describing presentation constraints for rendering the plurality of audio items, the presentation constraints constraining rendering parameters that can be adapted when rendering the plurality of audio items; an audio encoder for generating encoded audio data for the audio scene by encoding the plurality of audio items in response to the input presentation metadata; a metadata circuit for generating output presentation metadata from the input presentation metadata, the output presentation metadata comprising data for the encoded audio items that constrains ranges to which adaptable rendering parameters can be adapted when rendering the encoded audio items; and an output circuit for generating an encoded audio data stream comprising the encoded audio data and the output presentation metadata.

[0015] The present invention provides improved and / or more flexible encoding in many scenarios. This approach allows, in many embodiments, to generate coded audio data streams that offer improved quality to bitrate ratios. The coded audio data streams are generated in a way that allows some flexibility in rendering, while also allowing some control of rendering from the source / decoder side.

[0016] Presentation metadata for an audio item constrains spatial and / or loudness parameters for the rendering of the audio item, including, for example, constraining the rendering position, gain level, signal level, spatial distribution, or reverberation characteristics.

[0017] The audio encoder is configured to adapt the encoding of the audio item based on the input presentation metadata, and in particular based on the input presentation metadata for the audio item. This adaptation adapts the bit / data (rate) compression for decoding the audio item. The bit rate resulting from encoding the audio item is adapted based on the input presentation metadata.

[0018] The input presentation metadata describes presentation / rendering constraints for the received plurality of audio items. The encoded audio data comprises audio data for the plurality of encoded audio items. The plurality of encoded audio items are generated by encoding the received plurality of audio items. The output presentation metadata describes presentation / rendering constraints for the rendering of the plurality of encoded audio items.

[0019] A presentation constraint may be a rendering constraint, which constrains the rendering parameters for an audio item. The rendering parameters are parameters of the rendering process and / or characteristics of the rendered signal.

[0020] Output presentation metadata is any data associated with / linked to / provided for an encoded audio item produced by an audio encoder that specifically constrains the extent to which one or more adaptable / variable aspects / characteristics / parameters of the presentation / rendering may be adapted when rendering the encoded audio item.

[0021] Output presentation metadata, and in particular data for the encoded audio items that constrain the range to which adaptable rendering parameters can be adapted when rendering the encoded audio items, is generated by a metadata circuit in response to presentation constraints that constrain the rendering parameters that can be adapted when rendering the plurality of audio items.

[0022] The audio encoder generates encoded audio data (by encoding a plurality of audio items) to include a plurality of encoded audio items.

[0023] According to an optional feature of the invention, the audio encoder has a combiner for generating a synthesized audio item by combining at least a first audio item and a second audio item of the plurality of audio items in response to input presentation metadata for the first audio item and input presentation metadata for the second audio item, and the audio encoder is configured to generate synthesized audio encoding data for the first and second audio items by encoding the synthesized audio items, and to include the synthesized audio encoding data in the encoded audio data.

[0024] This provides particularly efficient coding and / or flexibility in many embodiments, and in particular provides efficient bitrate compression with reduced perceptual degradation in many embodiments.

[0025] According to an optional feature of the invention, the combiner is configured to select a first audio item and a second audio item from the plurality of audio items in response to input presentation metadata for the first audio item and the second audio item.

[0026] This provides particularly efficient coding and / or flexibility in many embodiments.

[0027] According to an optional feature of the invention, the combiner is configured to select the first audio item and the second audio item in response to determining that at least some of the input presentation metadata for the first audio item and the input presentation metadata for the second audio item satisfy a similarity criterion.

[0028] This provides for particularly efficient coding and / or flexibility in many embodiments.The similarity criterion comprises the requirement that the rendering constraints for the rendering parameters constrained by the presentation metadata satisfy the similarity criterion.

[0029] According to an optional feature of the invention, the input presentation metadata for the first audio item and the input presentation metadata for the second audio item have at least one of a gain constraint and a position constraint.

[0030] This provides a particularly efficient operation in many embodiments.

[0031] According to an optional feature of the invention, the audio encoder is further configured to generate synthesized presentation metadata for the synthesized audio item in response to the input presentation metadata for the first audio item and the input presentation metadata for the second audio item, and to include the synthesized presentation metadata in the output presentation metadata.

[0032] This provides improved operability in many embodiments, and in particular allows the encoder in many embodiments to process synthesized audio items and encoded input audio items in the same manner, without any knowledge as to whether the individual audio items are synthesized audio items or not.

[0033] According to an optional feature of the invention, the audio encoder is configured to generate at least some of the synthesized presentation metadata to reflect constraints on presentation parameters for the synthesized audio items that are determined to satisfy both constraints for the first audio item indicated by the input presentation metadata for the first audio item and constraints for the second audio item indicated by the input presentation metadata for the second audio item.

[0034] This provides improved performance in many scenarios and applications.

[0035] According to an optional feature of the invention, the audio encoder is configured to adapt compression of the first audio item in response to input presentation metadata for the second audio item.

[0036] This approach typically allows for improved compression and encoding of audio items. Compression is a reduction in bit rate, and increased compression results in a reduction in the data rate of the encoded audio items. Compression is a reduction / compression of the bit rate. The audio coding may be such that an encoded audio item, representing one or more input audio items, is represented by fewer bits than the input audio items.

[0037] According to an optional feature of the invention, the audio encoder is configured to estimate a masking effect from the second audio item to the first audio item in response to input presentation metadata for the second audio item, and to adapt compression of the first audio item in response to the masking effect.

[0038] This provides particularly efficient operation and improved performance in many embodiments.

[0039] According to an optional feature of the invention, the audio encoder is configured to estimate a masking effect from the second audio item on the first audio item in response to at least one of a gain constraint and a position constraint on the second audio item indicated by the input presentation metadata for the second audio item.

[0040] This provides particularly efficient operation and improved performance in many embodiments.

[0041] According to an optional feature of the invention, the audio encoder is further configured to adapt compression of the first audio item in response to input presentation metadata for the first audio item.

[0042] This provides particularly advantageous usability and / or performance in many embodiments.

[0043] According to an optional feature of the invention, the input presentation metadata comprises priority data for at least some of the audio items, and the audio encoder is configured to adapt compression for a first audio item in response to an indication of priority for the first audio item in the input presentation metadata.

[0044] This provides particularly advantageous usability and / or performance in many embodiments.

[0045] According to an optional feature of the invention, the audio encoder is configured to generate encoding adaptation data indicating how the encoding is adapted in response to the input presentation metadata, and to include the encoding adaptation data in the stream of encoded audio data.

[0046] This provides particularly advantageous operability and / or performance in many embodiments, in particular allowing for improved adaptation by the decoder to match the encoding process.

[0047] According to an aspect of the present invention, there is provided a method for encoding audio, the method comprising the steps of receiving a plurality of audio items representing an audio scene; receiving input presentation metadata for the plurality of audio items, the input presentation metadata describing presentation constraints for rendering the plurality of audio items, the input presentation metadata constraining rendering parameters that can be adapted when rendering the plurality of audio items; generating encoded audio data for the audio scene by encoding the plurality of audio items in response to the input presentation metadata; generating output presentation metadata from the input presentation metadata, the output presentation metadata comprising data for the encoded audio items that constrains ranges to which adaptable rendering parameters can be adapted when rendering the encoded audio items; and generating an encoded audio data stream comprising the encoded audio data and the output presentation metadata.

[0048] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter.

[0049] Embodiments of the present invention will now be described, by way of example only, with reference to the following drawings, in which: [Brief explanation of the drawings]

[0050] [Figure 1] 1 is an illustration of example elements of an audio distribution system according to some embodiments of the present invention. [Figure 2]1 is an illustration of example elements of an audio encoding device according to some embodiments of the present invention. [Figure 3] 1 is an illustration of example elements of an audio decoding device according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0051] The following description focuses on an audio encoding and decoding system that is compatible with the 3GPP® Immersive Voice and Audio Services (IVAS) codec, but it will be understood that the principles and concepts described can be used in many other applications and embodiments.

[0052] Figure 1 illustrates an example of an audio coding system, in which an audio source 101 provides audio data to an audio encoder unit 103. The audio data comprises audio data for a number of audio items representing the audio of an audio scene. The audio items are provided as different types, in particular including:

[0053] Channel-based audio items: For such audio items, 1D (mono), 2D or 3D spatial audio content is typically represented as discrete signals intended to be presented via loudspeakers at predetermined positions relative to a listener. Common loudspeaker configurations are, for example, two-channel stereo (also known as "2.0") or five channels surrounding the listener plus a low-frequency effects channel (also referred to as "5.1"). Binaural audio is also considered to be channel-based audio, consisting of two audio signal channels intended to be presented directly to each ear of the listener (usually via headphones).

[0054] Object-based audio items: In such audio items, individual audio signals are typically used to represent distinct sound sources. These sound sources often relate to real objects or people, for example participants in a conference call. The signals are typically mono, although other representations are also used. Object-based audio signals are often accompanied by metadata that describes further characteristics such as the range (spatial extent), directionality or diffuseness of the object audio.

[0055] Scene-based audio items: For such audio items, the original 2D or 3D spatial audio scene is typically represented as several audio signals related by certain spherical harmonic functions. By synthesizing these scene-based audio signals, a presentable audio signal can be constructed at any 2D or 3D position, such as the position of an actual loudspeaker in an audio playback setting. An exemplary implementation of scene-based audio is Ambisonics. Scene-based audio uses a sound field technique called "Higher Order Ambisonics" (HOA) to produce a holistic description of both live-captured sound scenes and artificially created sound scenes that are independent of the specific loudspeaker layout.

[0056] In addition to the audio data, the audio source provides presentation metadata for the audio items, which describes the presentation constraints for the rendering of the audio scene, and thus provides presentation / rendering constraints for multiple audio items.

[0057] Presentation metadata describes constraints on how the rendering of an audio item is performed by a renderer. Presentation metadata defines constraints on one or more rendering parameters / properties. The parameters / properties in particular affect the perceptual properties of the rendering of an audio item. The constraints are constraints that affect the spatial perception and / or (relative) signal levels of the audio items in a scene. Presentation metadata in particular constrains spatial and / or gain / signal level parameters for one or more audio items. This metadata is, for example, constraints on the position and / or gain for each audio item.

[0058] This metadata may, for example, describe a range or a set of allowable values ​​for one or more parameters of one or more audio items. The rendering of the audio items is free to occur within the constraints, i.e. the rendering may be such that the constrained parameter has any of the indicated allowable values, but must not be such that the constrained parameter does not have this value.

[0059] For example, the presentation metadata may describe a region and / or (relative) gain range for one or more of the audio items, such that the audio items must be rendered at a perceived position within that region and / or with a gain within that gain range.

[0060] The presentation metadata thus constrains the rendering while still allowing some flexibility to adapt and customize the local rendering.

[0061] Examples of parameter or characteristic rendering constraints provided by presentation metadata include:

[0062] Position constraints for one or more audio items, which for example define a spatial region or volume in the audio scene from which the audio items must be rendered.

[0063] A reverberation constraint for one or more audio items, which may for example define a minimum or maximum reverberation time. This constraint may, for example, ensure that the audio items are rendered with a desired degree of distraction. For example, an audio item representing general ambient background sounds may be required to be rendered with a minimum amount of reverberation, while an audio item representing a main speaker may be required to be rendered below a given threshold of reverberation.

[0064] Gain constraints. The rendering of an audio item is adapted by the renderer to be louder or quieter according to the particular preferences of the rendering process. For example, the gain for a speaker relative to ambient background sounds may in some cases be increased or decreased based on the preferences of the listener. However, gain constraints constrain how much the gain can be modified, thereby ensuring, for example, that the speaker is always audible above the ambient noise.

[0065] Loudness constraints. The rendering of an audio item can be adapted by the renderer to be louder or quieter according to the particular preferences of the rendering process. For example, the gain for participants in a conference call can in some cases be increased or decreased based on the listener's preferences. However, loudness constraints constrain how much the perceived loudness of a participant can be modified, for example, thereby ensuring that the conference chair always has sufficient loudness even in the presence of other speakers or background noise.

[0066] Dynamic range constraints. The dynamic range of an audio item can be adapted in size by the renderer, e.g. reduced so that the audio remains audible even for periods of lower levels when background noise is present at the listener's position. For example, a violin sound will automatically have greater loudness at low levels. However, dynamic range control constraints limit how much the dynamic range can be reduced, thus ensuring a sufficiently natural perception of the normal dynamics of, e.g., a violin.

[0067] Presentation metadata describing presentation constraints for the rendering of a number of audio items is in particular data that provides constraints on rendering parameters or properties that can be met when rendering the audio items (for which the presentation metadata is provided). The rendering parameters or properties are parameters / properties of the rendering operation and / or parameters / properties of the generated and rendered / presented signal and / or audio.

[0068] Input presentation metadata is specifically any data associated with / linked to / provided to an input audio item for audio encoder 205 that constrains the extent to which one or more adaptable / variable aspects / characteristics / parameters of presentation / rendering can be adapted when rendering the input audio item.

[0069] The audio encoder unit 103 is configured to generate an encoded audio data stream containing encoded audio data for an audio scene. The encoded audio data is generated by encoding an audio item (i.e., the received audio data represents an audio item). In addition, the audio encoder unit 103 generates output presentation metadata for the encoded audio item and includes this metadata in the encoded audio data stream. The output presentation metadata describes rendering constraints for the encoded audio item.

[0070] Output presentation metadata is specifically any data associated with / linked to / provided to an encoded audio item generated by audio encoder 205 that constrains the extent to which one or more adaptable / variable aspects / characteristics / parameters of presentation / rendering can be adapted when rendering the encoded audio item.

[0071] The output presentation metadata, and in particular data for the encoded audio items that constrain the range to which adaptable rendering parameters can be adapted when rendering the encoded audio items, is generated by a metadata circuit in response to (input) presentation constraints that constrain the rendering parameters that can be adapted when rendering a plurality of (input) audio items.

[0072] The audio encoder unit 103 is coupled to a transmitter 105 to which the encoded audio data stream is provided, which in this example is configured to transmit / distribute the encoded audio data stream to one or more clients which render an audio scene based on the encoded audio data stream.

[0073] In this example, the encoded audio data stream is distributed over a network 107, which is specifically or includes the Internet. The sender 105 is configured to potentially support a large number of clients simultaneously, and the audio data is typically distributed to multiple clients.

[0074] In this particular example, the encoded audio data stream is transmitted to one or more rendering devices 109. The rendering devices 109 include a receiver 111 that receives the encoded audio data stream from the network 107.

[0075] It is understood that the transmitter 105 and receiver 111 communicate in any suitable form and using any suitable communication protocols, standards, techniques, and capabilities. In this example, the transmitter 105 and receiver 111 have suitable network interface capabilities, but in other embodiments, the transmitter 105 / receiver 111 may include, for example, wireless communication capabilities, fiber optic communication capabilities, etc.

[0076] The receiver 111 is coupled to a decoder 113 to which the received encoded audio data stream is provided. The decoder 113 is configured to decode the encoded audio data stream in order to reproduce the audio item. The decoder 113 further decodes the presentation metadata from the encoded audio data stream.

[0077] The decoder 113 is coupled to a renderer 115, which is provided with decoded audio data and presentation metadata for the audio items. The renderer 115 renders the audio scene by rendering the audio items based on the received presentation metadata. The rendering by the renderer 115 is targeted to the particular playback system being used. For example, in the case of a 5.1 surround sound system, audio signals for the individual channels are generated, since binaural signals for headphone systems are generated using, for example, HRTF filters. It will be appreciated that many different possible audio rendering algorithms and techniques are known, and any suitable approach may be used without detracting from the invention.

[0078] The renderer 115 generates output audio signals for playback, particularly so that the combined playback provides the perception of an audio scene when perceived by a listener. The renderer typically processes different audio items separately and differently according to the specific characteristics of each individual audio item, and then combines the resulting signal components for each output channel. For example, in the case of an audio object audio item, signal components are generated for each output channel according to the desired position in the audio scene for the audio source corresponding to the audio object. An audio channel audio item is rendered, for example, by generating signal components for the corresponding output playback channel, or, for example, by multiple playback channels if they do not map exactly to one of the playback channels (e.g., using panning or upmixing techniques, if appropriate).

[0079] Representing an audio scene with several, typically different types of audio items, allows the renderer 115 a high degree of flexibility and adaptability in rendering the scene. This can be used, for example, by the renderer to adapt and customize the rendered audio scene. For example, the relative gain and / or position of different audio objects can be adapted, the frequency content of the audio items can be modified, the dynamic range of the audio items can be controlled, the reverberation characteristics can be changed, etc. Thus, the renderer 115 generates output where the audio scene is adapted to the specific preferences for the current application / rendering, including adaptation to the particular playback system being used and / or to the personal preferences of the listener. This approach also allows, for example, for the rendered audio scene to be efficiently and locally adapted to changes in the virtual listening position in the audio scene. For example, to support virtual reality applications, the renderer 115 can dynamically and continuously receive user position data input and adapt the rendering in response to changes in the user's indicated virtual position in the audio scene.

[0080] The renderer 115 is configured to render the audio item based on the received presentation metadata, in particular the presentation metadata indicating constraints on variable aspects / properties / parameters of the rendering of the encoded / decoded audio item, and the renderer 115 follows these constraints during rendering.

[0081] The output audio signal from the renderer 115 / rendering device 109 results from applying rendering operations to the decoded audio item generated by the decoder 113 from the received encoded audio data stream. The rendering operations have several parameters that can be adapted externally or locally and that perceptually affect (aspects of) the rendered output audio. Presentation metadata describing presentation constraints for the rendering is specifically data that limits the set (i.e., range of values ​​in the case of continuously adaptable parameters, or discrete set of values ​​in the case of enumerated parameters) to which the rendering parameters can be adapted during rendering.

[0082] 2 shows example elements of the audio encoder unit 103 in more detail. In this example, the audio encoder unit 103 comprises an audio receiver 201 that receives input audio data describing a scene. In the current example, the audio scene is represented by three different types of audio data: channel-based audio items C, object-based audio items O, and scene-based audio items S. The audio items may be provided by audio data in any suitable format, for example, providing the audio items as raw WAV files or as audio encoded according to any suitable format. Typically, the input audio items have high audio quality and data rate.

[0083] The audio encoder unit 103 further comprises a metadata receiver 203 configured to receive presentation metadata for the input audio item. As mentioned above, the presentation metadata provides constraints for the rendering of the audio item.

[0084] The audio receiver 201 and the metadata receiver 203 are coupled to an audio encoder 205 that is configured to generate encoded audio data for an audio scene by encoding the received audio items. The audio encoder 205 in this example generates, in particular, encoded audio items, i.e. audio items represented by the encoded audio data. Relative to the input audio items, the output / encoded audio items are also of different types of audio items, in particular in the particular example a channel-based audio item C', an object-based audio item O' and a scene-based audio item S'.

[0085] One, some or all of the encoded audio items may be generated by independently encoding an input audio item, i.e. the encoded audio items are encoded input audio items, however in some scenarios one or more of the encoded audio items may be generated to represent multiple input audio items, or an input audio item may be represented as / by multiple encoded audio items.

[0086] It will be understood that many encoding algorithms and techniques are known, and that any suitable algorithm, standard, and approach may be used. It will also be understood that different algorithms and techniques may be used for different audio items. For example, an audio item corresponding to music may be encoded using an AAA encoding approach, an audio item corresponding to speech may be encoded using a CELP coding approach, etc. For audio items already received in a coded format, the encoding by the audio encoder 205 may be transcoding to a different coding format or may simply be a conversion of the data rate (e.g., by modifying quantization and / or clipping levels). Typically, the encoding includes a bit rate compression, where the coded audio item is represented by fewer bits than the input audio item.

[0087] The audio encoder unit 103 further comprises a metadata circuit 207 configured to generate output presentation metadata for the encoded audio item. The presentation metadata circuit 207 is configured to generate this output presentation metadata from the received input presentation metadata. In practice, for many audio items, the output presentation metadata is identical to the input presentation metadata. For one or more audio items, the output presentation metadata is modified, as will be explained in more detail below.

[0088] The audio encoder 205 and the metadata circuit 207 are coupled to an output circuit 209 configured to generate an encoded audio data stream having encoded audio data and output presentation metadata. The output circuit 209 is specifically a bitstream packer that generates an encoded audio data stream including both the encoded audio data and the output metadata. The encoded audio data stream is generated according to a standardized format so that it can be interpreted by a range of receivers.

[0089] The output circuit 209 thus acts as a bitstream packer that accepts bitrate reduced / encoded audio items and output presentation metadata and combines them into a bitstream that can be carried over a suitable communication channel, such as a 5G network.

[0090] 3 illustrates a particular example of elements of a rendering device 109 that receives and processes the encoded audio data stream from the audio encoder unit 103. The rendering device 109 has a receiver 111 in the form of a bitstream unpacker that receives the encoded audio data stream from the audio encoder unit 103 and separates and extracts different data from the received data stream. In particular, the receiver 111 separates and extracts the individual audio data for the encoded audio items and provides these to a decoder 113.

[0091] The decoder 113 is configured to decode received encoded audio items, in particular to generate typically unencoded representations of channel, object and scene-based audio items.

[0092] For many audio items, the decoder 113 reverses the encoding performed by the audio encoder 205. For other audio items, the decoding may, for example, only partially reverse the encoding operation. For example, if the audio encoder 205 combined the audio items into a single combined audio item, the decoder 113 decodes only the combined audio item and does not fully generate the individual audio items. It will be understood that any suitable decoding algorithms and techniques may be used, depending on the particular preferences and requirements of each implementation.

[0093] The decoded audio items are provided to a renderer 115 which is configured to render an audio scene, for example by rendering the audio items as a binaural or surround sound signal as described above.

[0094] The rendering device 109 further comprises a metadata controller / circuitry 301 that is provided with presentation metadata from the receiver 111. In this example, the metadata controller 301 also receives local presentation metadata that reflects local preferences or requirements, such as individual user preferences or the characteristics of the playback system being used.

[0095] Thus, in addition to audio presentation metadata unpacked from a received bitstream, the rendering device 109 also receives local audio presentation metadata provided, for example, via one or more input interfaces. This data provides information about the context in which the audio is being presented that is not available at the encoder side, such as: - Desired presentation (loudspeaker) settings - User preferences (e.g., audio level and orientation of participant audio in a virtual meeting) - Local acoustic characteristics, e.g. room reverberation This allows the renderer to determine which environmental effects and characteristics should be applied to the audio item, such as: - local audio signals (e.g. to be taken into account when selecting gains for audio items) - the position of the listener, and - Listener's head direction

[0096] The Metadata Controller 301 merges the received metadata with the local metadata and provides it to the Renderer 115 which proceeds to render the audio item according to the constraints of the presentation metadata.

[0097] The renderer 115 synthesizes the audio items C'', O'', and S'' produced by the decoder 113 into presentable audio in a desired presentation setting (eg, binaural or surround sound).

[0098] The renderer 115 generates the audio presentation according to, among other things, the metadata received from the metadata controller 301 and the rendered audio constrained by the constraints of the received presentation metadata, i.e., constrained from the encoder side. This provides source-side / content provider control over the audio rendering and the presented audio scene, while still allowing some flexibility on the client side. This can be used, for example, to provide a service or application where the content author retains control of an immersive application that is designed to provide some limited control to the end user, etc.

[0099] More specifically, the metadata controller 301 processes the received metadata and therefore the local metadata, e.g., the suppression of an audio item, etc. The metadata controller 301 constrains the local metadata and therefore the received metadata, e.g., the range of rotation or elevation.

[0100] In some embodiments, the renderer 115 is a different device or functional entity than the rendering device 109. For example, a standard such as the envisioned 3GPP IVAS codec defines the behavior of the decoder 113 but allows the renderer 115 to be proprietary and more freely adaptable. In some embodiments, the metadata controller 301 is part of a different device or functional entity.

[0101] In such an embodiment, an external renderer is therefore required to process and interpret the decoded O", C", S" and the received presentation metadata. Rendering operations by the external renderer must still follow the constraints provided by the presentation metadata.

[0102] Presentation metadata is thus data used by the source / content provider to control rendering behavior at the client: rendering must be adapted / restricted according to the presentation metadata.

[0103] However, in addition to the presentation metadata being used to control rendering by the client-side renderer 115, the audio encoder 205 of the audio encoder unit 103 is also configured to adapt its encoding in response to input presentation metadata. The input presentation metadata is provided to the audio encoder 205, which modifies the encoding of one or more audio items based on the presentation metadata (typically for that audio item or items). The audio encoder 205 is thus an adaptable encoder that is responsive to presentation metadata received along with the audio items.

[0104] The audio encoder 205 specifically comprises an encoding circuit 211 configured to perform encoding of an audio item and an encoding adapter 213 configured to adapt the encoding by the encoding circuit 211 based on presentation metadata.

[0105] The encoding adapter 213 is configured to set the parameters of the encoding for a given audio item based on the presentation metadata for that audio item, for example it is configured to set the bitrate allocation / target for encoding, quantization levels, masking thresholds, frequency ranges, etc. based on, for example, gain ranges or position ranges indicated as acceptable for that audio item by the presentation metadata.

[0106] In many embodiments, the encoding circuit 211 is a bitrate compressor configured to encode the audio item using a reduced number of bits compared to the received input audio item. This encoding is therefore bitrate compressed, making the resulting encoded audio data stream more efficient and easier to distribute. In such embodiments, the encoding adapter 213 adapts the bitrate reduction of the encoding circuit 211 based on the presentation metadata (so as to optimize the quality of the rendered audio according to appropriate optimization criteria / algorithms).

[0107] The encoding adapter 213 performs a coding analysis process that, for example, analyzes the presentation metadata and makes decisions about how to best perform bitrate reduction for various input audio items. Examples of operations and adaptations performed by the encoding adapter 213 include: - to inform the encoding circuit 211 of the (minimum) masking level to be respected for bitrate reduction. The encoding adapter 213 has information about which audio items are presented together at which level and in which direction. This allows it to adapt the masking level for each individual audio item with the masking level currently used by the encoding. - Transforming audio items, for example moving audio objects to channel or scene-based audio. - selecting audio items for downmixing (with associated upmix parameters), where the downmix is ​​upmixed to reconstruct immersive audio at the decoder side, while ensuring that any parametric downmix coding artifacts are sufficiently masked by the various audio items presented together. - Optimizing downmixing / upmixing gain for maximum performance / minimum artifacts; It is possible to select upmixing parameters with optimal time / frequency characteristics. - lossily compositing audio items into a composite audio item that is rendered as a single audio item by the renderer 115. This takes advantage of the fact that there is no inherent need for all audio information to be available individually at the rendering side. For example, if separate adaptation of several input audio items is not allowed (e.g. they are required to be rendered at the same position), it is not necessary for the audio items to be available individually. For example, multiple input audio objects with similar orientation and gain adaptation constraints can be composited into one scene-based audio item, in which case it is still possible to adapt the gain and orientation as a whole for the scene during rendering, but previous objects have modified their relative audio levels and relative positions in the scene. - Allocating different bitrate budgets to different audio items depending on the presentation metadata for the audio items, e.g., bitrate is allocated to audio items based on the amount of unmasked information each represents.

[0108] The encoding circuit 211 then employs coding of the audio items in accordance with the coding control data generated by the encoding adapter 213. For example, the encoding circuit 211 generates reduced bitrate versions (e.g. quantized, parameterized, etc.) of some channel, object and scene-based audio items. Furthermore, due to, for example, synthesis or transformation as part of the encoding of different audio items, at least some of the encoded audio items may represent different audio information than the input audio items, i.e. there may not be a direct correspondence between the input audio items and the encoded audio items.

[0109] In some embodiments, the audio encoder 205 comprises a combiner 215 configured, inter alia, to combine multiple input audio items into one or more combined audio items. The combiner 215, inter alia, combines first and second input audio items into a single combined audio item. The combined audio item is then encoded to generate a combined encoded audio item, which is included in the encoded audio data stream, typically replacing the first and second audio items. Thus, rather than encoding the first and second audio items separately, the combiner 215 combines them into a single combined audio item, which is then included in the encoded audio data stream, while separate encoded audio items are not included for each of the first and second audio items.

[0110] The composition of the audio items is performed in response to the received presentation metadata. In many embodiments, the audio items selected for composition are selected based on the presentation metadata. For example, the encoding adapter 213 selects audio items for composition in response to criteria that include a requirement that constraints on the audio items satisfy a similarity criterion.

[0111] For example, for audio items to be combined, it may be required that the constraints for the audio items indicated by the presentation metadata must not be contradictory, i.e., it must be possible to satisfy both constraints. Thus, it may be required that the constraints indicated by the presentation metadata are not contradictory, e.g., that the constraints have at least an overlap, such that there is at least one rendering parameter that allows the rendering constraints for both (or all) audio items to be combined to be satisfied. The encoding adapter 213 requires that the presentation metadata do not describe incompatible constraints on common rendering parameters.

[0112] For example, presentation metadata may describe multiple constraints on the position of audio items in an audio scene, in which case it is required that these position constraints overlap and that there be some common allowed positions.

[0113] The selection of audio items to combine is based on the presentation metadata for those audio items. Thus, the selection of first and second audio items to combine is based on the presentation metadata for those first and second audio items. For example, as described above, it is required that the presentation metadata for the first and second audio items do not define conflicting constraints.

[0114] In some embodiments, the first and second audio items are selected as those audio items that have, for example, the most similar constraints on the same parameter, for example, audio items that have substantially the same position constraints.

[0115] Specifically, a similarity measure for the two audio items is determined to reflect the overlap between the allowable positions, e.g., the similarity measure is generated as the ratio of the volume of the area of ​​overlapping allowable positions to the sum of the volumes of the individual allowable positions for the two audio items.

[0116] As another example, multiple audio objects that satisfy a similarity criterion for a positional matching constraint can be combined into a scene-based audio item even if their respective position ranges or spatial volumes do not overlap, and although the audio sources will have fixed relative orientations to each other in the scene-based audio item (i.e., they are not individually matchable), their orientations can still be matched as a whole.

[0117] As another example, a similarity measure may be generated to reflect the size of the overlapping gain range for the two audio items: the larger the common allowable gain range, the higher the similarity.

[0118] The encoding adapter 213 may evaluate such similarity measures for different pairs of audio items and, for example, select pairs with a similarity measure higher than a given threshold, which are then combined into a single combined audio item.

[0119] In many embodiments, the encoding adapter 213 is further configured to generate synthesized presentation metadata for the synthesized audio item from the input presentation metadata, which is then provided to the bitstream packer 209, which includes it in the output encoded audio data stream.

[0120] The metadata circuit 207, among other things, generates synthesized presentation metadata that is linked to the synthesized audio item and provides rendering constraints for the synthesized presentation metadata. The generated synthesized audio item, along with its associated synthesized presentation metadata, is then processed as any other audio item, although in fact the client / decoder / renderer is not even aware that the synthesized audio item is in fact produced by synthesis of an input audio item by the audio encoder 205. Rather, the synthesized audio item and associated presentation metadata are indistinguishable from the input audio item and associated presentation metadata to the client, and are rendered as any other audio item.

[0121] In many embodiments, the synthesized presentation metadata is generated to reflect, for example, constraints on presentation parameters for synthesized audio items, which are determined to satisfy the individual constraints for the audio items being synthesized as indicated by the input presentation metadata for those audio items. Specifically, the synthesized audio item constraints for a first and a second audio item are determined as constraints that satisfy both the constraints for the first audio item indicated by the input presentation metadata for the first audio item and the constraints for the second audio item indicated by the input presentation metadata for the second audio item. Thus, the synthesized presentation metadata is generated to provide one or more constraints that ensure that the individual constraints for the individual audio items are satisfied if the synthesized constraints are satisfied.

[0122] For example, if the first audio item is an audio object, the input presentation metadata may indicate that it should be rendered at a position within a coordinate volume (azimuth, elevation, radius) of, say, ([0,100],[-40,60],[0.5,1.5]) with a relative gain ranging from, say, -6 dB to 0 dB. If the second audio item is an audio object, the input presentation metadata may indicate that it should be rendered at a position within a coordinate volume (azimuth, elevation, radius) of, say, ([-100,80],[-20,70],[0.2,1.0]) with a relative gain ranging from, say, -3 dB to 3 dB. In this case, synthesized presentation metadata is generated to indicate that the synthesized audio item, which is an audio object, should be rendered at a position within a coordinate volume of (azimuth, elevation, radius), e.g., ([0,80],[-20,60],[-0.5,1.0]), with a relative gain in the range of e.g., -3 dB to 0 dB, thereby ensuring that the synthesized audio item is rendered acceptable relative to both the first and second audio items.

[0123] In some embodiments, the audio encoder 205 is configured to adapt the compression of one audio item based on presentation metadata for another audio item.

[0124] As a less complex example, the compression of one audio item may depend on its proximity to another audio item and its gain / level. For example, if the presentation metadata for the current audio item indicates a position range and a level range, this is compared with the position range and level range for the second audio item. If the second audio item is constrained to be positioned close to the first audio item and constrained to be rendered at a significantly higher level than the first audio item, the first audio item may be perceived only slightly by the listener. Therefore, encoding the first audio item will involve a higher compression / bitrate reduction than if the other audio items were not present. Specifically, the bitrate allocation for encoding the first audio item depends on the distance to one or more other audio items and their levels.

[0125] In some embodiments, the encoding adapter 213 is configured to estimate the masking effect of the second audio item on the first audio item. The masking effect is represented by a masking measure that indicates the degree of masking that results from rendering of the second audio item on the first audio item. The masking measure thus indicates the perceptual importance of the first audio item in the presence of the second audio item.

[0126] The masking measure is generated specifically as an indication of the audio level received from the second audio item relative to the audio level received from the first audio item when the second audio item is rendered in accordance with the constraints indicated by the presentation metadata.

[0127] For example, the masking effect of a first audio item at its lowest gain on a second audio item at its highest gain is taken to estimate the masking level of the second audio item, and vice versa.

[0128] As another example, the furthest (or, for example, average) distance between a first audio item and a second audio item is determined and the attenuation between them is estimated, and then the masking effect can be estimated based on the relative level difference after compensation for the attenuation.

[0129] As another example, if the system uses a nominal listening position, the signal levels at the listening position from each of the first and second audio items are determined based on their relative gain or signal levels and the difference in attenuation from the position of the sound source, and the positions of the audio items are selected from the allowable positions (closest allowable position for the first audio item, furthest position for the second audio item) so that, for example, masking effects are minimized.

[0130] In this way, the encoding adapter 213 estimates the masking effect from the second audio item to the first audio item based on the gain / level and position constraints for the second audio item indicated by the input presentation metadata for the second audio item, and in many cases also based on the gain / level and position constraints for the first audio item indicated by the input presentation metadata for the first audio item.

[0131] In some embodiments, the encoding adapter 213 determines the masking threshold for the first audio item directly based on the presentation metadata for the second audio item, and the encoding circuit 211 subsequently encodes the first audio item using the determined masking threshold.

[0132] In some embodiments, the adaptation of encoding by audio encoder 205 is an internal process without other functions being adapted accordingly, for example, the lossy composition of multiple audio items into a single synthesized audio item is performed without the synthesized audio items being included in the encoded audio data stream and without any indication of how the synthesized audio items are created, i.e., without a rendering device performing any specific processing of the synthesized audio items.

[0133] However, in many embodiments, the audio encoder 205 generates encoding adaptation data that indicates how the encoding is adapted in response to the input presentation metadata. This encoding adaptation data is then included in the encoded audio data stream. Thus, in this approach, the rendering device 109 has information about the encoding adaptation and is configured to adapt its decoding and / or rendering accordingly.

[0134] For example, the audio encoder 205 generates data indicating which audio items in the acoustic environment data are actually synthesized audio items, which in turn indicates some parameters of the synthesis, which in many embodiments actually enable the rendering device 109 to generate a representation of the original synthesized audio items. In fact, in some embodiments, the synthesized audio items are generated as a downmix of the input audio items, and the audio encoder 205 generates parametric upmix data and includes this in the encoded audio data stream, thereby enabling the rendering device to perform a rational upmixing.

[0135] As another example, the decoding is not adapted per se, but the information is used for interaction with the listener / end user. For example, multiple audio objects whose adaptation constraints are considered to be close may be combined by the encoder into a single scene-based audio item, while their existence as "virtual objects" is announced to the decoder in the encoded adaptation data. The user would then be given this information and would be offered to manually control the "virtual sound sources" (only as a whole, since they have been combined as scene-based audio items), rather than being informed / knowing about scene-based audio items as carriers of virtual objects.

[0136] In some embodiments, the presentation metadata includes priority data for one or more audio items, and the audio encoder 205 is configured to adapt compression for a first audio item in response to an indication of priority for the first audio item.

[0137] A priority indication is a rendering priority indication that indicates the perceptual significance or importance of an audio item in an audio scene, e.g., it may be used to indicate that an audio item representing a main speaker is more significant than, for example, an audio item representing birdsong in the background.

[0138] The renderer 115 adapts the rendering based on the priority indications. For example, for a listener with poor hearing, the renderer 115 can make the speech more intelligible by increasing the gain for the high priority main dialogue relative to the low priority background noise.

[0139] Additionally, audio encoder 205 may increase compression to lower priorities. For example, it may be required that audio items must have a priority level lower than a given level in order to be synthesized. As another example, audio encoder 205 may synthesize all audio items with a priority level lower than a given level.

[0140] In some embodiments, the bit allocation for each audio item depends on the level of priority. For example, the bit allocation for different audio items may be based on an algorithm or formula that takes into account multiple parameters, including priority. It may also be that the bit allocation for a given audio item increases monotonically with increasing priority.

[0141] It will be understood that in the above description, for clarity, embodiments of the invention have been described with reference to different functional circuits, units, and processors. However, it will be apparent that any suitable distribution of functionality between different functional circuits, units, or processors may be used without detracting from the invention. For example, functionality illustrated as being performed by separate processors or controllers may be performed by the same processor or controller. References to specific functional units or circuits should therefore be seen as references to suitable means for providing the described functionality, rather than to indicate a strict logical or physical structure or organization.

[0142] The invention may be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention is optionally implemented, at least partly, as computer software running on one or more data processors and / or digital signal processors. The elements and components of an embodiment of the invention may be physically, functionally and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in multiple units or as part of other functional units. Thus, the invention may be implemented in a single unit, or may be physically and functionally distributed between different units, circuits and processors.

[0143] Generally, examples of an audio encoding device, a method for encoding audio, and a computer program product implementing the method are illustrated by the following embodiments.

[0144] 1. an audio receiver (201) for receiving a plurality of audio items representing an audio scene; a metadata receiver (203) for receiving input presentation metadata for a plurality of audio items, the input presentation metadata describing presentation constraints for rendering the plurality of audio items; an audio encoder (205) for generating encoded audio data for an audio scene by encoding a plurality of audio items in response to input presentation metadata; a metadata circuit (207) for generating output presentation metadata from input presentation metadata; an output circuit (209) for generating an encoded audio data stream having the encoded audio data and the output presentation metadata; An audio encoding device comprising:

[0145] 2. The audio encoding device described in above 1, wherein the audio encoder (205) has a combiner (215) for generating a synthesized audio item by synthesizing at least a first audio item and a second audio item among the plurality of audio items in response to input presentation metadata for the first audio item and input presentation metadata for the second audio item, and the audio encoder (205) is configured to generate synthesized audio encoding data for the first and second audio items by encoding the synthesized audio items and include the synthesized audio encoding data in the encoded audio data.

[0146] 3. An audio encoding device as described in claim 2, wherein the combiner (215) is configured to select a first audio item and a second audio item from a plurality of audio items in response to input presentation metadata for the first audio item and the second audio item.

[0147] 4. An audio encoding device as described in 2 or 3 above, wherein the combiner (215) is configured to select the first audio item and the second audio item in response to determining that at least some of the input presentation metadata for the first audio item and the input presentation metadata for the second audio item satisfy a similarity criterion.

[0148] 5. An audio encoding device as described in any one of 2 to 4 above, wherein the input presentation metadata for the first audio item and the input presentation metadata for the second audio item have at least one of a gain constraint and a position constraint.

[0149] 6. An audio decoding device described in any of 2 to 5 above, wherein the audio encoder (205) is further configured to generate synthesized presentation metadata for the synthesized audio item in response to input presentation metadata for the first audio item and input presentation metadata for the second audio item, and include the synthesized presentation metadata in the output presentation metadata.

[0150] 7. An audio encoding device as described in claim 6, wherein the audio encoder (205) is configured to generate at least some synthesized presentation metadata to reflect constraints on presentation parameters for the synthesized audio item that are determined to satisfy both constraints on the first audio item indicated by the input presentation metadata for the first audio item and constraints on the second audio item indicated by the input presentation metadata for the second audio item.

[0151] 8. An audio encoding device according to any one of claims 1 to 7, wherein the audio encoder (205) is configured to adapt compression of a first audio item in response to input presentation metadata for a second audio item.

[0152] 9. An audio encoding device as described in claim 8, wherein the audio encoder (205) is configured to estimate a masking effect from the second audio item to the first audio item in response to input presentation metadata for the second audio item, and to adapt compression of the first audio item in response to the masking effect.

[0153] 10. An audio encoding device as described in claim 9, wherein the audio encoder (205) is configured to estimate a masking effect from the second audio item to the first audio item in response to at least one of a gain constraint and a position constraint for the second audio item indicated by input presentation metadata for the second audio item.

[0154] 11. An audio encoding device described in any of claims 8 to 10, wherein the audio encoder (205) is further configured to adapt compression of the first audio item in response to input presentation metadata for the first audio item.

[0155] 12. An audio encoding device described in any of claims 1 to 11, wherein the input presentation metadata includes priority data for at least some audio items, and the audio encoder is configured to adapt compression for a first audio item in response to an indication of priority for the first audio item in the input presentation metadata.

[0156] 13. An audio encoding device described in any of 1 to 12 above, wherein the audio encoder (205) is configured to generate encoding adaptation data indicating how the encoding is adapted in response to the input presentation metadata, and to include the encoding adaptation data in the stream of encoded audio data.

[0157] 14. Receiving a plurality of audio items representing an audio scene; receiving input presentation metadata for a plurality of audio items, the input presentation metadata describing presentation constraints for rendering the plurality of audio items; generating encoded audio data for an audio scene by encoding a plurality of audio items in response to input presentation metadata; generating output presentation metadata from the input presentation metadata; generating an encoded audio data stream having encoded audio data and output presentation metadata; 2. A method for encoding audio having:

[0158] 15. A computer program product having computer program code means adapted to perform all the steps of the method set out in 14 above when the program is run on a computer.

[0159] More particularly, the present invention is defined by the following claims.

[0160] While the present invention has been described above in connection with several embodiments, it is not intended that the present invention be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while certain features may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0161] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Moreover, although individual features may be included in different claims, they may, in some cases, be advantageously combined, and their inclusion in different claims does not indicate that a combination of features is not feasible and / or advantageous. Furthermore, the inclusion of a feature in one category of claims does not imply a limitation to that category, but rather indicates that the feature is equally applicable to other claim categories, as appropriate. Furthermore, the order of features in the claims does not imply any particular order in which those features must function, and in particular the order of individual steps in method claims does not imply that those steps must be performed in that order. Rather, those steps may be performed in any suitable order. Furthermore, a reference to the singular does not exclude a plurality. Thus, reference to the singular terms "first," "second," etc. does not exclude a plurality. Reference signs in the claims are provided merely as a clarifying example and should not be construed as limiting the scope of the claims in any way.

Claims

1. an audio receiver for receiving a plurality of audio items representing an audio scene; - a metadata receiver for receiving input presentation metadata for said plurality of audio items, said input presentation metadata describing presentation constraints for rendering of said plurality of audio items, said presentation constraints constraining rendering parameters that can be adapted when rendering said plurality of audio items; an audio encoder for generating encoded audio data for the audio scene by encoding the plurality of audio items in response to the input presentation metadata; a metadata circuit for generating output presentation metadata from the input presentation metadata, the output presentation metadata comprising data for an encoded audio item that constrains ranges within which rendering adaptable parameters can be adapted when rendering the encoded audio item; and an output circuit for generating an encoded audio data stream having the encoded audio data and the output presentation metadata; An audio encoding device comprising:

2. 2. The audio encoding device of claim 1, wherein the audio encoder comprises a combiner for generating a synthesized audio item by combining at least a first audio item and a second audio item among the plurality of audio items in response to input presentation metadata for the first audio item and input presentation metadata for the second audio item, and the audio encoder generates synthesized audio encoding data for the first and second audio items by encoding the synthesized audio items and includes the synthesized audio encoding data in the encoded audio data.

3. 3. The audio encoding device of claim 2, wherein the combiner selects the first audio item and the second audio item from the plurality of audio items in response to the input presentation metadata for the first audio item and the second audio item.

4. 4. The audio encoding device of claim 2, wherein the combiner selects the first audio item and the second audio item in response to determining that at least some of the input presentation metadata for the first audio item and the input presentation metadata for the second audio item satisfy a similarity criterion.

5. 5. An audio encoding device according to claim 2, wherein the input presentation metadata for the first audio item and the input presentation metadata for the second audio item have at least one of a gain constraint and a position constraint.

6. 6. The audio encoding device of claim 2, wherein the audio encoder is further configured to generate synthesized presentation metadata for the synthesized audio item in response to the input presentation metadata for the first audio item and the input presentation metadata for the second audio item, and to include the synthesized presentation metadata in the output presentation metadata.

7. 7. The audio encoding device of claim 6, wherein the audio encoder generates at least some synthesized presentation metadata to reflect constraints on presentation parameters for the synthesized audio item that are determined to satisfy both constraints for the first audio item indicated by input presentation metadata for the first audio item and constraints for the second audio item indicated by input presentation metadata for the second audio item.

8. 8. Audio encoding apparatus according to any one of claims 1 to 7, wherein the audio encoder is responsive to input presentation metadata for a second audio item to adapt compression of a first audio item.

9. 9. The audio encoding device of claim 8, wherein the audio encoder estimates a masking effect from the second audio item to the first audio item in response to input presentation metadata for the second audio item, and adapts the compression of the first audio item in response to the masking effect.

10. 10. The audio encoding apparatus of claim 9, wherein the audio encoder estimates the masking effect from the second audio item to the first audio item in response to at least one of a gain constraint and a position constraint for the second audio item indicated by the input presentation metadata for the second audio item.

11. 11. An audio encoding apparatus according to claim 8, wherein the audio encoder is further configured to adapt the compression of the first audio item in response to input presentation metadata for the first audio item.

12. 12. An audio encoding apparatus according to claim 1, wherein the input presentation metadata comprises priority data for at least some audio items, and wherein the audio encoder adapts compression for a first audio item in response to an indication of priority for the first audio item in the input presentation metadata.

13. 13. The audio encoding device of claim 1, wherein the audio encoder generates encoding adaptation data indicating how the encoding is adapted in response to the input presentation metadata, and includes the encoding adaptation data in the stream of encoded audio data.

14. receiving a plurality of audio items representing an audio scene; receiving input presentation metadata for the plurality of audio items, the input presentation metadata describing presentation constraints for rendering the plurality of audio items, the presentation constraints constraining rendering parameters that can be adapted when rendering the plurality of audio items; generating encoded audio data for the audio scene by encoding the plurality of audio items in response to the input presentation metadata; generating output presentation metadata from the input presentation metadata, the output presentation metadata comprising data for an encoded audio item that constrains the range to which rendering adaptable parameters can be adapted when rendering the encoded audio item; generating an encoded audio data stream comprising the encoded audio data and the output presentation metadata; 1. A method for encoding audio, comprising:

15. 15. A computer program having computer program code means adapted to perform all the steps of the method of claim 14 when said program is run on a computer.