Apparatus and method for audio coding

Through the audio encoding device and method, combining the combiner and metadata circuit to generate an encoded audio data stream, the problem of poor flexibility and quality trade-offs in the prior art is solved, flexible audio encoding and rendering control is realized, and audio quality and data efficiency are improved.

CN114600188BActive Publication Date: 2025-07-08KONINKLIJKE PHILIPS NV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080072214.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-14
Filing Date
2020-10-08
Publication Date
2025-07-08
Estimated Expiration
2040-10-08

AI Technical Summary

Technical Problem

Existing audio encoding methods are difficult to achieve a good trade-off between flexibility and audio quality and complexity, resulting in insufficient rendering control and may reduce the quality and flexibility of the audio scene.

Method used

By adopting an audio encoding device and method, by receiving the audio items and their corresponding rendering constraint metadata, an encoded audio data stream is generated, and combined with a combiner and metadata circuit during the encoding process, the output presentation metadata is generated to control the adjustment of rendering parameters, and flexible audio encoding and rendering is realized.

Benefits of technology

Improved audio quality to bitrate ratios are provided, allowing for some flexibility in rendering while maintaining source-side control over rendering, reducing perceived degradation and data rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114600188B_ABST
    Figure CN114600188B_ABST
Patent Text Reader

Abstract

An audio coding apparatus includes an audio receiver (201) that receives an audio item representing an audio scene, and a metadata receiver (203) that receives input presentation metadata for the audio item that describes presentation constraints for rendering of the audio item. The presentation constraints constrain rendering parameters that can be adjusted when rendering the audio item. An audio encoder (205) generates coded audio data for the audio scene by coding a plurality of audio items, wherein the coding is adjusted in response to the input presentation metadata. A metadata circuit (207) generates output presentation metadata based on the input presentation metadata. The output presentation metadata includes data for the coded audio item that constrains the extent to which adjustable parameters of rendering can be adjusted when rendering the coded audio item. An output (209) generates a coded audio data stream that includes the coded audio data and the output presentation metadata.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to apparatus and methods for audio coding, and in particular but not exclusively, apparatus and methods for coding audio for dynamic applications such as virtual reality applications. Background Art

[0002] In recent years, the variety and scope of audio and video applications have substantially increased, and new services as well as ways of utilizing and consuming audio, images, and video have been continuously developed and introduced.

[0003] For example, a service that has become increasingly popular is to provide audio and images in such a way that an observer can actively and dynamically interact with the system to change rendering parameters. A very attractive feature in many applications is the ability to change the effective viewing / listening position. Such a feature can specifically allow a virtual reality experience to be provided to a user.

[0004] The trend is to provide increasing flexibility, which allows for rendering-side adjustment of the scene. To provide increased rendering-side flexibility for the rendering of an audio scene, multiple audio coding and distribution methods have been proposed, in which the audio scene can be represented by a combination of different audio items. For example, an audio item can represent a separate sound source, such as a particular speaker, etc. In some methods, all audio items are of the same type, but an increasing number of systems are being developed that allow the simultaneous use and support of multiple different audio types. For example, some audio items can be audio channels, others can be individual audio objects, and still others can be scene-based, such as Ambisonic audio items. In many systems, metadata can be provided together with the audio data representing the audio items. Such metadata can, for example, indicate the nominal position in the scene of the audio source for the audio item.

[0005] Such methods can achieve a high degree of client / rendering-side customization and adjustment. For example, the audio scene can be locally adapted to changes in the virtual position of a listener in the audio scene or to the specific preferences of an individual listener.

[0006] As a specific example, the 3GPP consortium is currently developing a so-called Immersive Voice and Audio Service (IVAS) codec. This codec will be able to encode audio content in various configurations, such as channel-based, object-based, or scene-based (especially Ambisonic) configurations. The aim of the encoding is to convey the audio information with the smallest amount of data.

[0007] The IVAS codec should also include a renderer for converting the various audio streams into a form suitable for reproduction at the receiving end. For example, the audio can be mapped into a known speaker configuration, or the audio can be rendered in a binaural format for reproduction via headphones.

[0008] Within the scope of the 3GPP IVAS codec, work is underway to collect potential use cases. For these, it is envisioned that the codec should provide interactivity to modulate rendering. For example, headphone audio may have to be rendered independently of head position and translation, which means that head movement has to be compensated for. As another example, it may be made possible for the user to spatially localize audio items, such as (re)positioning an object carrying the audio of a participant in a virtual meeting.

[0009] The renderer is considered part of the 3GPP IVAS codec work item and is considered to be within the IVAS codec. However, it has been proposed that the codec also includes a pass-through mode. This mode will allow audio items (i.e., as channel, object- and scene-based audio items in a 1:1 correspondence) to be represented at the decoder output in the same (one or more) configuration as they were input at its encoder input. Via a dedicated external rendering interface, an external renderer can have access to these items and can implement an alternative rendering to the internal IVAS renderer.

[0010] This approach can provide additional flexibility and increase the scope of customization and adaptation at the receiving end. However, this approach can also have associated drawbacks. For example, there is a trade-off between flexibility and audio quality and complexity. It is often useful to limit the degrees of freedom to allow the content provider to retain some control over rendering at the client side. This can not only help with rendering and produce a more realistic rendered audio scene, but can also allow the content provider to retain some control over the experience being provided to the user. For example, it can prevent the renderer from generating an unrealistic audio scene that may reflect poorly on the content and the content provider.

[0011] It is envisioned that encoded audio items can be supplemented with metadata that limits how the renderer is allowed to render the audio items. In many cases, this can allow for an improved trade-off between different requirements. However, it may not be optimal in all cases and may, for example, require an increased data rate and may result in reduced flexibility and / or quality for the rendered audio scene.

[0012] Accordingly, an improved method would be desirable. In particular, a method that allows for improved operation, increased flexibility, convenient implementation, convenient operation, reduced cost, reduced complexity, reduced data rate, improved perceived audio quality, improved rendering control, improved trade-off and / or improved performance would be advantageous. SUMMARY OF THE INVENTION

[0013] Accordingly, the present invention seeks to alleviate, mitigate or eliminate, preferably singly or in any combination, one or more of the above disadvantages.

[0014] According to one aspect of the present invention, there is provided an audio encoding apparatus, comprising: an audio receiver for receiving a plurality of audio items representing an audio scene; a metadata receiver for receiving input rendering metadata for the plurality of audio items, the input rendering metadata describing rendering constraints for rendering the plurality of audio items, the rendering constraints constraining rendering parameters that can be adjusted when rendering the plurality of audio items, the rendering constraints including at least one constraint from the group having: a reverberation constraint; a gain constraint; and a dynamic range control constraint; an audio encoder for generating encoded audio data for the audio scene by encoding the plurality of audio items, the encoding including a combiner (215) in response to the input rendering metadata, the combiner for generating a combined audio item by combining the first audio item and the second audio item in response to at least the input rendering metadata for the first audio item and the input rendering metadata for the second audio item among the plurality of audio items, and the audio encoder (205) being arranged to generate combined audio encoded data for the first audio item and the second audio item by encoding the combined audio item and including the combined audio encoded data in the encoded audio data; a metadata circuit for generating output rendering metadata based on the input rendering metadata, the output rendering metadata including data for the encoded audio item, the data constraining the degree to which adjustable parameters of rendering can be adjusted when rendering the encoded audio item; and an output circuit for generating an encoded audio data stream including the encoded audio data and the output rendering metadata.

[0015] The present invention can provide improved and / or more flexible encoding in many scenarios. In many embodiments, the method can allow the generation of an encoded audio data stream that provides an improved quality-to-bitrate ratio. An encoded audio data stream can be generated to allow some flexibility in rendering while also allowing some control over rendering from the source / encoding side.

[0016] The rendering metadata for an audio item can constrain at least one of the spatial parameters and volume parameters for rendering the audio item, including, for example, constraining the rendering position, gain level, signal level, spatial distribution, or reverberation properties.

[0017] The audio encoder can be arranged to adjust the encoding of the audio item based on the input rendering metadata, and specifically based on the input rendering metadata for the audio item. This adjustment can adjust the bit / data (rate) compression for encoding the audio item. The bitrate resulting from encoding the audio item can be adjusted based on the input rendering metadata.

[0018] Input presentation metadata may describe presentation / rendering constraints for a plurality of received audio items. Encoded audio data may include audio data for a plurality of encoded audio items. The plurality of encoded audio items may be generated by encoding the received plurality of audio items. Output presentation metadata that describes presentation / rendering constraints for the rendering of the plurality of encoded audio items.

[0019] The presentation constraints may be rendering constraints and may constrain rendering parameters for the audio items. The rendering parameters may be parameters of the rendering process and / or the nature of the rendered signal.

[0020] The output presentation metadata may specifically be any data associated with / linked to / provided for the encoded audio items generated by an audio encoder, which constrains the extent to which one or more adjustable / variable aspects / natures / parameters of the presentation / rendering can be adjusted when rendering the encoded audio items.

[0021] The output presentation metadata, and in particular the data for the encoded audio items, which constrains the extent to which adjustable parameters of the rendering can be adjusted when rendering the encoded audio items, may be generated by a metadata circuit in response to presentation constraints that constrain the rendering parameters that can be adjusted when rendering the plurality of audio items.

[0022] The audio encoder may generate encoded audio data to include a plurality of encoded audio items (by encoding a plurality of audio items).

[0023] The audio encoder includes a combiner that is configured to generate a combined audio item by combining at least the first audio item and the second audio item in response to input presentation metadata for the first audio item among the plurality of audio items and input presentation metadata for the second audio item among the plurality of audio items, and the audio encoder is arranged to generate combined audio encoding data for the first audio item and the second audio item by encoding the combined audio item, and include the combined audio encoding data in the encoded audio data.

[0024] This may provide particularly effective encoding and / or flexibility in many embodiments. In many embodiments, it may particularly provide effective bitrate compression with reduced perceptual degradation.

[0025] According to an optional feature of the present invention, the combiner is arranged to select the first audio item and the second audio item from the plurality of audio items in response to the input presentation metadata for the first audio item and the second audio item.

[0026] This may provide particularly effective encoding and / or flexibility in many embodiments.

[0027] According to an optional feature of the invention, the combiner is arranged to select the first audio item and the second audio item in response to a determination that at least some of the input presentation metadata for the first audio item and the input presentation metadata for the second audio item satisfy a similarity criterion.

[0028] This can provide particularly effective encoding and / or flexibility in many embodiments. The similarity criterion can include a requirement that the rendering constraints on the rendering parameters constrained by the presentation metadata satisfy a similarity criterion.

[0029] According to an optional feature of the invention, the input presentation metadata for the first audio item and the input presentation metadata for the second audio item include at least one of a gain constraint and a position constraint.

[0030] This can provide particularly effective operation in many embodiments.

[0031] According to an optional feature of the invention, the audio encoder is further arranged to generate combined presentation metadata for the combined audio item in response to the input presentation metadata for the first audio item and the input presentation metadata for the second audio item; and include the combined presentation metadata in the output presentation metadata.

[0032] This can provide improved operation in many embodiments and in particular can allow the encoder to handle the combined audio item and the encoded input audio item in the same way and effectively not know whether an individual audio item is a combined audio item.

[0033] According to an optional feature of the invention, the audio encoder is arranged to generate at least some combined presentation metadata to reflect constraints on the presentation parameters for the combined audio item, the constraints being determined to satisfy both the constraints of the first audio item indicated by the input presentation metadata for the first audio item and the constraints of the second audio item indicated by the input presentation metadata for the second audio item.

[0034] This can provide improved performance in many situations and applications.

[0035] According to an optional feature of the invention, the audio encoder is arranged to adjust the compression of the first audio item in response to the input presentation metadata for the second audio item.

[0036] This method can generally allow for improved compression and encoding of audio items. The compression can be a bitrate reduction and increasing the compression can result in a reduced data rate for the encoded audio item. The compression can be a bitrate reduction / compression. Audio encoding can enable the encoded audio item representing one or more input audio items to be represented by fewer bits than the one or more input audio items.

[0037] According to an optional feature of the present invention, the audio encoder is arranged to estimate the masking effect of the second audio item on the first audio item in response to input presentation metadata for the second audio item, and to adjust the compression of the first audio item in response to the masking effect.

[0038] This can provide particularly efficient operation and improved performance in many embodiments.

[0039] According to an optional feature of the present invention, the audio encoder is arranged to estimate the masking effect of the second audio item on the first audio item in response to at least one of a gain constraint and a position constraint of the second audio item indicated by the input presentation metadata for the second audio item.

[0040] This can provide particularly efficient operation and improved performance in many embodiments.

[0041] According to an optional feature of the present invention, the audio encoder is further arranged to adjust the compression of the first audio item in response to input presentation metadata for the first audio item.

[0042] This can provide particularly advantageous operation and / or performance in many embodiments.

[0043] According to an optional feature of the present invention, the input presentation metadata includes priority data for at least some audio items, and the encoder is arranged to adjust the compression for the first audio item in response to a priority indication for the first audio item in the input presentation metadata.

[0044] This can provide particularly advantageous operation and / or performance in many embodiments.

[0045] According to an optional feature of the present invention, the audio encoder is arranged to generate encoding adjustment data indicating how to adjust the encoding in response to the input presentation metadata, and to include the encoding adjustment data in the encoded audio data stream.

[0046] This can provide particularly advantageous operation and / or performance in many embodiments. In particular, it can allow for improved adjustment by the decoder to match the encoding process.

[0047] According to one aspect of the present invention, there is provided a method for encoding audio, the method comprising: receiving a plurality of audio items representing an audio scene; receiving input presentation metadata for the plurality of audio items, the input presentation metadata describing presentation constraints for rendering the plurality of audio items, the presentation constraints constraining rendering parameters that can be adjusted when rendering the audio items, the presentation constraints including at least one constraint from a group having: a reverb constraint; a gain constraint; and a dynamic range control constraint; generating encoded audio data for the audio scene by encoding the plurality of audio items, the encoding being responsive to the input presentation metadata by: generating a combined audio item by combining a first audio item and a second audio item in the plurality of audio items at least in response to the input presentation metadata for the first audio item and the input presentation metadata for the second audio item in the plurality of audio items, and generating combined audio encoded data for the first audio item and the second audio item by encoding the combined audio item, and including the combined audio encoded data in the encoded audio data; generating output presentation metadata according to the input presentation metadata, the output presentation metadata including data for the encoded audio items, the data constraining the degree to which adjustable parameters of rendering can be adjusted when rendering the encoded audio items; and generating an encoded audio data stream including the encoded audio data and the output presentation metadata.

[0048] These and other aspects, features and advantages of the present invention will be apparent from and elucidated with reference to the (one or more) embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Embodiments of the present invention will be described by way of example only with reference to the accompanying drawings, in which

[0050] Figure 1 illustrates an example of elements of an audio distribution system according to some embodiments of the present invention;

[0051] Figure 2 illustrates an example of elements of an audio encoding apparatus according to some embodiments of the present invention; and

[0052] Figure 3 illustrates an example of elements of an audio decoding apparatus according to some embodiments of the present invention. DETAILED DESCRIPTION

[0053] The following description will focus on audio encoding and decoding systems that can be compatible with 3GPP immersive voice and audio service (IVAS) codecs, but it will be appreciated that the principles and concepts described can be used in many other applications and embodiments.

[0054] Figure 1 An example of an audio encoding system is illustrated. In this system, an audio source 101 provides audio data to an audio encoder unit 103. The audio data includes audio data for a plurality of audio items representing audio for an audio scene. The audio items can be provided in different types, specifically including:

[0055] Channel-based audio items: For such audio items, 1D (mono), 2D or 3D spatial audio content is typically represented as a discrete signal, intended to be presented via speakers at a predefined position relative to a listener. Well-known speaker setups are for example two-channel stereo (also known as "2.0"), or 5 channels around the listener plus a low-frequency effects channel (also known as "5.1"). Moreover, binaural audio is typically considered to be channel-based audio, including two audio signal channels, which are intended to be presented directly to the respective ears of the listener (usually via headphones).

[0056] Object-based audio items: For such audio items, individual audio signals are typically used to represent different sound sources. These sound sources are often related to actual objects or people, such as participants in a conference call. The signals are typically mono, but other representations can also be used. Object-based audio signals are often accompanied by metadata describing additional properties, such as the extent (spatial extent) of the object audio, directionality or diffusivity.

[0057] Scene-based audio items: For such audio items, the original 2D or 3D spatial audio scene is typically represented as a plurality of audio signals related to certain spherical harmonic functions. By combining these scene-based audio signals, an audio signal can be presented that can be constructed at any 2D or 3D position, for example at the position of the actual speakers in an audio reproduction setup. An example implementation of scene-based audio is multichannel analog stereo. Scene-based audio uses a sound field technique called "higher-order multichannel analog stereo" (HOA) to create an overall description of both live capture and artistic creation of sound scenes independent of a specific speaker layout.

[0058] In addition to the audio data, the audio source can also provide presentation metadata for the audio items. The presentation metadata can describe presentation constraints for the rendering of the audio scene and can thus provide presentation / rendering constraints for a plurality of audio items.

[0059] Presentation metadata may describe constraints on how an audio item is to be rendered by a renderer. The presentation metadata may define constraints on one or more rendering parameters / properties. The parameter / property may specifically be a parameter / property that affects the perceived properties of the rendering of the audio item. The constraints may be constraints that affect the spatial perception and / or (relative) signal level of the audio item in a scene. The presentation metadata may specifically constrain the spatial and / or gain / signal level parameters for one or more audio items. The metadata may for example be a constraint on the position and / or gain of each audio item.

[0060] The metadata may for example describe a range or a set of allowable values for one or more parameters for one or more audio items. The rendering of the (one or more) audio items may be freely performed within the constraints, i.e., the rendering may be such that the constrained parameters have any of the indicated allowable values, but may not be such that the constrained parameters do not have that value.

[0061] As an example, the presentation metadata may describe an area and / or (relative) gain range for one or more of the audio items. The audio item must then be rendered with a perceived position within the area and / or with a gain within the gain range.

[0062] The presentation metadata can thus constrain the rendering while still allowing some flexibility to adjust and customize the local rendering.

[0063] Examples of rendering constraints on parameters or properties that may be provided by the presentation metadata include:

[0064] Position constraints for one or more audio items. For example, this may define the spatial area or volume from which the audio item must be rendered in the audio scene.

[0065] Reverb constraints for one or more audio items. This may for example define a minimum or maximum reverb time. The constraint may for example ensure that the audio item is rendered with a desired diffuseness. For example, an audio item representing general ambient background sound may need to be rendered with a minimum amount of reverb, while an audio item representing a main speaker may need to be rendered with less than a given reverb threshold.

[0066] Gain constraints. The renderer may adjust the rendering of the audio item to be louder or quieter according to specific preferences of the rendering process. For example, the gain of a speaker relative to the ambient background sound may be increased or decreased in some cases based on listener preferences. However, the gain constraint can constrain how much the gain can be modified, for example to ensure that the speaker can always be heard over the ambient noise.

[0067] Loudness constraint. The rendering of an audio item can be adjusted by a renderer to be louder or quieter according to specific preferences of the rendering process. For example, in some cases, the gain of a conference call participant can be increased or decreased based on the preferences of the listener. However, a loudness constraint can restrict how much the perceived loudness of certain participants can be modified, for example to ensure that, for instance, in the presence of other speakers or background noise, the voice of the conference chairperson is always loud enough.

[0068] Dynamic range control constraint. The dynamic range of an audio item can be adjusted by a renderer to be louder, for example it may be reduced such that the audio remains audible during periods of lower levels in the presence of background noise at the listener's location. For example, the sound of a violin can be made louder automatically at lower levels. However, a dynamic range control constraint can restrict how much the dynamic range can be reduced, thus ensuring, for example, a sufficiently natural perception of the normal dynamics of the violin.

[0069] Presentation metadata that describes presentation constraints for the rendering of multiple audio items can specifically be data that provides constraints on rendering parameters or properties that can be adjusted when rendering an audio item (for which it provides the presentation metadata). The rendering parameters or properties can be parameters / properties of the rendering operation and / or the generated rendered / presented signal and / or parameters or properties of the audio.

[0070] Input presentation metadata can specifically be any data associated with / linked to / provided for an input audio item for an audio encoder 205 that constrains the extent to which one or more adjustable / variable aspects / properties / parameters of the presentation / rendering can be adjusted when rendering the input audio item.

[0071] The audio encoder unit 103 is arranged to generate an encoded audio data stream that includes encoded audio data for an audio scene. The encoded audio data is generated by encoding an audio item (i.e., the received audio data representing the audio item). In addition, the audio encoder unit 103 generates output presentation metadata for the encoded audio item and includes this metadata in the encoded audio data stream. The output presentation metadata describes the rendering constraints for the encoded audio item.

[0072] Output presentation metadata can specifically be any data associated with / linked to / provided for an encoded audio item generated by an audio encoder 205 that constrains the extent to which one or more adjustable / variable aspects / properties / parameters of the presentation / rendering can be adjusted when rendering the encoded audio item.

[0073] In response to a (input) rendering constraint, output rendering metadata can be generated by metadata circuitry, and in particular data for an encoded audio item that encodes how adjustable parameters for constrained rendering can be adjusted when rendering the encoded audio item, the (input) rendering constraint constraining rendering parameters that can be adjusted when rendering multiple (input) audio items.

[0074] The audio encoder unit 103 is coupled to a transmitter 105, and the transmitter 105 is fed an encoded audio data stream. The transmitter 105 is arranged in the example to send / distribute the encoded audio data stream to one or more clients, which can render an audio scene based on the encoded audio data stream.

[0075] In this example, the encoded audio data stream is distributed via a network 107, which can in particular be or include the Internet. The transmitter 105 can be arranged to support potentially a large number of clients simultaneously, and the audio data can generally be distributed to multiple clients.

[0076] In a particular example, the encoded audio data stream can be sent to one or more rendering devices 109. The rendering device 109 can include a receiver 111 that receives the encoded audio data stream from the network 107.

[0077] It will be appreciated that the transmitter 105 and the receiver 111 can communicate in any suitable form and using any suitable communication protocols, standards, technologies, and functions. In the example, the transmitter 105 and the receiver 111 can include appropriate network interface functionality, but it will be appreciated that in other embodiments, the transmitter 105 and / or the receiver 111 can include, for example, radio communication functionality, fiber optic communication functionality, etc.

[0078] The receiver 111 is coupled to a decoder 113, and the decoder 113 is fed the received encoded audio data stream. The decoder 113 is arranged to decode the encoded audio data stream to recreate the audio item. The decoder 113 can also decode rendering metadata from the encoded audio data stream.

[0079] The decoder 113 is coupled to a renderer 115, and the renderer 115 is fed the decoded audio data and the rendering metadata for the audio item. The renderer 115 can render an audio scene by rendering the audio item based on the received rendering metadata. The rendering by the renderer 115 can be for a particular audio reproduction system in use. For example, for a 5.1 surround sound system, audio signals for individual channels can be generated, and for a headphone system, binaural signals can be generated using, for example, HRTF filters, etc. It will be appreciated that many different possible audio rendering algorithms and techniques are known, and any suitable method can be used without departing from the present invention.

[0080] Renderer 115 can specifically generate an output audio signal for reproduction such that the combined reproduction provides a perception of an audio scene as perceived by a listener. The renderer typically processes different audio items separately and differently according to the specific characteristics of the individual audio items, and then combines the resulting signal components for each output channel. For example, for an audio object audio item, signal components can be generated for each output channel according to the desired position of the audio source corresponding to the audio object in the audio scene. An audio channel audio item can be rendered, for example, by generating signal components for the corresponding output reproduction channel, or, for example, by multiple reproduction channels if it does not map precisely to one of the reproduction channels (e.g., using panning or upmixing techniques if appropriate).

[0081] The representation of an audio scene by multiple audio items of generally different types can allow renderer 115 to have a high degree of flexibility and adaptability in the rendering of the scene. For example, this can be used by the renderer to adjust and customize the rendered audio scene. For example, the relative gain and / or position of different audio objects can be adjusted, the frequency content of an audio item can be modified, the dynamic range of an audio item can be controlled, the reverb properties can be changed, and so on. Thus, renderer 115 can generate an output in which the audio scene is adapted to the specific preferences for the current application / rendering, including adaptation to the specific reproduction system used and / or the personal preferences of the listener. For example, the method can also allow the rendered audio scene to effectively adapt locally to changes in the virtual listening position in the audio scene. For example, to support virtual reality applications, renderer 115 can dynamically and continuously receive user position data input and adjust the rendering in response to changes in the indicated virtual position of the user in the audio scene.

[0082] Renderer 115 is arranged to render audio items based on received rendering metadata. In particular, the rendering metadata can indicate constraints on variable aspects / natures / parameters of the rendering of the encoded / decoded audio items, and renderer 115 can comply with these constraints during rendering.

[0083] The output audio signal from renderer 115 / rendering device 109 is produced by a rendering operation applied to the decoded audio items generated by decoder 113 from the received encoded audio data stream. The rendering operation may have some parameters that can be adjusted externally or locally and that perceptually affect aspects of the rendered output audio. The rendering metadata that describes the rendering constraints can specifically be data that limits the set of rendering parameters that can be adjusted during rendering (i.e., for continuously adjustable parameters, a range of values, or for enumerated parameters, a set of discrete values).

[0084] Figure 2An example of the elements of the audio encoder unit 103 is shown in more detail. In this example, the audio encoder unit 103 includes an audio receiver 201 that receives input audio data describing a scene. In this example, the audio scene is represented by three different types of audio data, namely channel-based audio items C, object-based audio items O, and scene-based audio items S. The audio items are provided by audio data that can take any suitable form. The audio data can, for example, provide the audio items as raw WAV files or as audio encoded according to any suitable format. Generally, the input audio items will be of high audio quality and high data rate.

[0085] The audio encoder unit 103 also includes a metadata receiver 203 that is arranged to receive rendering metadata for the input audio items. As previously described, the rendering metadata can provide constraints on the rendering of the audio items.

[0086] The audio receiver 201 and the metadata receiver 203 are coupled to an audio encoder 205 that is arranged to generate encoded audio data for the audio scene by encoding the received audio items. The audio encoder 205 in this example specifically generates encoded audio items, that is, audio items represented by the encoded audio data. As for the input audio items, the output / encoded audio items can also be different types of audio items and can specifically be channel-based audio C’, object-based audio items O’, and scene-based audio items S’ in a specific example.

[0087] One, some, or all of the encoded audio items can be generated by independently encoding the input audio items, that is, the encoded audio items can be encoded input audio items. However, in some cases, one or more of the encoded audio items can be generated to represent multiple input audio items, or the input audio items can be represented in / by multiple encoded audio items.

[0088] It will be appreciated that many encoding algorithms and techniques are known and any suitable algorithms, standards, and methods can be used. It will also be appreciated that different algorithms and techniques can be used for different audio items. For example, audio items corresponding to music can be encoded using the AAC encoding method, audio items corresponding to speech can be encoded using the CELP encoding method, and so on. For audio items that have been received in an encoded format, encoding by the audio encoder 205 can be a transcoding to a different encoding format or can, for example, just be a data rate conversion (e.g., by modifying the quantization and / or clipping levels). Generally, encoding includes bitrate compression and the encoded audio items are represented by fewer bits than the input audio items.

[0089] The audio encoder unit 103 also includes a metadata circuit 207, which is arranged to generate output presentation metadata for the encoded audio item. The presentation metadata circuit 207 is arranged to generate the output presentation metadata based on the received input presentation metadata. In fact, for many audio items, the output presentation metadata may be the same as the input presentation metadata. For one or more audio items, the output presentation metadata can be modified, as will be described in more detail later.

[0090] The audio encoder 205 and the metadata circuit 207 are coupled to an output circuit 209, which is arranged to generate an encoded audio data stream including the encoded audio data and the output presentation metadata. The output circuit 209 can specifically be a bitstream packer, which generates an encoded audio data stream including both the encoded audio data and the output metadata. The encoded audio data stream can be generated according to a standardized format, thus allowing it to be interpreted by a series of receivers.

[0091] Therefore, the output circuit 209 operates as a bitstream packer, which receives the bitrate-reduced / encoded audio item and the output presentation metadata and combines them into a bitstream that can be communicated via a suitable communication channel (such as, for example, via a 5G network).

[0092] Figure 3 Illustrated is a specific example of the components of a rendering device 109 that can receive and process the encoded audio data stream from the audio encoder unit 103. The rendering device 109 includes a receiver 111 in the form of a bitstream unpacker, which receives the encoded audio data stream from the audio encoder unit 103 and separates different data from the received data stream. Specifically, the receiver 111 can separate the individual audio data for the encoded audio item and feed these to a decoder 113.

[0093] The decoder 113 is specifically arranged to decode the received encoded audio item to generate a generally unencoded representation of the audio item based on channels, objects, and scenes.

[0094] For many audio items, the decoder 113 can invert the encoding performed by the audio encoder 205. For other audio items, the decoding can, for example, only partially invert the encoding operation. For example, if the audio encoder 205 has combined the audio items into a single combined audio item, the decoder 113 can only decode the combined audio item without fully generating the individual audio items. It will be appreciated that any suitable decoding algorithms and techniques can be used according to the specific preferences and requirements of individual embodiments.

[0095] The decoded audio item is fed to a renderer 115, which is arranged to render the audio scene by rendering the audio item, for example, as a binaural signal or a surround sound signal as described above.

[0096] The rendering device 109 also includes a metadata controller / circuit 301 that is fed presentation metadata from the receiver 111. In this example, the metadata controller 301 may also receive local presentation metadata that may reflect local preferences or requirements, such as, for example, individual user preferences or the nature of the reproduction system being used.

[0097] Thus, in addition to the audio presentation metadata unpacked from the received bitstream, the rendering device 109 may also accept local audio presentation metadata that may be provided, for example, via one or more input interfaces. This data may provide information about the context in which the audio is to be presented, information not available at the encoder side, such as, for example:

[0098] – Desired presentation (speaker) configuration;

[0099] - User preferences (e.g., audio levels and orientations of participants' audio in a virtual meeting);

[0100] - Nature of the local acoustics, such as the reverberation of the room. This may allow the renderer to determine which environmental effects and properties to apply to the audio item;

[0101] - Local audio signals (e.g., considered when selecting a gain for an audio item;

[0102] - Listener location; and

[0103] - Listener head orientation.

[0104] The metadata controller 301 may combine the received metadata and the local metadata together and provide it to the renderer 115, which may proceed to render the audio item according to the constraints of the presentation metadata.

[0105] The renderer 115 may combine the audio items C”, O” and S” generated by the decoder 113 into presentable audio in a desired presentation configuration (e.g., binaural or surround sound).

[0106] The renderer 115 may specifically generate an audio presentation according to the metadata received from the metadata controller 301 and in cases where the rendered audio is constrained by the received presentation metadata (i.e., constrained from the encoder side). This provides source side / content provider control over the audio rendering and the presented audio scene while still allowing some flexibility on the client side. This can be used, for example, to provide services or applications where the content author retains control over an immersive application that is designed to provide a certain limited control to end users, etc.

[0107] More specifically, the metadata controller 301 can process the received metadata, such as suppression of audio items, and accordingly the local metadata. The metadata controller 301 can, for example, limit the local metadata, such as the range of rotation or elevation angle, and accordingly the received metadata.

[0108] In some embodiments, the renderer 115 can be a different device or functional entity from the rendering device 109. For example, standards such as the envisioned 3GPP IVAS codec can specify the operation of the decoder 113 but allow the renderer 115 to be proprietary and more freely adjustable. In some embodiments, the metadata controller 301 can be part of a different device or functional entity.

[0109] In such embodiments, an external renderer is therefore needed to process and interpret the decoded audio items O”, C”, S” and the received rendering metadata. The rendering operations of the external renderer must still comply with the constraints provided by the rendering metadata.

[0110] The rendering metadata can thus be data used by the source side / content provider to control the rendering operations at the client side. The rendering must be adjusted / limited according to the rendering metadata.

[0111] However, in addition to the rendering metadata for controlling the rendering of the client-side renderer 115, the audio encoder 205 of the audio encoder unit 103 is also arranged to adjust the encoding in response to the input rendering metadata. The input rendering metadata is fed to the audio encoder 205, and this can modify the encoding of one or more audio items (typically for one or more audio items) based on the rendering metadata. The audio encoder 205 is thus an adjustable encoder responsive to the rendering metadata received together with the audio items.

[0112] The audio encoder 205 specifically includes an encoding circuit 211 and an encoding adapter 213. The encoding circuit 211 is arranged to perform the encoding of the audio items, and the encoding adapter 213 is arranged to adjust the encoding of the encoding circuit 211 based on the rendering metadata.

[0113] The encoding adapter 213 can be arranged to set the parameters of the encoding for a given audio item based on the rendering metadata for that audio item. For example, it can be arranged to set the bitrate allocation / target, quantization level, masking threshold, frequency range, etc. based on, for example, the gain range or position range indicated as permissible for the audio item by the rendering metadata.

[0114] In many embodiments, the encoding circuit 211 is a bitrate compressor that is arranged to encode audio items with a reduced number of bits compared to the received input audio items. The encoding can thus be bitrate compression, allowing for a more efficient and easier distribution of the encoded audio data stream to be generated. In such embodiments, the encoding adapter 213 can adjust the bitrate reduction of the encoding circuit 211 based on the presentation metadata (in order to optimize the quality of the rendered audio according to suitable optimization criteria / algorithms).

[0115] The encoding adapter 213 can, for example, perform an encoding analysis process that analyzes the presentation metadata and decides how best to perform bitrate reduction of the various input audio items. Examples of operations and adjustments that can be performed by the encoding adapter 213 include:

[0116] - Signaling the (minimum) masking level of the encoding circuit 211 to comply with the bitrate reduction. The encoding adapter 213 has information regarding which audio items are co-presented and at what level and in which orientation. This can allow it to adjust the masking level for individual audio items, which is then used by the encoding for masking.

[0117] - Transforming audio items, e.g., moving audio objects into channel-based or scene-based audio.

[0118] - Selecting audio items for downmixing (with associated upmix parameters), where the downmix can be upmixed to reconstruct immersive audio at the decoder side while ensuring that artifacts of the parametric downmix encoding are sufficiently masked by the various audio items that are co-presented. As a further improvement, the encoding adapter 213 can

[0119] - Optimize the downmix / upmix gain to obtain maximum performance / minimum artifacts;

[0120] - Select upmix parameters with the best time / frequency characteristics.

[0121] - Irreversibly combine audio items into a combined audio item, which can then be rendered by the renderer 115 as a single audio item. This can take advantage of the inherent need that not all audio information needs to be made individually available at the rendering side. For example, if separate adjustment of some input audio items is not allowed (e.g., they may be required to be rendered in the same position), then the audio items do not have to be made individually available. For example, multiple input audio objects with similar orientation and gain adjustment constraints can be combined into a single scene-based audio item, where the gain and orientation of the entire scene can still be adjusted during rendering, but the previous objects will have fixed relative audio levels and fixed relative positions in the scene.

[0122] - Different bitrate budgets are allocated to different audio items according to the presentation metadata for the audio items. For example, the bitrate can be allocated to the audio items based on the amount of unmasked information they each represent.

[0123] The encoding circuit 211 can then adopt the encoding of the audio items according to the encoding control data generated by the encoding adapter 213. For example, the encoding circuit 211 can generate some bitrate-reduced (e.g., quantization, parameterization, etc.) versions of the audio items based on channels, objects, and scenes. In addition, due to, for example, combinations or conversions as part of the encoding of different audio items, at least some of the encoded audio items can represent audio information different from the input audio items, that is, there may be no direct correspondence between the input audio items and the encoded audio items.

[0124] In some embodiments, the audio encoder 205 can specifically include a combiner 215, which is arranged to combine the input audio items into one or more combined audio items. The combiner 215 can specifically combine the first and second input audio items into a combined audio item. The combined audio item can then be encoded to generate a combined encoded audio item, and the combined encoded audio item can be included in the encoded audio data stream, typically replacing the first and second audio items. Thus, instead of encoding the first and second audio items individually, the combiner 215 can combine them into a single encoded audio item, which is then included in the encoded audio data stream, without including individual encoded audio data for the first or second audio item separately.

[0125] The combination of the audio items is performed in response to the received presentation metadata. In many embodiments, the audio items selected for combination are selected based on the presentation metadata. For example, the encoding adapter 213 can select the audio items for combination in response to a criterion that includes a requirement for the constraints on the audio items to meet a similarity criterion.

[0126] For example, for the audio items to be combined, it may be required that the constraints on the audio items as indicated by the presentation metadata do not contradict each other, that is, it must be possible to satisfy both constraints. Therefore, it may be required that the constraints indicated by the presentation metadata do not conflict, and for example, the constraints at least have an overlap such that there is at least one rendering parameter value that allows the rendering constraints on the two (or all) audio items being combined to be satisfied. The encoding adapter 213 can require that the presentation metadata does not describe incompatible constraints on common rendering parameters.

[0127] For example, the presentation metadata can describe the constraints on the positions of the audio items in an audio scene. In this case, it may be required that the position constraints must overlap and there must be some common allowed positions.

[0128] The selection of audio items to be combined can be based on the presentation metadata for the audio items. Thus, the selection of the first and second audio items for combination can be based on the presentation metadata for the first and second audio items. For example, as described above, it may be required that the presentation metadata for the first and second audio items does not define conflicting constraints.

[0129] In some embodiments, the first and second audio items can be selected, for example, as audio items having constraints on, for example, the most similar identical parameters. For example, audio items having substantially the same position constraints can be selected.

[0130] Specifically, a similarity measure for the two audio items can be determined to reflect the overlap between the allowable positions. For example, the similarity measure can be generated as the ratio of the volume of the region of the overlapping allowable positions to the sum of the volumes of the individual allowable positions for the two audio items.

[0131] As another example, multiple audio objects that meet the similarity criteria for position adjustment constraints can be combined into a scene-based audio item, even when the corresponding position ranges or spatial volumes may not overlap, where the audio sources will have a fixed relative orientation to each other (i.e., not adjustable individually) in the scene-based audio, but their orientation can still be adjusted together as a whole.

[0132] As another example, a similarity measure can be generated to reflect the size of the overlapping gain ranges of the two audio items. The larger the common allowable gain range, the greater the similarity.

[0133] The encoding adapter 213 can evaluate such similarity measures for different pairs of audio items and select, for example, pairs for which the similarity measure is higher than a given threshold. These audio items can then be combined into a single combined audio item.

[0134] In many embodiments, the encoding adapter 213 is also arranged to generate combined presentation metadata for the combined audio item from the input presentation metadata. This presentation metadata is then fed to the bitstream packer 209, which includes it in the output encoded audio data stream.

[0135] The metadata circuit 207 can specifically generate composite presentation metadata that links to the composite audio item and provides rendering constraints for the composite presentation metadata. The generated composite audio item with the associated composite presentation metadata can then be treated as any other audio item, and in fact the client / decoder / renderer may not even know that the composite audio item was indeed generated by combining input audio items by the audio encoder 205. Instead, the composite audio item and the associated presentation metadata may be indistinguishable from the input audio item and the associated presentation metadata to the client side and can be rendered as any other audio item.

[0136] In many embodiments, for example, composite presentation metadata can be generated to reflect constraints on the presentation parameters for the composite audio item. The constraints can be determined such that they satisfy the individual constraints of the audio items being combined, as indicated by the input presentation metadata for those audio items. Specifically, the constraints for the composite audio item of the first and second audio items can be determined to satisfy both the constraints of the first audio item indicated by the input presentation metadata for the first audio and the constraints of the second audio item indicated by the input presentation metadata for the second audio item. Thus, the composite presentation metadata is generated to provide one or more constraints that, if the composite constraints are satisfied, ensure that the individual constraints on the individual audio items are satisfied.

[0137] For example, for a first audio item that is an audio object, the input presentation metadata can indicate that it must be rendered with a relative gain in a range, for example, from -6 dB to 0 dB and at a position within a coordinate volume of (azimuth, elevation, radius), for example, ([0,100],[-40,60],[0.5,1.5]). For a second audio item that is an audio object, the input presentation metadata can indicate that it must be rendered with a relative gain in a range, for example, from -3 dB to 3 dB and at a position within a coordinate volume of (azimuth, elevation, radius), for example, ([-100,80],[-20,70],[0.2,1.0]). In this case, the composite presentation metadata can be generated to indicate that the composite audio item, which is an audio object, must be rendered with a relative gain in a range, for example, from -3 dB to 0 dB and at a position within a coordinate volume (azimuth, elevation, radius), for example, ([0,80],[-20,60],[-0.5,1.0]). This will ensure that the composite audio item is rendered in a way that is acceptable for both the first audio item and the second audio item.

[0138] In some embodiments, the audio encoder 205 can be arranged to adjust the compression of one audio item based on the presentation metadata for another audio item.

[0139] As a low-complexity example, the compression of one audio item may depend on the proximity and gain / level of another audio item. For example, if the presentation metadata for the current audio item indicates a position range and a level range, this can be compared with the position range and level range of a second audio item. If the second audio item is constrained to be positioned close to the first audio item and is constrained to be rendered at a substantially higher level than the first audio item, the first audio item is likely to be barely perceptible to the listener. Thus, the encoding of the first audio item can have a higher compression / bitrate reduction compared to the case where no other audio items are present. Specifically, the bitrate allocation for encoding the first audio item can depend on the distance and level to one or more other audio items.

[0140] In some embodiments, the encoding adapter 213 may be arranged to estimate the masking effect of the second audio item on the first audio item. The masking effect may be represented by a masking metric that indicates the degree of masking introduced into the first audio item from the rendering of the second audio item. Thus, the masking metric indicates the perceived importance of the first audio item in the presence of the second audio item.

[0141] The masking metric may specifically be generated based on the constraints indicated by the presentation metadata as an indication of the level of the sound received from the second audio item relative to the level of the sound received from the first audio item when the second audio item is rendered.

[0142] For example, the masking effect of the first audio item at its lowest gain on the second audio item at its highest gain may be employed to estimate the masking level of the second item, and vice versa.

[0143] As another example, the furthest (or, for example, average) distance between the first and second audio items may be determined and the attenuation between them estimated. The masking effect may then be estimated based on the relative level difference after compensation for the attenuation.

[0144] As another example, if the system employs a nominal listening position, the signal levels at the listening position from the first and second audio items, respectively, may be determined based on the relative gain levels or signal levels and the attenuation difference from the positions of the sound sources. The audio item positions may be selected from the allowable positions, for example, such that the masking effect is minimized (the closest allowable position for the first audio item and the furthest position for the second audio item).

[0145] Thus, the encoding adapter 213 can estimate the masking effect of the second audio item on the first audio item based on the gain / level constraint and the position constraint of the second audio item indicated by the input presentation metadata for the second audio item and often also based on the gain / level constraint and the position constraint of the first audio item indicated by the input presentation metadata for the first audio item.

[0146] In some embodiments, the encoding adapter 213 can directly determine the masking threshold for the first audio item based on the presentation metadata for the second audio item, and the encoding circuit 211 can proceed to encode the first audio item using the determined masking threshold.

[0147] In some embodiments, the adjustment of the encoding by the audio encoder 205 can be an internal process without other functions being adjusted accordingly. For example, multiple audio items can be irreversibly combined into a combined audio item, where the combined audio item is included in the encoded audio data stream and there is no indication of how the combined audio item was created, i.e., no particular processing of the combined audio item is performed by the rendering device.

[0148] However, in many embodiments, the audio encoder 205 can generate encoding adjustment data that indicates how to adjust the encoding in response to the input presentation metadata. This encoding adjustment data can then be included in the encoded audio data stream. In this method, the rendering device 109 can accordingly have information about the encoding adjustment and can be arranged to adjust decoding and / or rendering accordingly.

[0149] For example, the audio encoder 205 can generate data indicating which audio items in the acoustic environment data are actually combined audio items. It can also indicate some parameters of the combination and in many embodiments these can actually allow the rendering device 109 to generate a representation of the original audio items that were combined. In fact, in some embodiments, the combined audio item can be generated as a downmix of the input audio items, and the audio encoder 205 can generate parametric upmix data and include this in the encoded audio data stream, enabling the rendering device to perform a reasonable upmix.

[0150] As another example, such decoding may not be adjusted, but this information can be used for interaction with the listener / end-user. For example, multiple audio objects that are considered "close" in their adjustment constraints can be combined by the encoder into a scene-based audio item, and their existence as "virtual objects" is signaled to the decoder in the encoding adjustment data. This information can then be presented to the user, and the user can be provided with manual control of the "virtual sound sources" (although only as a whole, since they have been combined in the scene-based audio) rather than being informed / aware of the scene-based audio as a carrier for the virtual objects.

[0151] In some embodiments, the presented metadata may include priority data for one or more audio items, and the audio encoder 205 may be arranged to adjust the compression for a first audio item in response to a priority indication for the first audio item.

[0152] The priority indication may be a rendering priority indication indicating the perceived significance or importance of an audio item in an audio scene. For example, it may be used to indicate that an audio item representing a main speaker is more important than an audio item representing, for example, a bird chirping in the background.

[0153] The renderer 115 may adjust the rendering based on the priority indication. For example, for a listener with reduced hearing, the renderer 115 may increase the gain of a high-priority main conversation relative to low-priority background noise, making the speech more intelligible.

[0154] In addition, the audio encoder 205 may increase the compression to reduce the priority. For example, in order to combine audio items, it may be required that the priority level must be below a given level. As another example, the audio encoder 205 may combine all audio items with a priority level below a given level.

[0155] In some embodiments, the bit allocation for each audio item may depend on the priority level. For example, the bit allocation for different audio items may be based on an algorithm or formula that takes into account multiple parameters including the priority. The bit allocation for a given audio item may increase monotonically with an increase in the priority.

[0156] It will be appreciated that, for clarity, the above description has described embodiments of the invention with reference to different functional circuits, units, and processors. However, it is apparent that any suitable functional distribution between different functional circuits, units, or processors may be used without departing from the invention. For example, functions illustrated as being performed by separate processors or controllers may be performed by the same processor or controller. Thus, the reference to a particular functional unit or circuit is only considered as a reference to an appropriate module for providing the described functionality, rather than indicating a strict logical or physical structure or organization.

[0157] The present invention may be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. Optionally, the present invention may be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the present invention may be physically, functionally, and logically implemented in any suitable manner. In fact, the functions may be implemented in a single unit, in multiple units, or as part of other functional units. As such, the present invention may be implemented in a single unit, or may be physically and functionally distributed among different units, circuits, and processors.

[0158] Generally, examples of an audio encoding apparatus, an audio encoding method, and a computer program product implementing the method are indicated by the following embodiments.

[0159] Although the present invention has been described in connection with some embodiments, it is not intended to be limited to the specific forms set forth herein. Instead, the scope of the present invention is limited only by the appended claims. Additionally, although features may be described as being combined with specific embodiments, those skilled in the art will recognize that the various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0160] Furthermore, although multiple modules, elements, circuits, or method steps are listed individually, they may be implemented by, for example, a single circuit, unit, or processor. Additionally, although individual features may be included in different claims, these features may advantageously be combined, and the inclusion in different claims does not imply that the combination of features is not feasible and / or advantageous. Moreover, including a feature in one category of claims does not imply a limitation to that category, but rather indicates that the feature is equally applicable to other claim categories where appropriate. Additionally, the order of features in the claims does not imply any particular order in which the features must operate, and in particular, the order of individual steps in method claims does not imply that the steps must be performed in that order. Instead, the steps may be performed in any suitable order. Additionally, singular references do not exclude plural. Thus, references to "a", "an", "first", "second", etc. do not exclude plural. The reference numerals in the claims are provided merely as illustrative examples and should not be construed as limiting the scope of the claims in any way.

Claims

1. An audio encoding apparatus, comprising: An audio receiver (201) arranged to receive a plurality of input audio items representing an audio scene; A metadata receiver (203) arranged to receive input rendering metadata for the plurality of input audio items, the input rendering metadata describing rendering constraints for rendering the plurality of input audio items; An audio encoder (205) arranged to generate encoded audio data for the audio scene by encoding the plurality of input audio items, the audio encoder (205) including a combiner (215) arranged to generate a combined audio item by at least combining the first audio item and the second audio item in response to the input rendering metadata for the first audio item among the plurality of input audio items and the input rendering metadata for the second audio item among the plurality of input audio items, and wherein the audio encoder (205) is arranged to generate combined audio encoded data for the first audio item and the second audio item by encoding the combined audio item, and include the combined audio encoded data in the encoded audio data; A metadata circuit (207) arranged to generate output rendering metadata based on the input rendering metadata, the output rendering metadata including data for the encoded audio data, the data for the encoded audio data constraining the degree to which adjustable parameters of rendering at a client can be adjusted when rendering the encoded audio data; and An output circuit (209) arranged to generate an encoded audio data stream to be sent to the client, including the encoded audio data and the output rendering metadata; and Wherein the output rendering metadata includes at least one of the following: Reverberation constraints; and Audio item position constraints.

2. The audio coding device according to claim 1, wherein, The combiner (215) is arranged to select the first audio item and the second audio item from the plurality of input audio items in response to the input rendering metadata for the first audio item and the second audio item.

3. The audio encoding device according to claim 1 or 2, wherein The combiner (215) is arranged to select the first audio item and the second audio item in response to a determination that at least some of the input rendering metadata for the first audio item and the input rendering metadata for the second audio item satisfy a similarity criterion.

4. The audio coding device according to claim 1, wherein, The input rendering metadata for the first audio item and the input rendering metadata for the second audio item include at least one of a gain constraint and a position constraint.

5. The audio encoding device according to claim 1, wherein, The audio encoder (205) is further arranged to generate combined rendering metadata for the combined audio item in response to the input rendering metadata for the first audio item and the input rendering metadata for the second audio item; And include the combined rendering metadata in the output rendering metadata.

6. The audio coding device according to claim 5, wherein, The audio encoder (205) is arranged to generate at least some combined presentation metadata to reflect constraints on presentation parameters for the combined audio item, the constraints being determined to satisfy the constraints of the first audio item indicated by the input presentation metadata for the first audio item and the constraints of the second audio item indicated by the input presentation metadata for the second audio item both.

7. The audio coding device according to claim 1, wherein, The audio encoder (205) is arranged to adjust the compression of the first audio item in response to the input presentation metadata for the second audio item.

8. The audio encoding device according to claim 7, wherein, The audio encoder (205) is arranged to estimate the masking effect of the second audio item on the first audio item in response to the input presentation metadata for the second audio item, and to adjust the compression of the first audio item in response to the masking effect.

9. The audio encoding device according to claim 8, wherein, The audio encoder (205) is arranged to estimate the masking effect of the second audio item on the first audio item in response to at least one of the gain constraint and the position constraint of the second audio item indicated by the input presentation metadata for the second audio item.

10. The audio encoding device according to claim 7, wherein, The audio encoder (205) is further arranged to adjust the compression of the first audio item in response to the input presentation metadata for the first audio item.

11. The audio encoding device according to claim 1, wherein, The input presentation metadata includes priority data for at least some of the audio items, and the encoder is arranged to adjust the compression for the first audio item in response to the priority indication for the first audio item in the input presentation metadata.

12. The audio encoding device according to claim 1, wherein, The audio encoder (205) is arranged to generate encoding adjustment data indicating how to adjust the encoding in response to the input presentation metadata, and to include the encoding adjustment data in the encoded audio data stream.

13. A method of encoding audio, the method comprising: Receiving a plurality of input audio items representing an audio scene; Receiving input presentation metadata for the plurality of input audio items, the input presentation metadata describing presentation constraints for rendering the plurality of input audio items; Generating encoded audio data for the audio scene by encoding the plurality of input audio items, the encoding responding to the input presentation metadata by: generating a combined audio item by at least combining the first audio item and the second audio item in response to the input presentation metadata for the first audio item and the input presentation metadata for the second audio item among the plurality of input audio items; and generating combined audio encoded data for the first audio item and the second audio item by encoding the combined audio item, and including the combined audio encoded data in the encoded audio data; Generate output presentation metadata based on the input presentation metadata, the output presentation metadata including data for the encoded audio data, the data for the encoded audio data constraining the degree to which adjustable parameters of rendering at the client can be adjusted when rendering the encoded audio data; And Generate an encoded audio data stream to be sent to the client, the encoded audio data stream including the encoded audio data and the output presentation metadata; and wherein the output presentation metadata includes at least one of the following: Reverberation constraint; and Audio item position constraint.

14. A computer program product comprising computer program code modules adapted to perform all the steps of claim 13 when the program is run on a computer.

Citation Information

Patent Citations

  • System and Method for Non-destructively Normalizing Loudness of Audio Signals Within Portable Devices

    US20120310654A1

  • Sound processing device and method

    WO2018047667A1