Rendering of a spatial audio stream

The described system allows interactive gain adjustments for audio objects and MASA components in spatial audio rendering, addressing limitations in existing systems by enabling personalized audio experiences through metadata-based gain control.

WO2025201817A1PCT designated stage Publication Date: 2025-10-02NOKIA TECHNOLOGIES OY
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/055909
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2025-03-05
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing spatial audio rendering systems, such as those using the IVAS codec, do not adequately support interactive adjustments of gain factors for audio objects and MASA components, limiting user control over the audio experience.

Method used

A system and method for rendering spatial audio streams that allow interactive adjustment of gain factors based on metadata, enabling individual control over audio objects and MASA components by determining gain processing information using audio object portion energy proportions and spatial metadata.

Benefits of technology

Enables users to modify the loudness of individual audio objects and MASA parts, providing a personalized audio experience by adjusting gain values based on spatial properties and energetic information, enhancing audio quality and user interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025055909_02102025_PF_FP_ABST
    Figure EP2025055909_02102025_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus for rendering a spatial audio signal based on a spatial audio stream comprising: at least one audio signal; and metadata associated with the at least one audio signal, the apparatus comprising means configured to: obtain the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audio object portion energy proportion; obtain gain control information; determine gain processing information based on: the gain control information; and the at least one audio object portion energy proportion; and render the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]RENDERING OF A SPATIAL AUDIO STREAMField The present application relates to apparatus and methods for rendering of a spatial audio stream, but not exclusively for gain factors in a combined formatparametric spatial stream interactive audio rendering.Background Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An exampleof such a codec is the 3GPP Immersive Voice and Audio Services (IVAS) codecwhich is designed to be suitable for use over a communications network such as a 4G / 5G network including use in such immersive services as for example immersivevoice and audio for virtual reality (VR). This audio codec handles the encoding,decoding and rendering of speech, music and generic audio. It supports a varietyof input formats, such as channel-based audio, object-based audio, and scene-based audio inputs including spatial information about the sound field and soundsources, as well as MASA (metadata-assisted spatial audio) inputs. IVAS operateswith low latency to enable conversational services as well as supports high errorrobustness under various transmission conditions. The IVAS codec operates on awide range bit rates from very low (13.4 kb / s) to relatively high bit rates (512 kb / s).Input signals can be presented to the IVAS encoder in one of a number of supported formats (and in some allowed combinations of the formats). Forexample, a mono audio signal (without metadata) may be encoded using anEnhanced Voice Service (EVS) encoder. Other input formats utilize new IVAS encoding tools. One input format proposed for IVAS is the Metadata-assisted spatial audio (MASA) format, where the encoder may utilize, e.g., a combination of mono and stereo encoding tools and metadata encoding tools for efficient transmission of the format. The use of “Audio objects”, or Independent streams with metadata (ISM), is another example of an input format proposed for IVAS. In this input format thescene is defined by a number (1 to N) of audio objects (where N is, e.g., 4). Eachof the objects have an individual audio signal and some metadata describing its (spatial) features. The metadata may be a parametric representation of audio objectand may include such parameters as the direction of the audio object (e.g., azimuthand elevation angles). Furthermore, IVAS supports inputs in combined formats, e.g., the combinedformat input of audio object and MASA streams. This combination is referred to asOMASA. An example of such an input would be that the MASA format stream isobtained from the conferencing microphone in a room, capturing everything in thespace, and the object format streams may be obtained using lapel microphones ofthe talkers. The combined encoding of the two input streams allows improvementsin the quality of the coding. Being able to interact with an object or objects in the decoder side whenemploying a IVAS codec may be a desirable feature. For example, listener A maywant to change a gain related to an audio object, whereas listener B may want tochange the gain related to the same audio object. Furthermore listener B may wantto change the main gain differently to the audio object or the gain of either the objectrelative to some other audio object. Furthemore in a combined format example, thelistener may want to adjust the gain of the MASA component which may representthe signal background (and which may be relative to the object gain). Thus,rendering systems implementing codecs such as the above should be able toperform an interaction within the decoder / renderer so that each listener can havean individual experience. Summary There is provided according to a first aspect an apparatus for rendering aspatial audio signal based on a spatial audio stream comprising: at least one audio signal; and metadata associated with the at least one audio signal, the apparatus comprising means configured to: obtain the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio objectportion and at least one other spatial audio portion, and the associated metadatais at least configured to define at least one audio object portion position and at leastone audio object portion energy proportion; obtain gain control information;determine gain processing information based on: the gain control information; and the at least one audio object portion energy proportion; and render the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal. The gain control information may comprise at least one gain for at least oneof: the at least one audio object portion; and the at least one other spatial audioportion. The at least one gain for the at least one audio object signal may be based on a distance parameter associated with an at least one audio object of the at least one audio object portion. The means configured to determine the gain processing information based on the gain control information and the at least one audio object portion energyproportion may be configured to determine at least one first gain value based onthe at least one audio signal and the at least one audio object portion energy proportion. The means may be further configured to apply the at least one first gainvalue to one of: the at least one audio object portion; and the at least one other spatial audio portion. The means configured to determine gain processing information may be further configured to determine gain processing information based on the at least one audio object portion position. The at least one other spatial audio portion may comprise at least two transport audio signals. The means configured to obtain the spatial audio stream may be configuredto perform at least one of: receive information defining the at least one audio objectportion position and at least one audio object portion energy proportion; and receiveat least one parameter value defining the at least one audio object portion position and at least one audio object portion energy proportion associated with the at least one object. The means may be further configured to process the metadata associatedwith the at least one audio signal based on the gain control information.The means configured to process the metadata associated with the at leastone audio signal based on the gain control information may be configured to atleast one of: process metadata associated with the at least one audio object portion of the at least one audio signal; and process metadata associated with the at least other spatial audio portion of the at least one audio signal. The at least one audio object portion and at least one other spatial audio portion may be a mixture within the at least one audio signal. According to a second aspect there is provided a method for an apparatusfor rendering a spatial audio signal based on a spatial audio stream comprising: at least one audio signal; and metadata associated with the at least one audio signal, the method comprising: obtaining the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audioobject portion energy proportion; obtaining gain control information; determininggain processing information based on: the gain control information; and the at least one audio object portion energy proportion; and rendering the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal. The gain control informationn may comprise at least one gain for at least one of: the at least one audio object portion; and the at least one other spatial audio portion. The at least one gain for the at least one audio object signal is based on a distance parameter associated with an at least one audio object of the at least one audio object portion. Determining the gain processing information based on the gain controlinformation and the at least one audio object portion energy proportion maycomprise determining at least one first gain value based on the at least one audiosignal and the at least one audio object portion energy proportion.The method may further comprise applying the at least one first gain valueto one of: the at least one audio object portion; and the at least one other spatial audio portion. Determining gain processing information may further comprise determining gain processing information based on the at least one audio object portion position. The at least one other spatial audio portion may comprise at least two transport audio signals. Obtaining the spatial audio stream may comprise at least one of: receivinginformation defining the at least one audio object portion position and at least oneaudio object portion energy proportion; and receiving at least one parameter valuedefining the at least one audio object portion position and at least one audio object portion energy proportion associated with the at least one object. The method may further comprise proessing the metadata associated withthe at least one audio signal based on the gain control information.Processing the metadata associated with the at least one audio signal basedon the gain control information may comprise at least one of: processing metadataassociated with the at least one audio object portion of the at least one audio signal; and processing metadata associated with the at least other spatial audio portion of the at least one audio signal. The at least one audio object portion and at least one other spatial audioportion may be a mixture within the at least one audio signal.According to a third aspect there is provided an apparatus for rendering aspatial audio signal based on a spatial audio stream, the apparatus comprising atleast one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the atleast one processor, cause the apparatus at least to perform: obtaining the spatialaudio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audio object portion energyproportion; obtaining gain control information; determining gain processinginformation based on: the gain control information; and the at least one audio object portion energy proportion; and rendering the spatial audio signal based on the gainprocessing information, the at least one audio signal, and the metadata associatedwith the at least one audio signal. The gain control information may comprise at least one gain for at least one of: the at least one audio object portion; and the at least one other spatial audio portion. The at least one gain for the at least one audio object signal is based on a distance parameter associated with an at least one audio object of the at least one audio object portion. The apparatus caused to perform determining the gain processing information based on the gain control information and the at least one audio objectportion energy proportion may be caused to perform determining at least one firstgain value based on the at least one audio signal and the at least one audio objectportion energy proportion. The apparatus may further be caused to perform applying the at least one first gain value to one of: the at least one audio object portion; and the at least one other spatial audio portion. The apparatus caused to perform determining gain processing information may further be caused to perform determining gain processing information based on the at least one audio object portion position. The at least one other spatial audio portion may comprise at least two transport audio signals. The apparatus caused to perform obtaining the spatial audio stream may be caused to perform at least one of: receiving information defining the at least one audio object portion position and at least one audio object portion energy proportion; and receiving at least one parameter value defining the at least one audio object portion position and at least one audio object portion energy proportion associated with the at least one object. The apparatus may be caused to further perform proessing the metadataassociated with the at least one audio signal based on the gain control information.The apparatus caused to perform processing the metadata associated withthe at least one audio signal based on the gain control information may be causedto perform at least one of: processing metadata associated with the at least one audio object portion of the at least one audio signal; and processing metadata associated with the at least other spatial audio portion of the at least one audio signal. The at least one audio object portion and at least one other spatial audio portion may be a mixture within the at least one audio signal. According to a fourth aspect there is provided an apparatus for rendering aspatial audio signal based on a spatial audio stream comprising: at least one audio signal; and metadata associated with the at least one audio signal, the apparatuscomprising: means for obtaining the spatial audio stream comprising: the at leastone audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audioobject portion energy proportion; means for obtaining gain control information;means for determining gain processing information based on: the gain control information; and the at least one audio object portion energy proportion; and means for rendering the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal. According to a fifth aspect there is provided a computer program comprisinginstructions [or a computer readable medium comprising program instructions] for causing an apparatus for rendering a spatial audio signal based on a spatial audiostream to perform at least the following: obtaining the spatial audio streamcomprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion positionand at least one audio object portion energy proportion; obtaining gain controlinformation; determining gain processing information based on: the gain controlinformation; and the at least one audio object portion energy proportion; and rendering the spatial audio signal based on the gain processing information, the atleast one audio signal, and the metadata associated with the at least one audiosignal. According to a sixth aspect there is provided a non-transitory computerreadable medium comprising program instructions for causing an apparatus for rendering a spatial audio signal based on a spatial audio stream to perform at least the following: obtaining the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audio objectportion energy proportion; obtaining gain control information; determining gainprocessing information based on: the gain control information; and the at least one audio object portion energy proportion; and rendering the spatial audio signal basedon the gain processing information, the at least one audio signal, and the metadataassociated with the at least one audio signal. According to a seventh aspect there is provided an apparatus for renderinga spatial audio signal based on a spatial audio stream, the apparatus comprising:obtaining circuitry configured to obtain the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at leastone audio object portion energy proportion; obtaining circuitry configured to obtaingain control information; determining circuitry configured to determine gainprocessing information based on: the gain control information; and the at least one audio object portion energy proportion; and rendering circuitry configured to render the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal. According to an eighth aspect there is provided a computer readablemedium comprising program instructions for causing an apparatus for rendering a spatial audio signal based on a spatial audio stream to perform at least the following: obtaining the spatial audio stream comprising: the at least one audio signal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audio objectportion energy proportion; obtaining gain control information; determining gainprocessing information based on: the gain control information; and the at least one audio object portion energy proportion; and rendering the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal. An apparatus comprising means for performing the actions of the method as described above. An apparatus configured to perform the actions of the method as described above. A computer program comprising program instructions for causing a computer to perform the method as described above. A computer program product stored on a medium may cause an apparatus to perform the method as described herein. An electronic device may comprise apparatus as described herein. A chipset may comprise apparatus as described herein. Embodiments of the present application aim to address problems associated with the state of the art. Summary of the Figures For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which: Figure 1 shows schematically a system of apparatus suitable forimplementing some embodiments; Figure 2 shows a flow diagram of the operation of the apparatus shown in Figure 1 according to some embodiments; Figure 3 shows schematically an example of the encoder as shown in Figure1 according to some embodiments; Figure 4 shows a flow diagram of the operations of the example encodershown in Figure 3 according to some embodiments;Figure 5 shows schematically an example of the decoder as shown in Figure1 according to some embodiments; Figure 6 shows a flow diagram of the operations of the example decodershown in Figure 5 according to some embodiments;Figure 7 shows schematically an example of the spatial synthesizer asshown in Figure 5 according to some embodiments; Figure 8 shows a flow diagram of the operations of the example spatialsynthesizer shown in Figure 7 according to some embodiments; andFigure 9 shows schematically an example device suitable for implementingthe apparatus shown herein.Embodiments of the The concept as discussed herein in further detail with respect to the following embodiments is related to parametric spatial audio rendering. In the following examples an IVAS codec is used to show practical implementations or examplesof the concept. However, it would be appreciated that the embodiments presentedherein may be extended to other codecs without inventive input. The concept as discussed herein in further detail in the following embodiments is one of providing an ability to modify or edit the gains of differentaudio objects and of a MASA part within a suitable decoder / renderer for a combinedspatial audio format such as combined object and MASA formats (OMASA). As discussed earlier the metadata-assisted spatial audio (MASA) is one of the input formats supported by IVAS. It uses audio signal(s) together with corresponding spatial metadata (containing, e.g., directions and direct-to-total energy ratios in frequency bands). The MASA stream can, e.g., be obtained by capturing spatial audio with microphones of, e.g., a mobile device, where the set of spatial metadata is estimated based on the microphone signals. The MASA stream can be obtained also from other sources, such as specific spatial audio microphones (such as Ambisonics), studio mixes (e.g., 5.1 multichannel mix) or other content by means of a suitable format conversion. MASA spatial metadata values are available for each time-frequency tile(TF-tile) (there can, for example, be 24 frequency bands and 4 temporal sub-framesin each frame). The frame size in IVAS is 20 ms (and thus the temporal sub-frameis 5 ms). In addition, MASA supports 1 or 2 directions for each time-frequency tile(i.e., there are 1 or 2 direction index, direct-to-total energy ratio, and spread coherence parameters for each time-frequency tile). Additionally IVAS supports also audio objects (Independent streams with metadata, ISM) as an input. The audio objects contain for each object an audio signal and associated metadata (e.g., the direction of the object).Furthermore, IVAS supports a combined format input of audio object andMASA streams. This combination is referred to as OMASA. The combined encoding of the two input streams allows improvements in the quality of the coding. For example, at certain bitrates with a certain number of objects (e.g., 3-4 objects at 64 kbps), the audio signals of some of the objects and the audio signals of the MASA stream are combined to common stereo transport audio signals, which are then encoded. As fewer audio signals are encoded than there were in the input, significant bitrate savings can be made. As a result, the effective bitrate per encoded audio channel is higher than in case all the audio signals would have been separately encoded leading to improved audio quality There have been discussed methods, such as UK patent applicationGB2104309.6 where a position of an object in a renderer or decoder can bechanged for OMASA audio format signals in an IVAS codec. However, such an approach involved rotation only and did not propose methods for enabling changing the loudness (due to, e.g., distance or level) of an object. Spatial Audio Object Coding (SAOC) as described in Herre, J. et al. (2012). “MPEG spatial audio object coding—the ISO / MPEG standard for efficient coding ofinteractive audio scenes”, Journal of the Audio Engineering Society, 60(9), 655-673describes encoding objects as a downmix and spatial metadata and then decodingthem while allowing rendering-time spatial adjustments. Audio objects are downmixed for example to a stereo track and a set of spatial metadata is extractedin time-frequency regions. This metadata comprises Object level differences(OLD), inter-object cross coherences (IOC), downmix channel level differences (DCLD), downmix gains (DMG) and object energies (NRG). Such a set of metadataprovides the information to manipulate or render the multi-object mixture to aspatialized output. However, the metadata involved, and therefore the techniques related do not provide means to account for mixtures that prominently have non- object content. Similarly, MPEG-H methods such as described in Herre, J., Hilpert, J.,Kuntz, A., & Plogsties, J. (2015), ”MPEG-H 3D audio—The new standard for coding of immersive spatial audio”, IEEE Journal of selected topics in signalprocessing, 9(5), 770-779 applies the methods as described in context of SAOC,however in an extended form that is referred to as SAOC-3D (such as described in Murtaza, A., Herre, J., Paulus, J., Terentiv, L., Fuchs, H., & Disch, S. (2015,October), “ISO / MPEG-H 3D audio: SAOC 3D decoding and rendering”, AudioEngineering Society Convention 139. Audio Engineering Society. These describea system having audio objects and discrete surround audio channels at the same downmix stream, and effective decoding of them. However, this extension appliesonly to object-only coding as the original audio channels are conceptually close tostatic audio objects. In other words, SAOC-3D does not provide means to account for mixtures that prominently have non-object content, where “non-object content” is understood more broadly than loudspeaker channels, i.e., spatially static audio objects. The concept as discussed in the embodiments herein expands on these implementations in that the transport audio signals do not only include audio objects, but also audio signals (and associated spatial metadata) from othersources. The other sources may, for example, be spaced microphone capturedaudio signals (which could be captured using a mobile device), downmixed 5.1audio signals, or any other suitable audio input format.The embodiments herein thus improve on the methods disclosed by SAOC in being configured to provide means to effectively handle gain based modifications(for example, object distance modification) of such signals. In particular in certaincoding modes for the object and MASA combined input (OMASA) in IVAS, at least some of the object audio signals and the MASA audio signals are mixed together to a common stereo downmix signal. This enables obtaining higher bitrate per encoded audio channel, and thus improved audio quality. However, as a result of mixing the audio signals together, the object audio signals, nor the MASA audio signals, are no longer individually available in the decoder. This is not a problem, if the objects and the MASA stream are to be rendered according to the original input metadata without any modifications. However, the embodiments discussed herein enable the listener to modify the sound scene in some way before the rendering. This can be applied to OMASA(objects + MASA) but can also be applied to other combinations such as objects +Ambisonics and objects + multi-channel signals (e.g., 5.1 or 7.1+4). For example, the listener often wants to change the loudness of certain objects. For example, the objects may contain speech of different participants in a teleconference. In such a case, the listener may want to make a selected participant louder, and some other participant softer. The embodiments as described herein enable such a scenario to be implemented. Moreover, in some circumstances the listener may wants to change the loudness of the MASA part. For example, a recording of a call at a live concert can result in the MASA part may contain some background sounds (e.g., a live concert)and the object part may contain speech of the caller. In this case, the listener at theother end may want to make the MASA part louder to hear the concert better, or alternatively may want to make the MASA part softer, to hear the speech of the caller better. The embodiments as discussed in detail hereafter enable this to be implemented. The examples presented herein thus enable changing the loudness of the different objects or the different audio streams despite not having independentaudio signals for the objects and the MASA parts (for example, in certain codingmodes of the OMASA input in IVAS where there are 3-4 objects encoded at 64 kbps). In some embodiments the changing of the loudness of the objects and / orthe MASA part may be performed manually or automatically. For example, alistener or user may adjust the loudness of the different participants using a suitableuser interface (UI) on, for example, a mobile phone app. This may be performed,in some embodiments, during a call. Alternatively, in some embodiments theloudness adjustment may be performed automatically, or semi-automatically, forexample, based on the preferences of the user in the past manual adjustments.Hence in summary there is presented an interactive rendering of aparametric spatial audio stream, which contains at least one audio signal and associated spatial metadata, and where the audio signal is a mixture of audio objects and other audio content (e.g., MASA). The apparatus and methods as discussed herein enable the modification of the loudness of individual audio objects and the other audio content at the rendering and enabling, for example, different listeners having an individual balance between the different objects and the other spatial audio content. This can be achieved by determining information of the proportion (orenergetic information or values) of the audio objects and the other content (forexample, in a time-frequency domain representation) as well as their spatialproperties. Additionally there can be a determination of target proportions (or targetenergetic information or values) for the different audio objects and the other audiocontent based on the at least one audio signal, the spatial metadata and the values relared to the desired level modification. Furthermore there can be a modification of the at least one audio signal based on the determined energetic values and the determined target energetic values. Then there can be a modification of the spatial metadata based on the determined energetic values and the determined target energetic values. After this there can be a rendering of the spatial audio (for example, binaural audio signals) using the modified at least one audio signal and the modified spatial metadata. In such embodiments the values related to the desired level modification may be based on, for example, the desired levels, desired level changes, and / or the desired distances of the audio objects. The energetic values can furthermore be related, for example, to theenergies or the levels of the objects and the other audio content. The values can be absolute, or they can be relative. The energetic values may also be called level values. The other spatial audio content can in some embodiments be in any suitable format, for example, in MASA, Ambisonics, or multi-channel (e.g., 5.1 or 7.1+4). The mixture audio signal may, for example, have been created in an encoder, where the audio signals from the objects and the other spatial audio content have been at least partially mixed together. Before discussing the concept in further detail we will initially describe infurther detail some aspects of spatial encoding, decoding and reproduction whichmay be implemented in some embodiments. For example, with respect to Figure 1is shown an example system suitable for implementing embodiments as described herein. The system comprises an encoder 101 which is configured to receive anumber M of spatial audio signal streams. In Figure 1 is shown the spatial audiostream 1104, spatial audio stream 2106, and spatial audio stream M 108 which isinput to the encoder 101. The encoder 101 can in some embodiments comprise an IVAS encoder, though in other embodiments other suitable encoders can be employed. The spatial audio streams 104, 106, 108 can in some embodiments bedifferent kind of streams. For example, the streams can be MASA streams,multichannel loudspeaker signal streams, and / or object streams. The encoder 101is configured to generate an encoded bitstream 110. The encoded bitstream 110 inFigure 1 is shown being passed to a separate decoder 111. However in someembodiments the bitstream may be stored in a suitable storage medium for later retrieval. The system further comprises a decoder 111. The decoder 111 is configured to receive or retrieve the encoded bitstream 110. Additionally the decoder 111 isconfigured to receive gain control information 112. The decoder, which can be anIVAS decoder or any suitable format decoder (which matches the encoder) isconfigured to decode the bitstream 110 and render spatial audio signals 114 basedon the gain control information 112. The gain control information 112 can, forexample, comprise desired gains for the objects and / or the other streams (forexample, the user of the decoder 111 apparatus may have set them). The spatialaudio signals 114, in some embodiments, can be binaural audio signals.Figure 2 shows, for example, a flow diagram of the operation of the example system as shown in Figure 1. Thus the spatial audio streams are initially obtained as shown in step 201. Then the audio streams are encoded to generate the bitstream as shown in Figure 2 by step 203. The encoded bitstream is then transmitted to the decoder / received from the encoder (or stored / retrieved) as shown in Figure 2 by step 205. Additionally the gain control information is obtained as shown in Figure 2 bystep 206. The encoded audio streams in the form of the encoded bitstream is thendecoded and spatial audio signals are rendered based on the gain controlinformation as shown in Figure 2 by step 207. Finally the spatial audio signals is output as shown in Figure 2 by step 209.Figure 3 shows an example encoder 101 as shown in Figure 1 according tosome embodiments. In this example, there are shown two input streams. The first input stream is a MASA stream, which comprises a MASA transport audio signals302 and MASA metadata 300. The second input stream shown in Figure 3 is anobject audio stream 320 (containing a number of, for example, N objects).It should be noted that in some embodiments there can be any suitablenumber of input streams and these two input streams are an example only.The encoder 101 comprises an object analyser 301. The object analyser301 has an input which receives the object audio stream 320 and is configured toanalyse the object audio stream 320 to determine which object to separate fromthe other objects. This can, for example, be performed in a manner as presentedin WO2022 / 214730. The result of this analysis can be the separated object audiosignal 316 and separated object metadata 314 which are outputted from the objectanalyser to an object encoder 309. The separated object metadata 314 can, forexample, contain the object direction and the index of the object that wasseparated. In some embodiments the object encoder 309 is configured to receive theseparated object audio signal 316 and separated object metadata 314 and encodethese based on any suitable encoding mechanism. The resulting encodedseparated object audio signal 326 and encoded separated object metadata 324 areoutput to a multiplexer or Mux 307.Additionally the remaining objects (from the object audio stream 320) areanalyzed, to determine object transport audio signals 312 and object metadata 310.This remaining object analysis can be performed using any suitable method (forexample, in a manner similar to that described in WO2022 / 200666). Furthermore,the metadata can contain any suitable metadata. As an example, the object transport audio signals 312 can be a stereo downmix using amplitude panning based on the object directions, and the objectmetadata 310 may contain the object directions and time-frequency domain object-to-total energy ratios (or, ISM ratios, in other words), which are obtained by analyzing the energies of the different objects in frequency bands and comparingthem to the total energy of all objects in the corresponding bands and produceobject transport audio signals 312 and object metadata 310. The object transport audio signals 312 and the object metadata 310 can in some embodiments be passed to a metadata encoder 303 and the object transport audio signals 312 and MASA transport audio signals 302 passed to a transport audio signal combiner and encoder 305. In some embodiments the encoder 101 comprises a transport audio signal combiner and encoder 305. The transport audio signal combiner and encoder 305is configured to obtain the MASA transport audio signals 302 and object transportaudio signals 312 and combine and encode these inputs to generate encodedtransport audio signals 306. The combination in some embodiments may be by summing them. In some embodiments, the transport audio signal combiner andencoder 305 is configured to perform other processing on the obtained transportaudio signals or the combination of the transport signals. For example, in someembodiments the transport audio signal combiner and encoder 305 is configuredto adaptively equalize the resulting signals in order to have the same energy in thetime-frequency domain for the combined signals as the sum of the energies of the MASA and object transport audio signals. The encoding of the combined transport audio signals can employ anysuitable codec. For example, in some embodiments the transport audio signalcombiner and encoder 305 is configured to encode the combined transport audiosignals using a EVS or AAC codec. In some implementations the transport audiosignal combiner and encoder 305 is configured to encode the combined transport audio signals using a IVAS core coder. The encoded transport audio signals 306 can then be output to a multiplexer or Mux 307. In some embodiments the encoder 101 comprises a metadata encoder 303.The metadata encoder 303 is configured to receive the MASA metadata 300 andthe object metadata 310 (in some embodiments the metadata encoder 303 isfurther configuered to receive the MASA transport audio signals 302 and the objecttransport audio signals 312). The metadata encoder 303 is configured to apply asuitable encoding to the metadata. The implementation of the metadata encoding may be any suitable encodingmethod, a few examples of which are described hereafter.As a first example of metadata encoding, MASA-to-total energy ratios aredetermined using the MASA transport audio signals 302 and the object transportaudio signals 312, for example, by computing the energies of the MASA and theobject transport audio signals in time-frequency tiles, and then determining the MASA-to-total energy ratios by where ^^^^^(^, ^) is the energy of the MASA transport audio signals 302 forthe frequency band ^ and temporal subframe ^, and ^^^^(^, ^) the energy of theobject transport audio signals 312.Then, the MASA metadata, the object directions, the ISM ratios, and the MASA-to-total energy ratios are encoded using any suitable methods (for example,using the methods presented in WO2020 / 089510, PCT / FI2019 / 050675,GB1811071.8, WO2020 / 193865, GB1913274.5, WO2022 / 200666, GB2217884.2, GB2217905.5, GB2217928.7, GB2217884.2). The resulting encoded metadata304 is outputted from the block.The encoded transport audio signals 306, encoded metadata 304, encodedseparated object audio signal 326, and encoded separated object metadata 324are forwarded to the multiplexer or Mux 307, which multiplexes them to a bitstream110, which is output from the encoder 101.Figure 4 shows a flow diagram of the operation of the example encoder as shown in Figure 3. As such in some embodiments the object audio streams are obtained as shown in Figure 4 by step 401. The object audio streams are then analysed to generate the object transport audio signals, object metadata, separated object audio signal and separated objectmetadata as shown in Figure 4 by step 403.Additionally the MASA transport audio signals are obtained as shown in Figure 4 by step 402. The MASA metadata is furthermore obtained as shown in Figure 4 by step 404. Having obtained the MASA transport audio signals and the object transport audio signals these are combined and encoded to generate encoded combined transport audio signals as shown in Figure 4 by step 405. Having obtained the MASA metadata and the object metadata these are combined and encoded to generate encoded combined metadata as shown in Figure 4 by step 406. Furthermore the separated object audio signal and separated object metadata are encoded as shown in Figure 4 by step 407. Having generated the encoded separated object audio signal, the encodedseparated object metadata, the encoded combined metadata and the encodedcombined transport audio signals then these can be multiplexed as shown in Figure 4 by step 408. Then the bitstream (the multiplexed encoded signals) are output as shown in Figure 4 by step 409. Figure 5 shows an example decoder 111 as shown in Figure 1 according to some embodiments. In this example, there is shown the bitstream which is obtained by the decoder 111. The decoder 111 can in some embodiments comprise a demultiplexer or demux 501 which is configured to obtain the bitstream 110 and demultiplex thebitstream to generate:encoded metadata 502, which is passed to a metadata decoder and processor 503; encoded transport audio signals 512, which is passed to the transport audiosignal decoder 513; andencoded separated object audio signal 522 and encoded separated object metadata 532 which are passed to the object decoder 523. Furthermore the decoder 111 can in some embodiments comprise a transport audio signal decoder 513 configured to receive the encoded transport audio signals 512. The transport audio signal decoder 513 can then be configuredto decode the encoded transport audio signals 512 and generate decoded transportaudio signals 514 which can be passed to a spatial synthesizer 505. The decoder 111 furthermore, in some embodiments, comprises a metadata decoder and processor 503 configured to receive the encoded metadata 502. The metadata decoder and processor 503 furthermore is configured to decode and process the encoded metadata 502 and generate rendering metadata 504. Asmentioned above, there various ways to encode the metadata, and also different possible sets of metadata transmitted. Hence, the decoding and processing implemented in some embodiments can vary. In some embodiments the encoded metadata 502 is decoded to generate decoded metadata. The decoded metadata can comprise the followingparameters: decoded MASA metadata; MASA-to-total energy ratios; ISM ratios; andobject directions. In some embodiments the decoded metadata parameters can be processedor converted to a form that is more suitable for rendering to generate the renderingmetadata 504. For example firstly, the direct-to-total energy ratios ^^^^^(^, ^) in the MASAmetadata are modified by multiplying them with the MASA-to-total energy ratio The rest of the MASA metadata (directions ^^^^^^^(^, ^) , spreadcoherences ^^^^^(^, ^) , and surround coherences ^^^^^(^, ^) ) can be usedwithout modifications. The ISM ratios ^^^^(^, ^, ^) can, in some embodiments, be modified by^^^^,^^^^(^, ^, ^) = (1 − ^(^, ^)) ^^^^(^, ^, ^)where ^ is the object index. The object directions can be used withoutmodifications (directions ^^^^^^(^, ^)).In some embodiments the processing can comprise a selection of one or more of the original decoded metadata parameters. The resulting rendering metadata 504 ( ^^^^^^^(^, ^) , ^^^^^,^^^^(^, ^) ,^^^^^(^, ^), ^^^^^(^, ^), ^^^^^^(^, ^), ^^^^,^^^^(^, ^, ^)) can then be provided to thespatial synthesizer 505 as an output of the metadata decoder and processor 503.It should be noted that the rendering metadata 504 in some embodimentsdoes not necessarily directly correspond to the original MASA metadata and objectmetadata (that were input to the metadata encoder as shown in the exampleencoder), as the original metadata was related to the separate transport audio signals, whereas the decoded metadata is related to the combined transport audio signals. Furthermore the generation of the rendering metadata may be implementedin the encoder 101 (as was mentioned above), or it may be performed elsewhere.The rendering metadata 504 can be passed to the spatial synthesizer 505.The encoded separated object audio signal 522 and encoded separatedobject metadata 532 are forwarded to the object decoder 523. The object decoder523 is configured to decode the encoded separated object audio signal 522 andencoded separated object metadata 532 to generate decoded separated objectaudio signal 524 and decoded separated object metadata 534. The decodedseparated object audio signal can comprise, for example, direction ^^^^^^(^) andthe object index ^^^^(^). The decoded separated object audio signal 524 anddecoded separated object metadata 534 are then passed to the spatial synthesizer505. The decoder 111 in some embodiments comprises a spatial synthesizer 505. The spatial synthesizer 505 is configured to receive the rendering metadata 504, the decoded transport audio signals 514, the decoded separated object audiosignal 524 and decoded separated object metadata 534 and the gain controlinformation 112. The spatial synthesizer 505 can then be configured to generate the spatialaudio signals 114 based on the rendering metadata 504, the decoded transportaudio signals 514, the decoded separated object audio signal 524 and the decodedseparated object metadata 534 and the gain control information 112. The gaincontrol information 112, in some embodiments, comprises a gain ^^^^(^, ^) foreach object and ^^^^^(^) for the MASA audio. The spatial audio signals 114 canthen be output. With respect to Figure 6 is shown a flow diagram of the operations of the example decoder as shown in Figure 5. The bitstream is obtained as shown in Figure 6 by step 601. The bitstream is then demultiplexed to generate the encoded metadata, theencoded transport audio signals, the encoded separated object audio signal andthe encoded separated object metadata as shown in Figure 6 by step 603.The encoded transport audio signals are then decoded to generate the decoded transport audio signals as shown in Figure 6 by step 605. The encoded metadata furthermore is decoded and processed to generatethe rendering metadata as shown in Figure 6 by step 606.The encoded separated object audio signal and encoded separated objectmetadata is decoded to generate decoded separated object audio signal anddecoded separated object metadata as shown in Figure 6 by step 607.The gain control information is obtained as shown in Figure 6 by step 602.The spatial audio signals are generated from the decoded transport audiosignals, decoded rendering metadata, decoded separated object audio signal,decoded separated object metadata and gain control information as shown inFigure 6 by step 608. Then the spatial audio signals are output as shown in Figure 6 by step 609. Figure 7 shows in further detail a schematic view of an example spatialsynthesizer 505 as shown in Figure 5 according to some embodiments.The spatial synthesizer 505 is configured to receive the decoded transportaudio signals 514, the gain control information 112, the rendering metadata 504,the decoded separated object audio signal 524 and decoded separated objectmetadata 534.In some embodments the spatial synthesizer 505 comprises a forward filterbank 701 or analysis filter bank. The forward filter bank 701 is configured to receivethe decoded transport audio signals 514 and convert the signals to the time-frequency domain and generate time-frequency transport audio signals 702 to bepassed to the transport signal and metadata processor 705. The time-frequencytransport audio signals 702 ^(^, ^, ^) can be denoted as^ either in vector or scalar form, where ^ is the frequency bin index, ^ is thetime-frequency signal temporal index (or temporal slot index), and ^ is the channelindex. In this specific example, there are exactly two channels. Other number ofchannels can be employed in other examples. For example, the forward filter bank 701 comprises a short-time Fouriertransform (STFT), the complex low-delay filter bank (CLDFB) or complex-modulated quadrature mirror filter (QMF) bank. In some embodiments, the filter bank is configured to have 60 frequency bins, and sufficient stop-band attenuation to avoid significant aliasing to occur when the frequency bin signals are processed. In this configuration, all frequency bins can be processed independently from each other, except that some frequency bins may share the same spatial metadata. For example, the spatial metadata may consist of spatial parameters in a limitednumber of frequency bands, for example, 5 bands, and each of these bandscorrespond to a set of one or more frequency bins provided by the Forward filterbank 701. In the following, however, it is assumed that if the spatial metadata is oflower frequency resolution than, for example, 60 bins, it is mapped (and byrepetition when needed) to the appropriate, for example, 60 bins before theprocessing as described in the following. In some embodiments the spatial synthesiser 505 comprises a transportsignal and metadata processor 705 configued to process the time-frequencytransport audio signals 702 based on the gain control information 112 andrendering metadata 504 to provide processed time-frequency transport signals706, so that they attain level properties for the objects and / or other sounds asdetermined in the gain control information 112. The processed time-frequencytransport signals 706 can be denoted as^ The processed time-frequency transport signals 706 can be then providedto a decorrelator / mixer 707 and a mix matrix determiner 709.As described in detail further below, the transport signal and metadataprocessor 705 can also be configued to also modify the rendering metadata 504based on the gain control information 112 to obtain processed rendering metadata716 which is passed to the mix matrix determiner 709. The spatial synthesizer 505 in some embodiments comprises a mix matrixdeterminer 709 configured to receive the processed time-frequency transportsignals 706 ^(^, ^) , the processed rendering metadata 716 and the decodedseparated object metadata 534. The mix matrix determiner 709 is configured todetermine a mixing matrix that, when applied to the processed time-frequencytransport signals 706, enables a spatialized (e.g., binaural) output to be generated.In some embodiments the mix matrix determiner 709 is configured to initially determine first the processed transport signal covariance matrix where the superscript H indicates a conjugate transpose and ^^(^) and^^(^) are the first and last time-frequency signal temporal indices corresponding tosubframe ^. In this example, there are four time indices ^ at each subframe ^. Assaid, the covariance matrix is determined for each bin. In other embodiments, it could be also averaged (or summed) over multiple frequency bins, in a resolution that approximates human hearing resolutions, or in the resolution of the determined spatial metadata parameters, or any suitable resolution. The mix matrix determiner 709 can then be configured to determine anoverall energy value ^^(^, ^) as the sum of the diagonal values of ^^(^, ^).The mix matrix determiner 709 can furthermore be configured to determinea target covariance matrix, which consists of the levels and correlations for thespatial audio signals, which in this example is a binaural signal. To determine atarget covariance matrix in a binaural form, the mix matrix determiner 709 isconfigured to be able to determine (for example, employing lookup from a database) the head related transfer functions (HRTFs) for any direction of arrival(DOA). The HRTF is denoted ^(^^^, ^) which is a 2x1 column vector havingcomplex-valued gains for left and right ears for bin ^ and direction ^^^ . Thecorresponding HRTF covariance matrix is ^(^^^, ^) = ^(^^^, ^)^^(^^^, ^). Themix matrix determiner 709 can in some embodiments be configured to determineinformation of a diffuse-field covariance matrix ^^^^^(^), which may be formulated,for example, by selecting a spatially equally spaced set of directions ^^^^ where .The target covariance matrix can therefore in some embodiments be determined byThe parameters ^′^^^^,^^^^(^, ^) and ^′^^^,^^^^(^, ^, ^) used in this equationare described in further detail later, and are a part of the processed renderingmetadata. In this example implementation, there is one simultaneous MASAdirection ^^^^^^^(^, ^) and ^^ object directions ^^^^ . In other embodiments,there may be more than one MASA direction, and those directions can be straightforwardly added to the equation above. Similarly, if the metadata indicates,the target covariance matrix could be generated by taking into account variousother features such as coherent or incoherent spatial spreads, spatial coherences, or any other spatial features known in the art. The rendering based on spreadcoherences ^^^^^(^, ^) , surround coherences ^^^^^(^, ^) is described inGB2572650. The mix matrix determiner 709 can then be configured to employ anysuitable method to generate a mixing matrix ^(^, ^) based on the matrices ^^(^, ^)and ^^(^, ^). For example a suitable method has been described inVilkamo, J., Bäckström, T., & Kuntz, A. (2013). Optimized covariance domain framework for time–frequency processing of spatial audio. Journal of the Audio Engineering Society, 61(6), 403-411. The formula provided in the appendix of the above publication can be usedto formulate a mixing matrix ^(^, ^). The same notation for matrices has beenemployed as in the publication for assist the implementation of the publicationapproach to determine the mixing matrix In some embodiments there is also determined a prototype matrix ^ =^ 1 0.050.05 1 ^ that guides the generation of the mixing matrix. The rationale of thesematrices and the formula to obtain a mixing matrix ^(^, ^) based on them has beenthoroughly explained in the above cited publication. In summary, the method issuch that provides a mixing matrix ^(^, ^) that when applied to a signal with acovariance matrix ^^(^, ^) produces a signal with covariance matrix ^^(^, ^), in aleast-squares optimized way. In these examples the prototype matrix ^ is simplythe identity matrix with slight leakage for regularization purposes. Having an identity prototype matrix means that the processing aims to produce an output that is as similar as possible to the input (i.e., with respect to the prototype signals) whileobtaining the target covariance matrix .For simplicity, the above example does not account for head orientation orchanges in head orientation. However when head tracking is enabled, this couldpresent the situation where the user or listener is facing a ‘rear’ direction. In suchexamples the channels of the processed time-frequency transport audio signals 706 can be mutually flipped before the above processing. Furthermore, in some embodiments the processed rendering metadata canbe rotated based on any received head orientation data. Other processing can also be implemented, such as signal-dependently cross-mixing the transport audio signals before the presented rendering operations, if the user is facing side directions. Such and other procedures to account for head-tracked rendering has been described thoroughly in the cited documents and are not repeated here. The mixing matrix determiner 709 in some configurations also determines aresidual processing matrix ^^(^, ^) . In some situations, it is possible that theprocessed transport signals do not have suitable inter-channel incoherenceenabling rendering of incoherent outputs (for example, in situations where there areambience or spread sounds). The determination of the residual processing matrix was also described in the earlier cited publication. In summary the residualprocessing matrix can be determined, after implementing matrix regularizations, todetermine how the processing of the transport signals with ^(^, ^) falls short inobtaining the target covariance matrix . The residual processing matrix canthen be determined such that it is able to process a decorrelated version of theprocessed transport signals ^(^, ^) to obtain that missing portion of the targetcovariance matrix. In other words, the residual processing matrix achieves toproduce a signal with a covariance matrix .The mixing matrix determiner 709 furthemore can be configured to also determines HRTF gains for the separated object for each frequency bin as^^^^^^^^, ^^.The mixing matrix determiner 709 can then be configured to provide themixing matrix ^(^, ^), the residual mixing matrix ^^(^, ^), and the separated objectprocessing gains ^^^^^^^^, ^^ as processing matrices 710 to thedecorrelator / mixer 707. Additionally the spatial synthesizer 505 comprises a forward filter bank 721(or analysis filter bank) configured to receive the decoded separated object audio signal 524 and generate time-frequency object audio signal 722. The forward filter bank 721 in some embodiments can be similar to or implemented with the forward filter bank 701. The spatial synthesizer 505 furthermore comprises a separate object signal processor 703. The separate object signal processor 703 is configured to receive the time-frequency object audio signals 722 and the gain control information 112 and generate a processed time-frequency object audio signals 732 which can be passed to the decorrelator / mixer 707. The spatial synthesizer 505 furthemore comprises a decorrelator / mixer 707.The decorrelator / mixer 707 is configured to receive the processed time-frequencytransport audio signals 706 ^(^, ^), the processed time-frequency object audiosignal 732 ^^^^(^, ^) and the processing matrices 710 ^(^, ^) , ^^(^, ^) and^^^^^^^^, ^^ . The decorrelator / mixer 707 is configured to first process theprocessed time-frequency transport audio signals 706 with decorrelators togenerate decorrelated signals ^^(^, ^) . It then can be configued to apply thefollowing mixing procedure to generate the time-frequency spatial audio signals708 output by the decorrelator / mixer 707. In the above processing, although not explicitly written in the equation, theprocessing matrices may be linearly interpolated between subframes ^ such thatat each temporal index ^ of the time-frequency signal the matrices take a step from^(^, ^ − 1) towards ^(^, ^), and similarly for the other processing coefficients^^(^, ^) and ^^^^^^^^, ^^. The interpolation rate may be adjusted if an onset isdetected (fast interpolation) or not (normal interpolation). The time-frequencyspatial audio signals 708 ^(^, ^) can then be output to a inverse filter bank 711. In some embodiments the spatial synthesizer 505 comprises an inverse filterbank 711 which is configured to apply an inverse transform corresponding to thatused by the forward filter bank 701 / 721 to convert the time-frequency spatial audiosignals 708 to spatial audio signals 114, which is the output of the spatialsynthesizer 505 as shown in Figures 5 and 7 and decoder 111 of Figure 1 and 5, and also of Figure 3. It is understood that the mix matrix determiner 709 and decorrelator / mixer707 represent only one way to synthesize a spatial output signal based on transportsignals (in our example, the processed time-frequency transport signals) andspatial metadata (in our example, the processed rendering metadata), and othermethods of generating the processing matrices 710 and the time-frequency spatialaudio signals 708 (and the spatial audio signals 114) based on determinedprocessed T-F object audio signals 732 and processed T-F transport audio signals 706 are known in the literature. With respect to Figure 8 is shown a flow diagram showing the operations of the spatial synthesiser 505 shown in Figure 7. Thus the decoded transport audio streams are obtained as shown in Figure 8 by step 801. A forward filter bank is configured to time-frequency domain transform the decoded transport audio streams to generate time-frequency transport audio signals as shown in Figure 8 by step 807. Furthermore the rendering metadata is obtained as shown in Figure 8 by step 802. Additionally the decoded separated object metadata is obtained as shown in Figure 8 by step 804. Furthermore the gain control information is obtained as shown in Figure 8by step 803.Furthemore the separated object audio signal is obtained as shown inFigure 8 by step 805. A forward filter bank is configured to time-frequency domain transform the separated object audio signal to generate time-frequency separated object audio signal as shown in Figure 8 by step 806.A processed time-frequency transport signals and processed renderingmetadata based on the time-frequency transport signals, rendering metadata andgain control information is determined as shown in Figure 8 by step 809. Furthermore a processed time-frequency object signal based on a time- frequency object signal and gain control information is obtained as shown in Figure 8 by step 808. Then there is a determination of a mix matrix (or more generally processingmatrices) based on processed time-frequency transport signals, processedrendering metadata and decoded separated object metadata as shown in Figure 8 by step 810. Then decorrelation can be applied and processing matrices also applied tothe processed T-F transport signals and processed T-F object signal to generateT-F spatial audio signals as shown in Figure 8 by step 811.Then an inverse filter bank (or synthesis filter bank) is applied to the time- frequency domain spatial audio signals as shown in Figure 8 by step 813 to generate the spatial audio signals. The spatial audio signals can then be output as shown in Figure 8 by step 815. The transport signal and metadata processor 705 and separate object signalprocessor 703 are further described hereafter.As described above the transport signal and metadata processor 705 isconfigured to receive the time-frequency transport signals 702, gain controlinformation 112 and the rendering metadata 504. The aim of the transport signaland metadata processor 705 is to formulate target gain coefficients for every time-frequency instance based on the parametric representation of the objects andMASA in the rendering metadata 504 and target gain coefficients in gain controlinformation 112. In other words, the determined or formulated target gaincoefficients should amplify or attenuate only the intended energetic parts of thetime-frequency transport signals 702.The transport signal and metadata processor 705 thus in someembodiments attempts to generate processed time-frequency transport signals 706which have modified time-frequency energetic content based on the presentedgaining procedure below. Furthermore, the transport signal and metadataprocessor 705 is configured to produce processed rendering metadata 716 whichcomprises modified metadata values according to the performed gaining operations. The transport signal and metadata processor 705 receives the time-frequency transport signals 702 in frequency bins ^ and temporal indices ^, andfurther determines total channel energetic values and total energetic value per frame ^. The total channel energetic values can be determined by And the total energy value can be determined by The transport signal and metadata processor 705 can also receive therendering metadata which comprises the following parameters (as also referencedabove): MASA directions ^^^^^^^(^, ^);MASA direct-to-total energy ratio ^^^^^,^^^^(^, ^);spread coherences ;surround coherences ^^^^^(^, ^);Object directions ^^^^^^(^, ^);ISM .As stated above, in some embodiments, where the spatial metadata is of lower frequency resolution than the time-frequency audio signals (for example, fewer than the example 60 bins), the transport signal and metadata processor 705is configured to map metadata (by repetition when needed) to the appropriate binsbefore the processing. In addition, the transport signal and metadata processor 705 can receivegain control information, which comprises gain coefficients ^^^^(^, ^) and, for each object ^ and MASA. In some example embodiments, the gaincoefficient is a linear value, in other words, if the target amplification of the modified object is +6 dB, the related gain coefficient is 2.The transport signal and metadata processor 705 in some embodiments isconfigured to have (or have access to) the information of a panning function how the object signals have been mixed into the transport audio signals. The panningfunction provides panning gains ^(^^^, ^) for each channel ^ for any ^^^. Forexample, the panning function could be the tangent panning law for loudspeakersat ±30 degrees, such that any angle beyond this interval is hard panned to thenearest loudspeaker (except for the rear ±30 arc which could also use the samepanning rule). Another optional embodiment is that the panning follows a cardioidpattern shape towards left or right directions. Any further panning rule can be employed in some optional embodiments, as long as the decoder knows which panning rule was applied by the encoder. This may be known (e.g., fixed), or signalled among the spatial metadata, for example, as an index value to a table containing a set of pre-determinedpanning rules. Regardless of the panning rule, in the following example, thepanning gains are assumed to be limited between 0 and 1, and that the square sum of the panning gains is always 1. The transport signal and metadata processor 705 in some embodimentsapplies the following steps for each frequency bin ^ and subframe ^ in everychannel ^. First, for each object ^, the following steps are performed:- Determining total original object energetic value ^^^^(^, ^, ^) based onthe object ratio ^^^^,^^^^(^, ^, ^) and total energetic value ^(^, ^):^^^^(^, ^, ^) = ^^^^,^^^^(^, ^, ^)^(^, ^),- Determining object energetic pan values as^(^, ^, ^) = (^(^^^^^^(^, ^), ^))^ -Formulating target channel energetic value of object^^^^^ (^, ^, ^, ^) = ^ ^^^^ (^, ^) ^(^, ^, ^)^^^^(^, ^, ^)Then, for modifying the level of the MASA part, the following steps are performed:- Channel-specific MASA energetic value is determined as the remainderof the total original channel energy subtracted by the total original channel energy value of each object -Target channel energetic value for MASA is determined as^^^^^^ (^, ^, ^) = ^ ^^^^^ (^)^^^^^(^, ^, ^)The gain value ^^(^, ^, ^) is determined based on the square root ratio of totaltarget channel energy and total original channel energy The obtained total target and total original energy values may be temporallysmoothed, for example, using an infinite impulse response (IIR) or finite impulseresponse (FIR) filter, before the gains are computed.To obtain processed time-frequency transport audio signals 706 ^(^, ^), thedetermined gain values ^^(^, ^, 1) and ^^(^, ^, 2) are applied to the time-frequencytransport audio signals 702 ^(^, ^)^ It should be noted that, as the applied gain coefficients are determined within a temporal accuracy of a subframe ^, the same target gain coefficient is applied forevery temporal index ^ within the subframe ^. The gains may also be interpolatedfor different slots ^ between the values for the adjacent subframes so that thevalues change more smoothly. In some embodiments the energetic gain value ^^^(^, ^, ^) is determinedbased on the ratio of total target channel energy and total original channel energy The gain value ^^(^, ^, ^) is determined based on the energetic gain value, forexample, with the square root In some embodiments the square root is some other root, some other function, or may be omitted completely. The obtained total target and total original energy values may, in a similarmanner to the method above, be temporally smoothed, for example, using aninfinite impulse response (IIR) or finite impulse response (FIR) filter, before the gains are computed. For generating the processed rendering metadata 716, ISM ratios^^^^,^^^^(^, ^, ^) and MASA ratio ^^^^^,^^^^(^, ^) are modified according to the gaincontrol information 112.For example, firstly, the new total ratio ^′^^^^^is determined: where the diffuse-to-total energy ratio ^^^^^is defined as: The new ISM ratios and the MASA ratio are then obtained from the ratios between the original ratios and the new total ratio. Thus, the modified ISM ratios are: And, the modified MASA ratio is: In addition to the obtained new ISM ratios ^^^^^,^^^^(^, ^, ^) and new MASAdirect-to-total energy ratio ^^^^^^,^^^^(^, ^), the processed rendering metadata 716comprises unmodified metadata parameters of rendering metadata 504:Object directions ^^^^^^(^, ^);MASA directions ^^^^^^^(^, ^);spread coherences ;surround coherences ^^^^^(^, ^).In some embodiments the presented processing steps described above caninclude the option that if a certain object and / or MASA is not modified, thecorresponding ^^^^(^, ^) and / or ^^^^^(^) is 1.In some embodiments the separate object signal processor 703 receivestime-frequency object audio signal 722 ^^^^(^, ^) and gain control information 112.Based on the gain control information 112, the separate object signal processor703 can apply target amplification or attenuation to the time-frequency object audiosignal 722. Gain coefficients for separated objects can be obtained from the gaincontrol information 112 as ^^^^(^) = As the time-frequency object audio signal 722 comprises only the time-frequency signal of the separated object, the gain coefficient determined in the gaincontrol information 112 can be applied directly: As previously, the gains may also be temporally interpolated for differenttemporal indices ^.The separate object signal processor 703 can then be configured to produceprocessed time-frequency object audio signals 732 ^^^^(^, ^).In some alternative embodiments, no objects are separated from the mix.Thus, processing related to the separated object audio signal 524 and separatedobject metadata 534 can be omitted in these embodiments.In some alternative embodiments, there may be two directions in the MASA datastream. In some alternative embodiments, the metadata can comprise other parameters additionally or instead of the parameters presented above. In some alternative embodiments, the gain editing can be implemented jointly with the position edition according to GB2104309.6. In this case, the processing matrices to be applied on the audio signals can be combined, smoothed together, and applied only once to the audio signals. In a further alternative embodiment, the operations presented herein are combined with the operations of GB2104309.6 such that when determining the target channel energetic value, the object energetic pan values are determinedbased on panning gains ^(^^^, ^) when the direction of the object is not modifiedor using panning gains of the modified direction as described in GB2104309.6 forthe objects whose direction is modified. In a further alternative embodiment, the operations presented herein are combined with the operations of GB2104309.6 such that target energetic values are defined as presented above, and direction modifications and equalization asdescribed in GB2104309.6 are done based on the determined target energeticvalues. In such alternative embodiment, before determing gain values ^^(^, ^, ^)and applying them to the input signal, the processing steps described inGB2104309.6 are done based on the determined total target energies^^^^^^^(^, ^, ^) as described above. Gaining is then applied after the requiredprocessing steps described in GB2104309.6. In some alternative embodiments, the gain editing can be implemented to a mono transport signal. In this case, the target energetic values are determined based only on the rendering metadata and determined total energetic values. In some alternative embodiments, a new total ratio for obtaining the processed rendering metadata may comprise different weighting of applied ratio components for calculations. In some examples, the gain control information may be in part or fully based on a set distance parameter, so that higher distance entails lower gain values. Furthermore in some embodiments, the gain control information can be frequency-dependent, for example, when it is based on a set distance parameter, where a larger distance may cause more attenuation at high frequencies than at low frequencies. In some embodiments, the operations are performed on bands comprising one or more bins ^. In such embodiments the signal energy values are summed over the bins in the band, and the determined gain control information is expanded to all bins in the band before applying on the signal. In some further embodiments, the input may be a combination of the audioobjects and some other spatial audio format, instead of MASA. For example, aninput fomat can comprise: objects + Ambisonics or objects + multi-channel audiosignals (e.g., 5.1 or 7.1+4) or any suitable combination. In such embodiments, theAmbisonics or multi-channel audio signals may, for example, be first converted toa MASA stream, and then the processing can be applied as presented herein. Alternatively, the Ambisonics or multi-channel audio signals may be converted to some other suitable parametric spatial audio format, and the subsequent processing parts can be adjusted to be compatible with that format. With respect to Figure 9 an example electronic device which may be usedas the computer, encoder processor, decoder processor or any of the functionalblocks described herein is shown. The device may be any suitable electronicsdevice or apparatus. For example, in some embodiments the device 1600 is amobile device, user equipment, tablet computer, computer, audio playbackapparatus, a laptop, or a teleconferencing system. In some embodiments the device 1600 comprises at least one processor or central processing unit (CPU or processor) 1607. The processor 1607 can be configured to execute various program codes such as the methods such as described herein. The device 1600 furthermore comprises a transceiver 1609 which isconfigured to receive the bitstream and provide it to the processor 1607. Typically,the connection is wirelessly received data from a remote device or a server, however, in some embodiments the bitstream is received via a wired connection or read from a local memory of the device. The transceiver can communicate with further apparatus by any suitable known communications protocol. For example in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA). The device may furthermore comprise a user interface (UI) 1605 which maydisplay to the user an interface changing the loudness of the audio objects and the MASA stream, for example, using sliders. This loudness information is the gaincontrol information 1615 provided to the processor, CPU, 1607.The device 1600 may further comprise memory (MEM) 1611 which iscoupled to the processor 1607. In some embodiments the memory 1611 comprisesthe program code 1621 which is executed by the processor 1607. The programcode may involve instructions to perform the operations of the spatial synthesizerdescribed above. The processor 1607 can then be configured to output the spatialaudio signals, which in this example was a binaural output, to a digital to analogueconverter (DAC) / Bluetooth 1601 converter.The combination of the processor, CPU, 1607 and memory, MEM, can implement the IVAS decoder 1631 functionality described above. The DAC / Bluetooth 1601 is configured to convert the spatial audio signalsto an analogue form if the headphones are conventional wired (analogue)headphones. For wireless connections, the DAC / Bluetooth 1601 may be aBluetooth transceiver. The DAC / Bluetooth 1601 block provides (either wired or wirelessly) thespatial audio to be played back with the headphones 1603 to the user. In someembodiments, the headphones 1603 may have a head tracker which may provideorientation and / or position information of the user’s head to the processor 1607 ofthe rendering apparatus, so that user’s head orientation is accounted for at the spatial synthesizer. In some embodiments the remote device (not shown in Figure 11) may generate the bitstream in various ways. In one situation, the remote device consists of multiple devices, for example, a device with a microphone array at a room with multiple participants, and multiple other devices with near-microphones (e.g., headset microphones) of remote participants. The microphone array may generate the MASA stream, and the remote participants may generate single-channel audio streams treated as object signals. Depending on the bit rates, these streams may be combined by a server, and conveyed to the device of Figure 11. In another example, the MASA stream is a captured spatial stream, for example, an audio recording at a sports event, and the object stream would originate from a commentator. For the present invention, the bitstream may originate from any kind of a setting. The device of Figure 11 may also capture the audio locally, and transmit itto a remote device, where the remote device may perform the rendering similarly to the device of Figure 11. In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media, and optical media. The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples. Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate. Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication. The foregoing description has provided by way of exemplary and non- limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of thisinvention will still fall within the scope of this invention as defined in the appendedclaims.

Claims

CLAIMS:

1. An apparatus for rendering a spatial audio signal based on a spatial audiostream comprising: at least one audio signal; and metadata associated with the atleast one audio signal, the apparatus comprising means configured to:obtain the spatial audio stream comprising: the at least one audio signal;and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to defineat least one audio object portion position and at least one audio object portionenergy proportion; obtain gain control information; determine gain processing information based on:the gain control information; andthe at least one audio object portion energy proportion; and render the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal.

2. The apparatus as claimed in claim 1, wherein the gain control informationcomprises at least one gain for at least one of: the at least one audio object portion; and the at least one other spatial audio portion.

3. The apparatus as claimed in claim 2, wherein the at least one gain for the atleast one audio object signal is based on a distance parameter associated with an at least one audio object of the at least one audio object portion.

4. The apparatus as claimed in any of claims 1 or 2, wherein the meansconfigured to determine the gain processing information based on the gain controlinformation and the at least one audio object portion energy proportion is configuredto determine at least one first gain value based on the at least one audio signal andthe at least one audio object portion energy proportion.

5. The apparatus as claimed in claim 4, wherein the means is furtherconfigured to apply the at least one first gain value to one of: the at least one audioobject portion; and the at least one other spatial audio portion.

6. The apparatus as claimed in any of claims 4 or 5, wherein the meansconfigured to determine gain processing information is further configured to determine gain processing information based on the at least one audio object portion position.

7. The apparatus as claimed in any of claims 1 to 6, wherein the at least oneother spatial audio portion comprises at least two transport audio signals.

8. The apparatus as claimed in any of claims 1 to 7, wherein the meansconfigured to obtain the spatial audio stream comprising: the at least one audiosignal; and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is configured to define the at least one audio object portion position and at least one audio object portionenergy proportion is configured to perform at least one of:receive information defining the at least one audio object portion position and at least one audio object portion energy proportion; and receive at least one parameter value defining the at least one audio object portion position and at least one audio object portion energy proportion associated with the at least one object.

9. The apparatus as claimed in any of claims 1 to 8, wherein the means isfurther configured to process the metadata associated with the at least one audiosignal based on the gain control information.

10. The apparatus as claimed in claim 9, wherein the means configured toprocess the metadata associated with the at least one audio signal based on thegain control information is configured to at least one of:process metadata associated with the at least one audio object portion of the at least one audio signal; and process metadata associated with the at least other spatial audio portion ofthe at least one audio signal.

11. The apparatus as claimed in any of claims 1 to 10, wherein the at least oneaudio object portion and at least one other spatial audio portion are a mixture withinthe at least one audio signal.

12. A method for an apparatus for rendering a spatial audio signal based on aspatial audio stream comprising: at least one audio signal; and metadata associated with the at least one audio signal, the method comprising: obtaining the spatial audio stream comprising: the at least one audio signal;and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to define at least one audio object portion position and at least one audio object portion energy proportion; obtaining gain control information;determining gain processing information based on:the gain control information; and the at least one audio object portion energy proportion; and rendering the spatial audio signal based on the gain processing information,the at least one audio signal, and the metadata associated with the at least one audio signal.

13. The method as claimed in claim 12, wherein the gain control informationcomprises at least one gain for at least one of: the at least one audio object portion; and the at least one other spatial audio portion.

14. The method as claimed in claim 13, wherein the at least one gain for the atleast one audio object signal is based on a distance parameter associated with an at least one audio object of the at least one audio object portion.

15. The method as claimed in any of claims 12 or 13, wherein determining thegain processing information based on the gain control information and the at leastone audio object portion energy proportion comprises determining at least one firstgain value based on the at least one audio signal and the at least one audio objectportion energy proportion.

16. The method as claimed in claim 15, further comprising applying the at leastone first gain value to one of: the at least one audio object portion; and the at least one other spatial audio portion.

17. The method as claimed in any of claims 15 or 16, wherein determining gainprocessing information further comprises determining gain processing informationbased on the at least one audio object portion position.

18. The method as claimed in any of claims 12 to 17, wherein the at least oneother spatial audio portion comprises at least two transport audio signals.

19. The method as claimed in any of claims 12 to 18, wherein obtaining thespatial audio stream comprises at least one of: receiving information defining the at least one audio object portion position and at least one audio object portion energy proportion; and receiving at least one parameter value defining the at least one audio objectportion position and at least one audio object portion energy proportion associated with the at least one object.

20. The method as claimed in any of claims 12 to 19, further comprisingproessing the metadata associated with the at least one audio signal based on thegain control information.

21. The method as claimed in claim 20, wherein processing the metadataassociated with the at least one audio signal based on the gain control informationcomprises at least one of: processing metadata associated with the at least one audio object portion of the at least one audio signal; and processing metadata associated with the at least other spatial audio portion of the at least one audio signal.

22. The method as claimed in any of claims 12 to 21, wherein the at least oneaudio object portion and at least one other spatial audio portion are a mixture within the at least one audio signal.

23. An apparatus for rendering a spatial audio signal based on a spatial audiostream, the apparatus comprising at least one processor and at least one memoryincluding a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: obtain the spatial audio stream comprising: the at least one audio signal;and the metadata associated with the at least one audio signal, wherein the at least one audio signal comprises at least one audio object portion and at least one other spatial audio portion, and the associated metadata is at least configured to defineat least one audio object portion position and at least one audio object portionenergy proportion; obtain gain control information; determine gain processing information based on:the gain control information; andthe at least one audio object portion energy proportion; and render the spatial audio signal based on the gain processing information, the at least one audio signal, and the metadata associated with the at least one audio signal.

Citation Information

Patent Citations

  • Parametric spatial audio encoding

    GB202217905D0

  • Spatial audio parameters and associated spatial audio playback

    GB2572650A

  • Determination of spatial audio parameter encoding and associated decoding

    GB2575305A

  • Determination of spatial audio parameter encoding and associated decoding

    GB2587196A

  • Interactive audio rendering of a spatial stream

    GB2605190A