Spatial audio representation and rendering

By adjusting rendering parameters to avoid spectral artefacts, the spatial audio rendering system maintains audio quality and reverberance, addressing issues with short reverberation times in IVAS parametric binaural rendering.

WO2026057289A1PCT designated stage Publication Date: 2026-03-19NOKIA TECHNOLOGIES OY

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing spatial audio rendering systems, particularly those using the IVAS parametric binaural renderer, suffer from spectral artefacts such as sound coloration and comb filtering due to short reverberation times, leading to decreased audio quality and speech intelligibility.

Method used

Intelligently control the rendering parameters of direct-sound and late-reverberation by adjusting late-reverberation energy parameters and increasing direct-sound energy parameters when reverberation times are below a threshold, avoiding spectral artefacts while maintaining desired reverberance.

Benefits of technology

This approach mitigates spectral artefacts and maintains perceived reverberance, ensuring high-quality audio reproduction even with short reverberation times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025073452_19032026_PF_FP_ABST
    Figure EP2025073452_19032026_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus for rendering spatial audio with reverberation receive a spatial audio stream, the spatial audio stream comprising at least one audio signal and spatial metadata associated with the at least one audio signal The apparatus further obtains at least one room effect parameter comprising at least one reverberation time, adjusts at least one of the at least one room effect parameter when the at least one reverberation time is less than a determined threshold value and generates a spatial audio signal using the spatial audio stream and the adjusted at least one room effect parameter, the generated spatial audio signal comprising a first spatial audio portion and processed dependent on the spatial metadata and a second spatial audio portion generated based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]SPATIAL AUDIO REPRESENTATION AND RENDERINGField The present application relates to apparatus and methods for control of reverberation within spatial audio representation and rendering, but not exclusively for control of rendering within parametric spatial audio representation and rendering. Background Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An example ofsuch a codec is the Immersive Voice and Audio Services (IVAS) codec which is beingdesigned to be suitable for use over a communications network such as a 3GPP 4G / 5G network including use in such immersive services as for example immersive voice andaudio for virtual reality (VR). This audio codec is expected to handle the encoding,decoding and rendering of speech, music and generic audio. It is furthermore expected to support channel-based audio and scene-based audio inputs including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency to enable conversational services as well as support high error robustness under various transmission conditions. Input signals can be presented to the IVAS encoder in one of a number ofsupported formats (and in some allowed combinations of the formats). For example amono audio signal (without metadata) may be encoded using an Enhanced Voice Service(EVS) encoder. Other input formats may utilize new IVAS encoding tools. One inputformat proposed for IVAS is the Metadata-assisted spatial audio (MASA) format, wherethe encoder may utilize, e.g., a combination of mono and stereo encoding tools andmetadata encoding tools for efficient transmission of the format. MASA is a parametricspatial audio format suitable for spatial audio processing. Parametric spatial audio processing is a field of audio signal processing where the spatial aspect of the sound (or sound scene) is described using a set of parameters. For example, in parametric spatial audio capture from microphone arrays, it is a typical and an effective choice to estimate from the microphone array signals a set of parameters such as directions of the sound in frequency bands, and the relative energies of the directional and non-directional parts of the captured sound in frequency bands, expressed for example as a direct-to-total ratioor an ambient-to-total energy ratio in frequency bands. These parameters are known towell describe the perceptual spatial properties of the captured sound at the position of the microphone array. These parameters can be utilized in synthesis of the spatial sound accordingly, for headphones binaurally, for loudspeakers, or to other formats, such as Ambisonics. For example, there can be two channels (stereo) of audio signals and spatialmetadata. The spatial metadata may furthermore define parameters such as: Direction index, describing a direction of arrival of the sound at a time-frequency parameter interval; level / phase differences; Direct-to-total energy ratio, describing an energy ratio for the direction index; Diffuseness; Coherences such as Spread coherence describing a spread of energy for the direction index; Diffuse-to-total energy ratio, describing an energy ratio of non-directional sound over surrounding directions; Surround coherence describing a coherence of the non-directional sound over the surrounding directions; Remainder-to- total energy ratio, describing an energy ratio of the remainder (such as microphone noise) sound energy to fulfil requirement that sum of energy ratios is 1; Distance, describing a distance of the sound originating from the direction index in meters on a logarithmic scale; covariance matrices related to a multi-channel loudspeaker signal, or any data related to these covariance matrices; other parameters guiding a specific decoder, e.g., centre prediction coefficients and one-to-two decoding coefficients (used, e.g., in MPEGSurround). Any of these parameters can be determined in frequency bands.Listening to natural audio scenes in everyday environment is not only about sounds at particular directions. Even without background ambience, it is typical that the majority of the sound energy arriving to the ears is not from direct sounds but indirect sounds from the acoustic environment (i.e., reflections and reverberation). Based on theroom effect, involving discrete reflections and reverberation, the listener auditorilyperceives the source distance and room characteristics (small, big, damp, reverberant)among other features, and the room adds to the perceived feel of the audio content. In other words, the acoustic environment is an essential and perceptually relevant feature of spatial sound. The listener will listen to music in normal rooms (as opposed to, e.g. anechoicchambers), and music (e.g., stereo or 5.1 content) is typically produced in a way that it is expected to be listened in a room with normal reverberation, which creates envelopment and spaciousness to the sound. Listening to normal music in an anechoic chamber is known to be unpleasant due to lack of room effect. Hence, normal music should be (and basically always is) listened to in normal rooms with reverberation. Summary There is provided according to a first aspect an apparatus for rendering spatialaudio with reverberation, the apparatus comprising at least one processor and at leastone memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause theapparatus at least to: receive a spatial audio stream, the spatial audio stream comprisingat least one audio signal and spatial metadata associated with the at least one audio signal; obtain at least one room effect parameter comprising at least one reverberationtime; and adjust at least one of the at least one room effect parameter when the at leastone reverberation time is less than a determined threshold value; and generate a spatialaudio signal using the spatial audio stream and the adjusted at least one room effectparameter, the generated spatial audio signal comprising: a first spatial audio portiongenerated based on the at least one audio signal and processed dependent on the spatial metadata; and a second spatial audio portion generated based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter. The apparatus caused to generate the spatial audio signal using the spatial audio stream and the adjusted at least one room effect parameter may be further caused to: generate the first spatial audio portion based on the at least one audio signal and processed dependent on the spatial metadata; separately generate the second spatial audio portion based on the at least one audio signal and processed dependent on theadjusted at least one room effect parameter; and combine the first spatial audioportion and the second spatial audio portion to generate the spatial audio signal. The first spatial audio portion may be at least one of: a direct audio signal portionof the spatial audio signal; an early reverberation signal portion of the spatial audio signal.The first spatial audio portion may be further processed dependent on the adjustedat least one room effect parameter. The second spatial audio portion may be a late reverberation portion of the spatial audio signal. The second spatial audio portion may be further processed dependent on thespatial metadata. The adjusted at least one room effect parameter may comprise at least one of: atleast one adjusted early part energy correction coefficient; at least one adjusted late partenergy correction coefficient; at least one adjusted early part energy parameter; at leastone adjusted late part energy parameter; at least one adjusted early part energycoefficient; at least one adjusted late part energy coefficient; and at least one adjustedreverberation time. The apparatus caused to adjust at least one of the at least one room effectparameter when the at least one reverberation time is less than a determined thresholdvalue may be caused to: generate a reverberation time modification parameter based onthe reverberation time and the determined threshold value; generate an at least oneadjusted reverberation time by processing the at least one reverberation time with the reverberation time modification parameter. The apparatus caused to adjust at least one of the at least one room effectparameter when the at least one reverberation time is less than a determined thresholdvalue may be caused to: generate an energy modification value based on the at least oneadjusted reverberation time and the at least one reverberation time; generate at least oneadjusted early part energy correction coefficient based on an early part energy correctioncoefficient and the energy modification value.The apparatus caused to adjust at least one of the at least one room effectparameter when the at least one reverberation time is less than a determined threshold value may be further caused to generate at least one adjusted late part energy correctioncoefficient based on a late part energy correction coefficient and the energy modificationvalue. According to a second aspect there is provided a method for rendering spatialaudio with reverberation, the method comprising: receiving a spatial audio stream, thespatial audio stream comprising at least one audio signal and spatial metadata associatedwith the at least one audio signal; obtaining at least one room effect parameter comprisingat least one reverberation time; and adjusting at least one of the at least one room effect parameter when the at least one reverberation time is less than a determined thresholdvalue; and generating a spatial audio signal using the spatial audio stream and theadjusted at least one room effect parameter, the generated spatial audio signalcomprising: a first spatial audio portion generated based on the at least one audio signaland processed dependent on the spatial metadata; and a second spatial audio portiongenerated based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter. Generating the spatial audio signal using the spatial audio stream and the adjustedat least one room effect parameter may comprise: generating the first spatial audio portionbased on the at least one audio signal and processed dependent on the spatial metadata;separately generating the second spatial audio portion based on the at least one audiosignal and processed dependent on the adjusted at least one room effect parameter; andcombining the first spatial audio portion and the second spatial audio portion to generatethe spatial audio signal. The first spatial audio portion may comprise at least one of: a direct audio signalportion of the spatial audio signal; an early reverberation signal portion of the spatial audiosignal. The first spatial audio portion may be further processed dependent on the adjustedat least one room effect parameter. The second spatial audio portion may be a late reverberation portion of the spatial audio signal. The second spatial audio portion may be further processed dependent on thespatial metadata. The adjusted at least one room effect parameter may comprise at least one of: atleast one adjusted early part energy correction coefficient; at least one adjusted late partenergy correction coefficient; at least one adjusted early part energy parameter; at leastone adjusted late part energy parameter; at least one adjusted early part energycoefficient; at least one adjusted late part energy coefficient; and at least one adjustedreverberation time. Adjusting the at least one of the at least one room effect parameter when the atleast one reverberation time is less than a determined threshold value may comprise:generating a reverberation time modification parameter based on the reverberation timeand the determined threshold value; generating an at least one adjusted reverberationtime by processing the at least one reverberation time with the reverberation time modification parameter. Adjusting at least one of the at least one room effect parameter when the at leastone reverberation time is less than a determined threshold value may comprise:generating an energy modification value based on the at least one adjusted reverberationtime and the at least one reverberation time; generating at least one adjusted early partenergy correction coefficient based on an early part energy correction coefficient and theenergy modification value. Adjusting at least one of the at least one room effect parameter when the at leastone reverberation time is less than a determined threshold value may further comprisegenerating at least one adjusted late part energy correction coefficient based on a latepart energy correction coefficient and the energy modification value.According to a third aspect there is provided an apparatus comprising meansconfigured to: receive a spatial audio stream, the spatial audio stream comprising at leastone audio signal and spatial metadata associated with the at least one audio signal; obtain at least one room effect parameter comprising at least one reverberation time; and adjust at least one of the at least one room effect parameter when the at least onereverberation time is less than a determined threshold value; and generate a spatial audiosignal using the spatial audio stream and the adjusted at least one room effect parameter,the generated spatial audio signal comprising: a first spatial audio portion generatedbased on the at least one audio signal and processed dependent on the spatial metadata; and a second spatial audio portion generated based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter. The means configured to generate the spatial audio signal using the spatial audio stream and the adjusted at least one room effect parameter may be further configured to: generate the first spatial audio portion based on the at least one audio signal and processed dependent on the spatial metadata; separately generate the second spatial audio portion based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter; and combine the first spatial audio portion and the second spatial audio portion to generate the spatial audio signal. The first spatial audio portion may be at least one of: a direct audio signal portion of the spatial audio signal; an early reverberation signal portion of the spatial audio signal. The first spatial audio portion may be further processed dependent on the adjusted at least one room effect parameter. The second spatial audio portion may be a late reverberation portion of the spatial audio signal. The second spatial audio portion may be further processed dependent on the spatial metadata. The adjusted at least one room effect parameter may comprise at least one of: at least one adjusted early part energy correction coefficient; at least one adjusted late part energy correction coefficient; at least one adjusted early part energy parameter; at least one adjusted late part energy parameter; at least one adjusted early part energy coefficient; at least one adjusted late part energy coefficient; and at least one adjusted reverberation time. The means configured to adjust at least one of the at least one room effect parameter when the at least one reverberation time is less than a determined threshold value may be configured to: generate a reverberation time modification parameter based on the reverberation time and the determined threshold value; generate an at least one adjusted reverberation time by processing the at least one reverberation time with the reverberation time modification parameter. The means configured to adjust at least one of the at least one room effect parameter when the at least one reverberation time is less than a determined threshold value may be configured to: generate an energy modification value based on the at least one adjusted reverberation time and the at least one reverberation time; generate at least one adjusted early part energy correction coefficient based on an early part energycorrection coefficient and the energy modification value.The means configured to adjust at least one of the at least one room effect parameter when the at least one reverberation time is less than a determined thresholdvalue may be further configured to generate at least one adjusted late part energycorrection coefficient based on a late part energy correction coefficient and the energymodification value. According to a fourth aspect there is provided an apparatus for rendering spatialaudio with reverberation, the apparatus comprising: receiving circuitry configured toreceive a spatial audio stream, the spatial audio stream comprising at least one audiosignal and spatial metadata associated with the at least one audio signal; obtaining circuitry configured to obtain at least one room effect parameter comprising at least onereverberation time; and adjusting circuitry configured to adjust at least one of the at leastone room effect parameter when the at least one reverberation time is less than a determined threshold value; and generating circuitry configured to generate a spatial audio signal using the spatial audio stream and the adjusted at least one room effect parameter, the generated spatial audio signal comprising: a first spatial audio portion generated based on the at least one audio signal and processed dependent on the spatial metadata; and a second spatial audio portion generated based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter.. According to a fifth aspect there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] forcausing an apparatus for rendering spatial audio with reverberation to perform at leastthe following: receiving a spatial audio stream, the spatial audio stream comprising atleast one audio signal and spatial metadata associated with the at least one audio signal; obtaining at least one room effect parameter comprising at least one reverberation time; and adjusting at least one of the at least one room effect parameter when the at least one reverberation time is less than a determined threshold value; and generating a spatial audio signal using the spatial audio stream and the adjusted at least one room effect parameter, the generated spatial audio signal comprising: a first spatial audio portion generated based on the at least one audio signal and processed dependent on the spatial metadata; and a second spatial audio portion generated based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter. According to a sixth aspect there is provided a non-transitory computer readable medium comprising program instructions for causing an apparatus for rendering spatialaudio with reverberation to perform at least the following: receiving a spatial audio stream,the spatial audio stream comprising at least one audio signal and spatial metadata associated with the at least one audio signal; obtaining at least one room effect parameter comprising at least one reverberation time; and adjusting at least one of the at least one room effect parameter when the at least one reverberation time is less than a determined threshold value; and generating a spatial audio signal using the spatial audio stream and the adjusted at least one room effect parameter, the generated spatial audio signal comprising: a first spatial audio portion generated based on the at least one audio signal and processed dependent on the spatial metadata; and a second spatial audio portion generated based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter. According to a seventh aspect there is provided an apparatus for rendering spatial audio with reverberation, the apparatus comprising: means for receiving a spatial audio stream, the spatial audio stream comprising at least one audio signal and spatial metadata associated with the at least one audio signal; means for obtaining at least one room effect parameter comprising at least one reverberation time; and means for adjusting at least one of the at least one room effect parameter when the at least one reverberation time is less than a determined threshold value; and means for generating a spatial audio signal using the spatial audio stream and the adjusted at least one room effect parameter, the generated spatial audio signal comprising: a first spatial audio portion generated based on the at least one audio signal and processed dependent on the spatial metadata; and a second spatial audio portion generated based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter. According to an eighth aspect there is provided a computer readable medium comprising program instructions for causing an apparatus for rendering spatial audio with reverberation to perform at least the following: receiving a spatial audio stream, the spatial audio stream comprising at least one audio signal and spatial metadata associated with the at least one audio signal; obtaining at least one room effect parameter comprising at least one reverberation time; and adjusting at least one of the at least one room effect parameter when the at least one reverberation time is less than a determined threshold value; and generating a spatial audio signal using the spatial audio stream and the adjusted at least one room effect parameter, the generated spatial audio signal comprising: a first spatial audio portion generated based on the at least one audio signal and processed dependent on the spatial metadata; and a second spatial audio portion generated based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter. An apparatus comprising means for performing the actions of the method as described above. An apparatus configured to perform the actions of the method as described above. A computer program comprising program instructions for causing a computer to perform the method as described above. A computer program product stored on a medium may cause an apparatus to perform the method as described herein. An electronic device may comprise apparatus as described herein. A chipset may comprise apparatus as described herein. Embodiments of the present application aim to address problems associated with the state of the art. Summary of the Figures For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which: Fig.1 shows schematically a system of apparatus suitable for implementing someembodiments; Fig.2 shows a flow diagram of the operation of the example apparatus accordingto some embodiments; Fig.3 shows schematically a synthesis processor as shown in Fig.1 according tosome embodiments; Fig.4 shows a flow diagram of the operation of the example apparatus as shownin Fig.3 according to some embodiments; Fig.5 shows schematically an added room effect parameter controller as shown inFig.3 according to some embodiments;Fig.6 shows a flow diagram of the operation of the example apparatus as shownin Fig.5 according to some embodiments; Figs.7a and 7b show graphs of the operation of the added room effect parameter controller according to some embodiments; Figs.8a and 8b show further graphs of the operation of the added room effect parameter controller according to some embodiments; Fig.9 shows an example rendering device according to some embodiments; and Fig.10 shows an example device suitable for implementing the apparatus shownin previous figures. Embodiments of the Application The following describes in further detail suitable apparatus and possible mechanisms for the control and addition of room effect to rendered spatial metadataassisted audio signals.Although the following examples focus on MASA encoding and decoding, it shouldbe noted that the presented methods are applicable to any system that utilizes transport audio signals and spatial metadata. The spatial metadata may include, e.g., some of the following parameters in any kind of combination: Directions; Level / phase differences; Direct-to-total-energy ratios; Diffuseness; Coherences (such as spread and / surrounding coherences); and Distances. Typically, the parameters are given in the time-frequencydomain. Hence, when in the following the terms IVAS and / or MASA are used, it shouldbe understood that they can be replaced with any other suitable codec and / or metadata format and / or system. In the following examples the IVAS stream is described with respect to a decoding and rendering to a binaural output format. However it is appreciated that the decoding and rendering can implement one of any suitable output format, including multichannel, stereo and Ambisonic (FOA / HOA) outputs. In addition, there can be an interface for external rendering, where the output format(s) can correspond, e.g., to the input formats. The current approach for IVAS parametric binaural rendering (used, e.g., with the MASA format) operates in two modes. Afirst mode is to utilize rendering based on head-related transfer functions(HRTFs) without adding a room effect. This rendering corresponds to reproduction in an anechoic environment, and is most suitable, for example, for situations where the IVASaudio signal and spatial metadata already contains all necessary environmental soundelements. This is typically the case when the signal is based on a microphone array recording in a real acoustic space which either already contains a room effect, or for which the room effect is not desired (e.g. outdoor spatial recordings). In a second mode, the HRTF-based rendering additionally uses added reverberation (i.e., the added room effect). Such rendering is needed when the audio signals are intended to be listened in a listening space. For example, stereo or 5.1 sounds may be intended to be listened in a room. The added reverberation also adds to the immersion and externalization at the binaural rendering of such sounds. The operation of the IVAS decoder to render the sound in these two modes has been described for example in in GB application GB1914712.3. The following disclosure is with respect to the second mode, when an IVAS renderer adds reverberation (i.e., the room effect) to the HRTF based rendering. The application of reverberation requires the implementation of at least one reverberator, of which various kinds are described in the literature. The moststraightforward reverberator to implement is a convolution (or an FFT based convolution)to reverberate a signal based on an existing response. However such a kind of reverberator is not parametrically adjustable. For IVAS, in particular, the limiting factor of any convolutional reverberator is that even with optimized FFT-based solutions, the computational load is fairly high, and thus it is not typically used in IVAS for the rendering of the parametric spatial audio. Where the convolutional reverberator aims to reconstruct the original response exactly, another class of reverberators aims to recreate reverberation in a perceptual sense but with significantly lower computational and memory requirements. A typical example of such a reverberator implementation is one which employs delay buffers. Thedelay buffers circulate the audio signal in a processing system, and further employs filters,to create the desired frequency-dependent reverberation times. A typical target of these reverberators is that their response is noise-like, as it is known that for most reverb responses the late reverberation can be modeled with a noise response that matches the original late reverberation in terms of the spectrum and late part reverberation time. However, even if the response is noise-like, the response or the reverberation would not necessarily be time variant, as any such procedure would cause distortion at the signal.A conveniently implemented reverberator implementation in this class ofreverberators is the feedback delay network (FDN) such as disclosed from Jot, J. M., &Chaigne, A. (1991, February). Digital delay networks for designing artificial reverberators. In Audio Engineering Society Convention 90. Audio Engineering Society, which achieveshigh reverberation quality with reasonable reverberation times. There are various otherreverberators in the literature, for example those in the literature review in Gardner, W. G. (1998). Reverberation algorithms. In Applications of digital signal processing to audio and acoustics (pp.85-131). Boston, MA: Springer US. However, the aforementioned reverberators have been designed for time-domain signals. The IVAS parametric renderer implementation typically operates in the time- frequency domain. For that domain, a sparse frequency domain reverberator such as disclosed in Vilkamo, J., Neugebauer, B. and Plogsties, J., 2012. Sparse frequency- domain reverberator. Journal of the Audio Engineering Society, 59(12), pp.936-943, which achieves a comparable quality and complexity to FDN, but is designed to be used in the time-frequency domain. This enables the reverberator to operate in the same domain as the rest of the processing and thus avoid unnecessary time-frequency transforms. Like FDN, this type of reverberator also utilizes delay loops to achieve the infinite response. However, since the operation is in frequency bands, the decay is controlled by simple multiplications instead of filters. The reverberator also uses decorrelators to create the desired inter-channel incoherence and the smoothly decaying short-time temporal characteristics. These decorrelators are computationally efficient since the short-time decay characteristics are implemented by increasing filter sparsity, and the filter tap magnitudes are unity (though can have phases at n*90 degrees). This allows the decorrelators to have no multiplications but only additions (and routing of the real and imaginary parts based on the phases), and thus achieve a high degree of computational efficiency. GB application GB1914712.3 describes the operation of current IVAS parametric binaural renderer implementations that utilize late reverberation. The IVAS implementation uses the reverberator defined above in Vilkamo, J., Neugebauer, B. and Plogsties, J., 2012. Sparse frequency-domain reverberator. Journal of the Audio Engineering Society, 59(12), pp.936-943. It is possible to perform the rendering of the reverberation in its entirety, or the early part of it, with convolution-based methods. The principle of convolutional reverberation is discussed for example in Wefers, F., & Vorländer, M. (2012, March). Optimal filter partitions for non-uniformly partitioned convolution. In Audio Engineering Society Conference: 45th International Conference: Applications of Time-Frequency Processing in Audio. Audio Engineering Society. These methods however are noteffective in IVAS parametric renderer applications. If the entire early part of the roomresponse is rendered with convolution rendering, where the early part is rendered as pointsources (i.e. virtual loudspeakers) at a fixed set of positions. This however can result in ineffective operations for a parametric rendering systemin terms of computational complexity and audio signal quality. In contrast, an efficient and high-quality parametric renderer should perform parametric modifications (e.g., mixing and amplitude and phase modifications) of the transport audio signals, instead of using an intermediate representation of a virtual loudspeaker setup and convolutions. IVAS parametric rendering allows the user to configure the binaural rendering characteristics. As described in GB application GB1914712.3, the renderer is provided with configuration parameters (in frequency bands) determining the reverberation times (T60s) and the energies of the reverberator-based late part rendering, as well as the energies of the (HRTF-based) early-part rendering. The reverberator creates the late response that has noise-like characteristics, so that this noisy response has the desired T60s and spectral characteristics. Having a noise-like response however means that the fine spectral characteristics of the response are also noisy. As is known for any such response for filtering sounds, when the responses are relatively long (e.g., 0.5 seconds or more) the amplitude characteristics of the response vary rapidly across frequency. And conversely, when the noisy response becomes shorter, the spectral features become broader, meaning wider dips and peaks in the frequency response. When there is a large amount of these dips and peaks within a certain frequency range (which happens when the response is long), the response sounds spectrally smooth, as the loudness perception of human hearing has certain frequency limitations (inside so called “critical bands” the human hearing averages the spectrum, i.e., spectral variations inside the critical bands cannot be perceived). With short responses, the peaks and dips are broad with respect to the human hearing characteristics (i.e., the critical bands), which means that the response does not sound spectrally smooth but instead highly colored, in a random fashion due to the noisy characteristics of the response. Moreover, the late reverberation is added to the direct part that is rendered with HRTFs, with a certain delay with respect to it. Such a configuration is similar to comb filtering, again causing unwanted spectral characteristics to the rendered sound. The delay between the early part and the start of the reverberation is fairly short, roughly inthe range of 10 milliseconds, and therefore the comb filter dips are broad and thusaudible. Broad dips of a comb filter are perceived as prominent coloration. Prominent coloration significantly degrades the perceived audio quality. Therefore, the problem with the prior art systems utilizing HRTF-based direct part rendering combined with synthetic late-reverberation is that short reverberation times cause unwanted spectral artefacts (sound coloration and / or comb filtering). These spectral artefacts create significant deterioration of the perceived audio quality. Sound coloration and comb filtering cause similar perception as if the sound is reproduced with very-low-quality headphones or loudspeakers (which cause coloration inthe reproduction), even when good-quality headphones or loudspeakers are used.Listening to the resulting sound is unpleasant and speech intelligibility is decreased. An example of such a short reverberation time causing these problems is 0.1 seconds. The IVAS parametric binaural renderer is an example system that uses such rendering and thus has the aforementioned problems. As such the following examples and embodiments relate to spatial audio rendering (e.g., binaural rendering) from a parametric spatial audio stream (e.g., MASA) with added room effect (i.e., reverberation) with controllable reverberation times, utilizing direct- sound rendering combined with synthetic late-reverberation rendering. In these embodiments there is apparatus and methods that generate good audio quality (i.e., not having spectral coloration nor comb filtering) with any kind of reverberation times (including short reverberation times) with such a direct / early sound + late reverberation rendering system. This is achieved by selectively and intelligently controlling the rendering parameters of the direct-sound and the late-reverberation rendering based on the desired reverberation times, in order to avoid any spectral artefacts that would otherwise be caused due to the short reverberation times. Furthermore these embodiments can comprise apparatus or methods configured to: obtain added-room rendering parameters, containing at least reverberation times and direct (early-sound) and late-reverberation energy parameters (in frequency bands); and determine whether the reverberation times are, at least in part, smaller thandetermined threshold value(s). Additionally these embodiments can be configured to adjust the added-roomrendering parameters (e.g., decreasing the values of the late-reverberation energy parameters and increasing the values of the direct / early-sound energy parameters,and / or increasing the values of the reverberation times) if or when the reverberation timesare smaller than the threshold value(s). Furthermore the embodiments can be configured to obtain a parametric spatialaudio stream, and further render spatial audio (e.g., binaural audio) using the parametricspatial audio stream and the adjusted added-room rendering parameters. The concept as shown by these embodiments as discussed in further detail hereafter is one where the apparatus and methods intelligently controls the added-room rendering parameters by adjusting the parameters so that the spectral artefacts are avoided, but still the perception of correct amount of reverberance is maintained. This is achieved in these embodiments by observing which parameter values cancause spectral artefacts and adjusting them so that the spectral artefacts are avoided, and correspondingly adjust those and / or other parameters in such a way that the potential effects on the perceived reverberance are compensated for. In an example embodiment, when the reverberation times are below the threshold,the control is configured to decrease late-reverberation energy parameter values andincrease the direct (early-sound) energy parameter values, and correspondingly increase the reverberation times (towards the threshold value). This has two positive effects. Firstly, the spectral artefacts are mitigated, as less sound is rendered using the reverberator and the reverberation times are longer (and thus the spectral dips and peaks are denser and thus not perceivable). Secondly, while increasing of the direct energy and decreasing the late reverberation energy decreases the perceived reverberance, increasing the reverberation times increases the perceived reverberance. These two effects somewhat cancel out each other, thus keeping the perceived reverberance as desired. Thus, as a result, the spectral artefacts can be avoided while still maintaining the perception of the desired reverberance. With respect to Figure 1 an example apparatus and system for implementing audiocapture and rendering are shown according to some embodiments.The system 199 is shown with encoder / analyser 101 part and adecoder / synthesizer 105 part. The encoder / analyser 101 part in some embodiments comprises an audio signalsinput configured to receive input audio signals 110. The input audio signals can be fromany suitable source, for example: two or more microphones mounted on a mobile phone; other microphone arrays, e.g., B-format microphone or Eigenmike; Ambisonic signals, e.g., first-order Ambisonics (FOA), higher-order Ambisonics (HOA); Loudspeakersurround mix and / or objects. The input audio signals 110 may be provided to an analysisprocessor 111 and to a transport signal generator 113.The encoder / analyser 101 part may comprise an analysis processor 111. The analysis processor 111 is configured to perform spatial analysis on the input audio signalsyielding suitable metadata 112. The purpose of the analysis processor 111 is thus toestimate spatial metadata in frequency bands. For all of the aforementioned input types, there exists known methods to generate suitable spatial metadata, for example directions and direct-to-total energy ratios (or similar parameters such as diffuseness, i.e., ambient-to-total ratios) in frequency bands. These methods are not detailed herein, however,some examples may comprise the performing of a suitable time-frequency transform forthe input signals, and then in frequency bands when the input is a mobile phonemicrophone array, estimating delay-values between microphone pairs that maximize the inter-microphone correlation, and formulating the corresponding direction value to that delay (as described in GB Patent Application Number 1619573.7 and PCT Patent Application Number PCT / FI2017 / 050778), and formulating a ratio parameter based on the correlation value. The metadata can be of various forms and can contain spatial metadata and other metadata. A typical parameterization for the spatial metadata is one direction parameterin each frequency band ^(^, ^) and an associated direct-to-total energy ratio in eachfrequency band ^(^, ^), where ^ is the frequency band index and ^ is the temporal frameindex. Determining or estimating the directions and the ratios depends on the device or implementation from which the audio signals are obtained. For example the metadatamay be obtained or estimated using spatial audio capture (SPAC) using methodsdescribed in GB Patent Application Number 1619573.7 and PCT Patent ApplicationNumber PCT / FI2017 / 050778. In other words, in this particular context, the spatial audioparameters comprise parameters which aim to characterize the sound-field. In someembodiments the parameters generated may differ from frequency band to frequencyband. Thus for example in band X all of the parameters are generated and transmitted, whereas in band Y only one of the parameters is generated and transmitted, and furthermore in band Z no parameters are generated or transmitted. A practical example of this may be that for some frequency bands such as the highest band some of the parameters are not required for perceptual reasons. When the input is a FOA signal or B-format microphone the analysis processor111 can be configured to determine parameters such as an intensity vector, based onwhich the direction parameter is obtained, and comparing the intensity vector length tothe overall sound field energy estimate to determine the ratio parameter. This method is known in the literature as Directional Audio Coding (DirAC). When the input is HOA signal, the analysis processor 111 may either take the FOAsubset of the signals and use the method above, or divide the HOA signal into multiple sectors, in each of which the method above is utilized. This sector-based method is knownin the literature as higher order DirAC (HO-DirAC). In this case, there is more than onesimultaneous direction parameter per frequency band. When the input is loudspeaker surround mix and / or objects, the analysis processor111 may be configured to convert the signal into a FOA signal(s) (via use of sphericalharmonic encoding gains) and to analyse direction and ratio parameters as above. As such the output of the analysis processor 111 is spatial metadata determinedin frequency bands. The spatial metadata may involve directions and ratios in frequency bands but may also have any of the metadata types listed previously. The spatial metadata can vary over time and over frequency. In some embodiments the spatial analysis may be implemented external to the system 199. For example in some embodiments the spatial metadata associated with the audio signals may be provided to an encoder as a separate bit-stream. In someembodiments the direction parameters within the spatial metadata may be provided as aset of spatial (direction) index values. The encoder / analyser 101 part may comprise a transport signal generator 113. The transport signal generator 113 is configured to receive the input signals and generatea suitable transport audio signal 114. The transport audio signal may be a stereo or monoaudio signal. The generation of transport audio signal 114 can be implemented using aknown method such as summarised below. When the input is mobile phone microphone array audio signals, the transportsignal generator 113 may be configured to select a left-right microphone pair, and applying suitable processing to the signal pair, such as automatic gain control, microphone noise removal, wind noise removal, and equalization. When the input is a FOA / HOA signal or B-format microphone, the transport signalgenerator 113 may be configured to formulate directional beam signals towards left andright directions, such as two opposing cardioid signals. When the input is loudspeaker surround mix and / or objects, the transport signalgenerator 113 may be configured to generate a downmix signal that combines left sidechannels to left downmix channel, and same for right side, and adds centre channels to both transport channels with a suitable gain. In some embodiments the transport signal generator 113 is configured to bypassthe input. For example, in some situations, where the analysis and synthesis occurs atthe same device at a single processing step, without intermediate encoding. The numberof transport channels can also be any suitable number (rather the one or two channelsas discussed in the examples). In some embodiments the encoder / analyser part 101 may comprise anencoder / multiplexer 115. The encoder / multiplexer 115 can be configured to receive thetransport audio signals 114 and the metadata 112. The encoder / multiplexer 115 mayfurthermore be configured to generate an encoded or compressed form of the metadata information and transport audio signals. In some embodiments the encoder / multiplexer115 may further interleave, multiplex to a single data stream 116 or embed the metadatawithin encoded audio signals before transmission or storage. The multiplexing may beimplemented using any suitable scheme. The encoder / multiplexer 115 for example could be implemented as an IVASencoder, or any other suitable encoder. The encoder / multiplexer 115 thus is configuredto encode the audio signals and the metadata and form a bit stream 116 (e.g., an IVAS bit stream). In some embodiments the encoder / analyser part 101 is configured to receive an externally generated spatial audio signal stream, for example a MASA stream, a multichannel loudspeaker signal stream, or an object stream. In these embodiments the encoder / analyser part 101 comprises a suitable encoder, such as an IVAS encoderconfigured to generate the bitstream 116. In other words the encoder could be configuredto receive one of the following two input options: 1) MASA input, in which case the Analysis processor and Transport signal generator are not needed, as MASA stream already contains metadata 112 and transport audio signals 114. 2) Multichannel or object input, in which case the Analysis processor and Transport signal generator are essentially inside the Encoder. In some embodiments any suitable input format can be employed. This bitstream 116 may then be transmitted / stored 103 as shown by the dashed line. In some embodiments there is no encoder / multiplexer 115 (and thus no decoder / demultiplexer 121 as discussed hereafter). The system 199 furthermore may comprise a decoder / synthesizer part 105. Thedecoder / synthesizer part 105 is configured to receive, retrieve or otherwise obtain thebitstream 116, and from the bitstream generate suitable audio signals to be presented tothe listener for example via listener playback apparatus, such as headphones or multi-channel loudspeakers. The decoder / synthesizer part 105 may comprise a decoder / demultiplexer 121configured to receive the bitstream and demultiplex the encoded streams and thendecode the audio signals to obtain the transport signals 124 and metadata 122. Furthermore in some embodiments, as discussed above there may not be any demultiplexer / decoder 121 (for example where there is no associated encoder / multiplexer 115 as both the encoder / analyser part 101 and the decoder / synthesizer 105 are located within the same device). The decoder / synthesizer part 105 can comprise an added room effect parameter determiner 127. The added room effect parameter determiner 127 is configured to obtainadded room effect information 126. The added room effect information 126 can, forexample, comprise desired reverberation times and early and late part energy spectral coefficients (in frequency bands). In some embodiments the added room effect information 126 can comprise binaural room impulse responses (BRIRs). The addedroom effect parameter determiner 127 can then generate, based on the added room effectinformation 126 suitable added room effect parameters 128. In some embodiments, theadded room effect information 126 can already comprise the added room effectparameters 128. In these embodiments, the added room effect parameter determiner 127can operate a pass through mode and can pass through the added room effectparameters (in some embodiments with no processing of the parameters). In someembodiments, where the added room effect information 126 comprises information suchas BRIRs the added room effect parameter determiner 127 can be configured todetermine added room effect parameters using the methods described in GB applicationGB1914716.4. In some embodiments, the added room effect information 126 comprisesparameters, in a different format than the defined added room effect parameters 128, inwhich case the added room effect parameter determiner 127 is configured to perform asuitable mapping from the input to output format in order to generate parameters in thecorrect format for outputting.In some embodiments, the added room effect parameter determiner 127 isexternal to the decoder 105. For example the added room effect parameter determiner127 is implemented elsewhere and the added room effect parameters 128 are thenforwarded to the decoder 105.In this example embodiment, the added room effect parameters 128 comprise thefollowing parameters: reverberation times ^′^^(^);early part energy correction coefficients ^′^^^^^(^); andlate part energy correction coefficients ^′^^^^(^), where ^ denotes the bin index.The late part energy correction coefficients may also be called (late) reverberation energycorrection coefficients. In the present example, the energy correction coefficients ^′^^^^^(^) and ^′^^^^(^)indicate the energy that the early and late part responses should have in average infrequency bin ^, regardless of the ^′^^(^) values.For example, if ^′^^^^(^) = 1 then in that frequency the late reverb response hasunity energy, no matter which ^′^^(^) value is used. Similarly, the parameter ^′^^^^^(^) isintended to control the mean spectrum of the direct part rendering. If the direct partrendering is based on diffuse field equalized HRTFs, then ^′^^^^^(^) is intended to beapplied to modify the mean spectrum. The HRTFs may have different spectrum atdifferent directions, but the average spectrum of all directions is then ^′^^^^^(^). These added room effect parameters 128 can then be passed to an added room effect parameter controller 129. The decoder / synthesizer part 105 can comprise an added room effect parametercontroller 129. The added room effect parameter controller 129 is configured to receivethe added room effect parameters 128, such as the reverberation times ^′^^(^), early partenergy correction coefficients ^′^^^^^(^), and late part correction coefficients ^′^^^^(^)and modifies them to form adjusted added room effect parameters 130.The adjusted added room effect parameters 130 can thus comprise:adjusted reverberation times ^^^(^); adjusted early part energy correction coefficients ^^^^^^(^);adjusted late part energy correction coefficients ^^^^^(^).As discussed herein as there are no limits, in current implementations, regarding the energy correcting gain coefficients and reverberation times (apart from not havingnegative or infinite values), there are parameter configurations which provide undesiredaudio quality outputs and which the embodiments featuring the added room effectparameter controller 129 attempt to prevent occurring. In some embodiments the operations of added room effect parameter determiner127 and added room effect parameter controller 129 are typically implemented eitherduring an initialization of the decoder, or before the initialization occurs. The decoder / synthesizer part 105 may comprise a synthesisprocessor / synthesizer 123. The synthesis processor / synthesizer 123 is configured toobtain the transport audio signals 124, the spatial metadata 122 and an adjusted addedroom effect parameter 130 and produces a spatial audio output 128 (for example abinaural audio signal) that can be reproduced over headphones or any suitable outputdevice (for example multichannel loudspeakers). The operations of this system are summarized with respect to the flow diagram as shown in Fig.2. Fig.2 shows for example the receiving of the input audio signals as shown in step 201. Then the flow diagram shows the analysis (spatial) of the input audio signals togenerate the spatial metadata as shown in Fig.2 by step 203.The transport audio signals are then generated from the input audio signals as shown in Fig.2 by step 204. The generated transport audio signals and the metadata may then be encodedand multiplexed as shown in Fig.2 by step 205. The encoded signals can furthermore be demultiplexed and decoded to generatetransport audio signals and spatial metadata as shown in Fig.2 by step 207. This is alsoshown as an optional dashed box. Then spatial audio output (binaural audio signals) can be synthesized based onthe transport audio signals, spatial metadata and an added room effect information asshown in Fig.2 by step 209. This step can furthermore be divided into an operations of: obtaining the added room effect information; generating added room effect parameters based on the added room effect information; generating adjusted added room effect parameters based on the added room effect parameters; obtaining the transport signals and metadata; synthesizing the spatial audio output based on the transport audio signals, spatial metadata and the adjusted added room effect parameters. The synthesized spatial output (binaural audio signals) may then be output to asuitable output device, for example a set of headphones (or loudspeakers), as shown inFig.2 by step 211. With respect to Fig.3 is shown the synthesis processor 123 in further detail. In some embodiments the synthesis processor 123 comprises a time-frequency transformer 301. The time-frequency transformer 301 is configured to receive the (time-domain) transport audio signals 122, which converts them to the time-frequency domain.Suitable transforms include, e.g., short-time Fourier transform (STFT) and complex-modulated quadrature mirror filterbank (QMF) or a low delay variant of it, such as thecomplex low-delay filterbank (CLDFB). The resulting signals may be denoted as ^^(^, ^),where ^ is the channel index, ^ the frequency bin index of the time-frequency transform,and ^ the time index. The time-frequency signals are for example expressed here in avector form (for example for two channels the vector form is): In some embodiments the CLDFB is used and there are 60 frequency bins ^. The following processing operations may then be implemented within the time-frequencydomain and over frequency bands. A frequency band can be one or more frequency bins(individual frequency components) of the applied time-frequency transformer (filter bank). The frequency bands could in some embodiments approximate a perceptually relevant resolution such as the Bark frequency bands, which are spectrally more selective at low frequencies than at the high frequencies. Alternatively, in some implementations, frequency bands can correspond to the frequency bins. The frequency bands are typically those (or approximate those) where the spatial metadata has been determined by theanalysis processor. Each frequency band k may be defined in terms of a lowest frequencybin ^^^^(^) and a highest frequency bin ^^^^^(^). In the present example the processingis, however, performed in each frequency bin ^ independently.The index ^ denotes the transformed domain temporal index (sometimes referredto as “slot”). The time-frequency transport signals 302 in some embodiments may be providedto a covariance matrix estimator 307, a reverberator 351 and to a mixer 311.The synthesis processor 123 in some embodiments comprises a covariance matrixestimator 307. The covariance matrix estimator 307 is configured to receive the time-frequency domain transport signals 302 and estimates a covariance matrix of the time-frequency transport signals and their overall energy estimate (in frequency bands). Thecovariance matrix can for example in some embodiments be estimated as: where superscript H denotes the conjugate transpose. The estimated covariancematrix 310 may be output to a mixing rule determiner 309. The indices ^^(^) and ^^(^)are the first and last temporal indices of temporal subframe ^. The spatial metadata in thepresent example is defined in subframes ^. For example, one subframe ^ may have fourtemporal indices (slots) ^. The covariance matrix estimator 307 may also be configured to generate an overallenergy estimate ^(^, ^) 308, that is the sum of the diagonal values of ^^(^, ^) , andprovides this overall energy estimate to a target covariance matrix determiner 305. In some embodiments the synthesis processor 123 comprises a HRTF determiner303. The HRTF determiner 303 may comprise a suitably dense set of HRTFs or a HRTFinterpolator. The HRTF determiner is configured to determine a 2x1 complex-valuedhead-related transfer function (HRTF) ^(^(^, ^), ^) for an angle ^(^, ^) and frequency bin^. In some embodiments the HRTF determiner 303 is configured to receive the spatialmetadata 124 and from the angle ^(^, ^) (which is the direction parameter at the spatialmetadata) determine the output HRTF. Where the listener head-orientation tracking is involved, the directionparameters ^(^, ^) can be modified prior to obtaining the HRTFs to account for the currenthead orientation. The HRTF data set of the HRTF determiner 303 can in someembodiments be pre-formulated and fixed for the synthesis processor 123, and there canbe multiple HRTF data sets to select from.The HRTF data set of the HRTF determiner 303 in some embodiments also has adiffuse-field covariance matrix for each bin ^, which may be formulated for example bytaking an equally distributed set of directions ^^ where d = 1..D, and by estimating thediffuse-field covariance matrix as . The HRTF data may be rendered and interpolated by using any suitable method.For example, in some embodiments, a set of HRTFs is decomposed into inter-aural time differences and energies of left and right ears as a function of frequency. Then, when a HRTF at a given angle is needed, then the nearest existing data points at the HRTF set are found and the delays and energies at the given angle are interpolated. These energies and delays can be then converted as complex multipliers to be used. In some embodiments HRTFs are interpolated to convert a HRTF data set into aset of spherical harmonic binaural decoding matrices in frequency bands. Then, the HRTFfor any angle can be determined by formulating a spherical harmonic weight vector for that angle and multiplying it with that matrix. The result is again the 2x1 HRTF vector. In some embodiments interpolation of HRTFs can be implemented by treating them as virtual loudspeakers and to obtain the interpolated HRTFs, e.g., via amplitude panning. A HRTF, by definition, refers to a response from a certain direction to the ears in an anechoic space. However, it is entirely possible to use in place of a HRTF data set another data set which includes (additionally to the HRTF part) also early part of a binaural room impulse response. Such a data set also includes spectral and other features that are for example due to first floor or wall reflections. The HRTF data 304 (which consists of ^(^(^, ^), ^) and ^^(^)) can be output bythe HRTF determiner 303 and passed to a target covariance matrix determiner 305. In some embodiments the synthesis processor 123 comprises a target covariancematrix determiner 305. The target covariance matrix determiner 305 is configured toreceive the spatial metadata 124 which can in this example comprise at least onedirection parameter ^(^, ^) and at least one direct-to-total energy ratio parameter ^(^, ^),the HRTF data 304, the adjusted added room effect parameters 130 for example theadjusted early part energy correction coefficients ^^^^^^(^) , and the overall energyestimate ^(^, ^) 308.In some embodiments the target covariance matrix determiner 305 is configuredto modify the overall energy estimate according to the following expression The target covariance matrix determiner 305 is then configured to determine a target covariance matrix 306 based on the spatial metadata 124, the HRTF data 304 and the modified overall energy estimate. For example the target covariance matrix determiner 305 may formulate the targetcovariance matrix by The target covariance matrix 306 can then be provided to the mixing ruledeterminer 309. The synthesis processor 123 in some embodiments comprises a mixing ruledeterminer 309. The mixing rule determiner 309 is configured to receive the targetcovariance matrix 306 and the measured or estimated covariance matrix ^^(^, ^)310. The mixing rule determiner 309 is configured to generate a mixing matrix ^(^, ^) 312based on the target covariance matrix 306 and the estimated or measuredcovariance matrix ^^(^, ^) 310.In some embodiments the mixing matrix is generated based on a method described in “Optimized covariance domain framework for time–frequency processing ofspatial audio”, J Vilkamo, T Bäckström, A Kuntz, Journal of the Audio Engineering Society61, no.6 (2013): 403-411. In some embodiments the mixing rule determiner 309 is configured to determine aprototype matrix ^ = that guides the generation of the mixing matrix.In summary a mixing matrix ^(^, ^) may be provided that when applied to a signalwith a covariance matrix ^^(^, ^) it produces a signal with covariance matrix , ina least-squares optimized way. Matrix ^ guides the signal content in such mixing, and inthis example that matrix is (effectively) the identity matrix, since the left and right processed signals should resemble the original left and right signals. In other words, thedesign is to minimally alter the signals while obtaining for the processed output.The mixing matrix ^(^, ^) is formulated for each frequency bin b and is provided to themixer 311. In this example the mixing matrix is defined based on the input being a two channeltransport audio signal. However these methods can be adapted to embodiments for anynumber of transport audio channels. In some embodiments, the mixing rule determiner is further configured to formulatea residual mixing matrix ^^(^, ^) , the determination of which is also described in“Optimized covariance domain framework for time–frequency processing of spatial audio”, J Vilkamo, T Bäckström, A Kuntz, Journal of the Audio Engineering Society 61, no. 6 (2013): 403-411. In summary, the mixing method involves regularizations at the mixing matrix formulation, and as consequence the resulting mixing solution is such thatthe target covariance matrix ^^(^, ^) is not in all situations fully reachable with using themixing operation ^(^, ^) only. The residual mixing matrix is then determined such that itis for processing a decorrelated version of the transport audio signals. Adding this to the overall output then compensates for the effects of the necessary regularizations. Themixing matrices ^(^, ^) and ^^(^, ^) are then provided to the mixer 311.The synthesis processor 123 in some embodiments comprises a mixer 311. Themixer 311 receives the time-frequency transport audio signals 302 and the mixingmatrices 312. The mixer 311 is configured to process the time-frequency audio signals(input signal) in each frequency bin b to generate two processed (first or early part) time-frequency signals 314. This may, for example be formed based on the followingexpression: where ^[^(^, ^)] denotes a decorrelated version of ^(^, ^).In the above formula the time index for mixing matrices is the subframe ^ and forsignals it is the sample (or slot) ^. Although not shown in the equation, in the processingaccording to it, the mixing matrices may be temporally interpolated from ^(^, ^ − 1) to^(^, ^) during the four temporal steps ^. The interpolation may be linear, or adaptive sothat onsets cause faster interpolation. Also the decorrelated version of x(b, t) (i.e.,D[x(b,t)]) can have been obtained by applying decorrelation to x(b, t). For example thiscan be obtained within the mixer.The output of the mixer 311, for example the processed binaural (early part) time-frequency signal ^(^, ^) 314, is then passed to the combiner 313.The above part accounts for the early / dry part of the binaural processing, and thefollowing accounts for the late / wet part of the binaural processing. In this example, there is employed a reverberator 351 configured to receive the output of the T / F transformer301 to generate a late reverberation T / F signal or reverberation (late) part signal 318. Inthis example, the reverberator 351 employs the frequency-domain reverberator asdescribed earlier to reverberate the time-frequency transport audio signals 302.The reverberator 351 receives the Time-frequency transport audio signals 302 andadjusted added room effect parameters 130. The operation of the reverberator is onewhere, the reverberator operates in frequency bins, and each bin has a delay loop with a loop gain factor. The delay line length and the gain factor are adjusted based on ^^^(^), and this loop structure creates, as a response, an infinite decaying set of impulses. The delay line output is then processed with decorrelators which have energy-decaying responses according to ^^^(^). The purpose of the decorrelators is to create the short- time dense response for the reverberator and the desired incoherence between the channels. The decorrelator is implemented as a set of unity gains with potential phase shifts (0, 90, 180 or 270 degrees), where increasing sparsity of the response models the energy decay over time. This structure is computationally efficient and has no multiplications apart from the loop gain. The two mutually incoherent decorrelator output channels are then cross mixed in frequency bins to approximate the binaural diffuse field coherence. The reverberator is initialized based on the adjusted reverberation times ^^^(^)and adjusted late part energy correction gains ^^^^^(^, ^). As mentioned in the foregoing,the reverberator is configured so that the reverb tail approximates ^^^(^), and so that itsresponse has the energy determined in ^^^^^ (^, ^). The output of the reverberator is areverberation time-frequency signal 318 provided to the combiner 313.The synthesis processor 123 in some embodiments comprises a combiner 313 which is configured to receive the binaural T / F early part signal 314 and the reverberation(late) part signal 318 and combine or sum these together (for the left and right channelsseparately) in order to generate the binaural T / F signal 320. The binaural T / F signal 320can then be passed to an inverse T / F transformer 315. In some embodiments the synthesis processor 123 comprises an inverse T / Ftransformer 315 configured to receive the binaural T / F signal 320 and apply an inverse time-frequency transform corresponding to the applied time-frequency transform appliedby the T / F transformer 301. The output of the inverse T / F transformer 315 is a spatialaudio output (binaural audio signal) 128 suitable for output (for example overheadphones). With respect to Fig.4 a flow diagram showing the operation of the synthesis processor is shown. The flow diagram shows the operation of receiving such as the transport audiosignals, spatial metadata, and adjusted added room effect parameters as shown in Fig.4by step 401. Furthermore the HRTF data is determined as shown in Fig.4 by step 402. The generation of the time-frequency domain transport audio signals is shown inFig.4 by step 403. The generation of a reverberation (late) part signal 318 based on T / F transportaudio signals and adjusted added room effect parameters is shown in Fig.4 by step 405.The estimation of the covariance matrix based on the T / F transport audio signals and the overall energy based on the covariance matrix is shown in Fig.4 by step 407. The determination of the target covariance matrix based on HRTF data, spatialmetadata, adjusted added room effect parameters, and energy estimates is shown inFig.4 by step 409. Having determined the target covariance matrix and the estimated covariancematrix then the mixing rule is determined based on the estimated covariance matrix andtarget covariance matrix as shown in Fig.4 by step 411.The time-frequency transport signals can then be mixed based on mixing rule asshown in Fig.4 by step 413. The mixed audio signals (the Binaural T / F early part signal) and reverberation(late) part signal 318 are then combined to generate a binaural T / F signal as shown inFig.4 by step 415. Then the binaural T / F signal can be converted back to the time domain, or timedomain spatial audio output (binaural audio signals) are generated as shown in Fig.4 bystep 417. The spatial audio output (binaural audio signals) may then be output as shown inFig.4 by step 419. With respect to Fig.5 is shown the added room effect parameter controller 129 infurther detail. The added room effect parameter controller 129 as described earlier isconfigured to receive the added room effect parameters 128, such as the reverberationtimes ^′^^(^) 500, early part energy correction coefficients ^′^^^^^(^) 502, and late partenergy correction coefficients ^′^^^^(^) 504 and modifies them to form adjusted addedroom effect parameters 130 such as adjusted reverberation times ^^^(^) 550, adjustedearly part energy correction coefficients ^^^^^^(^) 552, and adjusted late part energycorrection coefficients ^^^^^(^) 554. These can also be known as adjusted early partenergy parameters, adjusted early part energy coefficients, adjusted late part energyparameters and adjusted late part energy coefficients. As described earlier theadjustments attempt to avoid undesired audio quality by avoiding parameter values whichare problematic because the late reverberator response is ultimately noisy (but time- invariant), and short reverberation times cause sound coloration issues. Therefore, theadded room effect parameter controller 129 attempts to modify the parameters, in orderto avoid the coloration issues. In some embodiments the added room effect parameter controller 129 comprisesa reverberation time comparer 501. The reverberation time comparer 501 is configuredto receive the reverberation times ^′^^(^) 500 and a T60 threshold parameter ^^^,^^^^^^^^^580 and then for each bin, compare the reverberation times ^′^^(^) 500 against the T60threshold parameter ^^^,^^^^^^^^^ 580 to determine a T60 modifier ^^^^(^) 592 value. Insome embodiments this T60 modifier ^^^^(^) 592 value is determined using the followingexpression where ^^^^^^,^is an operator truncating the values between 0 and 1. In someembodiments the value of the T60 threshold parameter ^^^,^^^^^^^^^ = 0.2 however othervalues could be used. In some embodiments the added room effect parameter controller 129 comprisesa reverberation time adjuster 503. The reverberation time adjuster 503 is configured toreceive the reverberation times ^′^^(^) 500 and T60 modifier ^^^^(^) 592 and from thisdetermine the adjusted reverberation times ^^^(^) 550. For example the reverberationtimes can be adjusted by the expression^^^(^) = ^1 − ^^^^(^)^^′^^(^) + ^^^^(^)(^′^^(^) + ^^^,^^^^^^^^^) / 2and output the adjusted reverberation times ^^^(^) 550.In some embodiments the added room effect parameter controller 129 comprises an energy modifier generator 505. The energy modifier generator 505 is configured toreceive the reverberation times ^′^^(^) 500 and adjusted reverberation times ^^^(^) 550and based on these generate an energy modifier value ^^^^(^) 594. The energy modifiervalue ^^^^(^) 594 can, for example, be formulated by In some embodiments the added room effect parameter controller 129 comprises an early energy corrector 507. The early energy corrector 507 is configured to receive theenergy modifier value ^^^^(^) 594, the early part energy correction coefficients ^′^^^^^(^)502 and late part energy correction coefficients ^′^^^^(^) 504 to generate the adjustedearly part energy correction coefficients ^^^^^^(^) 552. For example the adjusted earlypart energy correction coefficients ^^^^^^(^) 552 can be generated based on the followingformula The adjusted early part energy correction coefficients ^^^^^^(^) 552 can be thenoutput. In some embodiments the added room effect parameter controller 129 comprisesa late energy corrector 509. The late energy corrector 509 is configured to receive theenergy modifier value ^^^^(^) 594 and the late part energy correction coefficients^′^^^^(^) 504 to generate the adjusted late part energy correction coefficients ^^^^^(^)554. For example the adjusted late part energy correction coefficients ^^^^^ (^) 554 canbe generated based on the following formula The adjusted late part energy correction coefficients ^^^^^(^) 554 can be thenoutput.The added room effect parameter controller 129 then outputs the adjusted addedroom effect parameters. With respect to Fig.6 is shown a flow diagram showing the operations of the addedroom effect parameter controller 129 according to some embodiments.For example, there is shown in Fig.6 by 601 an operation of receiving or obtainingthe reverberation times ^′^^(^) 500, early part energy correction coefficients ^′^^^^^(^)502, and late part energy correction coefficients ^′^^^^(^) 504.Additionally as shown in Fig.6 by 603 there is an operation of comparing the reverberation time against a threshold value to generate a reverberation time modification factor. Then is shown an operation of generating the adjusted reverberation time ^^^(^)based on the reverberation time modification factor and reverberation time as shown in Fig.6 by 605. After this is an operation of generating an energy modifier factor based on theadjusted reverberation time and reverberation time as shown in Fig.6 by 607.Then are the operation of generating adjusted early and late part energy correction coefficients based on the energy modifier factor and also the original early and late energycorrection coefficients as shown in Fig.6 by 609.Finally is the operation of outputting the adjusted parameters as shown in Fig.6 by 611. The parameter modification according to the above example formulas is illustrated in Figs.7a and 7b. For example Fig.7a shows the graph of the adjustment of the T60 values and thedeviation below the 0.2 value (the threshold value selected in this example) to increase the T60 value when it is lower than the threshold values. Fig.7b shows the graph of the energy modifier values against T60 values showing the modifier value going from 1 to 0 from T60 values of 0 to the threshold value of 0.2.The energy modifier value m^^^(b) is illustrated, which moves more energy from latereverberation to direct sound rendering when m^^^(b) has higher values.It should be noted that the presented parameter modification formulas shown above are an example of modification. In other embodiments, other kind of modificationsor other suitable formulas can be implemented. A simulation was generated in order to examine the effect of the modification of the parameters proposed above. In this simulation a decoder with and without the addedroom effect parameter controller was generated. The testing was performed with speechand pink noise inputs, using a synthetic audio scene with the sound source directly at left. Speech inputs were used to validate, by listening, that the room effect parameter modification, as described in the foregoing, provided the desired advantages at different T60 ranges. The pink noise input was used to generate the illustration presented in the following. In Fig.8a and 8b, two plots are shown. The top plot Fig.8a shows the spectrum ofthe left ear binaural signal when the added room effect parameter controller 129 is in abypass mode and no modification are applied. The bottom plot Fig.8b shows the spectrumof the left ear binaural signal when the added room effect parameter controller 129operates as described in the foregoing. To provide a clearly visualizable example, the reverberation time in this example was 0.02 at all frequencies, and the direct and reverberant parts were adjusted to have the same overall energy. The results show that,as shown in Fig.8a without the added room effect parameter controller 129, there is astark comb filter effect present in the spectrum, causing highly colored characteristics to the sound. In the Fig.8b (bottom figure), where the added room effect parameter controller129 is operating, the comb filter artefact is not present or is suppressed, and the soundis not perceived to be colored. While the example shown in Figs.8a and 8b provides a clear spectral visualizationof the improvement, informal listening experiments were performed to evaluate theperformance of the reverberator with speech signals. It was found that there are clear audible benefits perceived with speech, also for other T60 values between 0 and 0.2. With T60 values larger than 0.2, the output with andwithout the added room effect parameter controller 129 was identical, as designed (withrespect to the threshold value being chosen or defined by 0.2). In some embodiments the transport audio signal and the parametric spatialmetadata format could be in some other formats than the MASA format without requiringsignificant redesign in the decoder and synthesis processor. Apparatus suitable for implementing some embodiments with respect to thedecoding and rendering a parametric audio stream (comprising metadata and audiosignals (e.g., in MASA format)) is shown with respect to Fig.9. The audio stream could, e.g., be an encoded IVAS audio stream. The example apparatus could be for example a mobile phone, a laptop, or a teleconferencing system. As shown in Fig.9, the example apparatus comprises a transceiver 903 which isconfigured to receive the bitstream 902 / 116 (e.g., an IVAS bitstream) and provides it to aprocessor 915. Typically, the transceiver 903 is configured to receive the bitstreamwirelessly from a remote device or a server. In some embodiments the bitstream is received via a wired connection or read from a local memory of the device. The example apparatus comprises an user interface 905 which may display to theuser an interface allowing the user to determine the desired reverberation characteristics (for example reverberation information), e.g., by selecting from pre-determined settings the desired one. Alternatively, in some embodiments the user could even hand-tune thereverberation times and energies. In this example this reverberation information is theadded room effect information provided to the processor 915.The apparatus further comprises a processor 915 coupled with a memory 917 andexecutes program code 919 that resides in the memory 917. The program code 919 mayinvolve instructions to perform the operations described above. The processor 915coupled with the memory 917 executing program code 919 can, e.g., be considered to bean IVAS decoder 907, implementing the embodiments described herein. The processor915 then outputs the spatial audio output 906 / 128, which in this example was a binauraloutput, to the DAC or Bluetooth 909.The DAC or Bluetooth 909 is a digital-to-analogue converter which converts thespatial audio output 906 / 128 to an analogue form if the headphones are conventionalwired (analogue) headphones 911. For wireless connections, the DAC or Bluetooth 909may be a Bluetooth transceiver. The DAC or Bluetooth 909 block provides (either wired or wirelessly) the spatialaudio 908 / 128 to be played back with the headphones 911 to the user. In someembodiments, the headphones 911 may comprise a head tracker which may provideorientation and / or position information of the user’s head to the processor 915 of the rendering apparatus, so that user’s head orientation is accounted for at the spatialsynthesizer as described above.The remote device (not shown in Fig.9) may be configured to generate thebitstream in various ways. In one example, the MASA stream is captured using a device with a microphone array (e.g., a mobile phone). Any other suitable capture method may be used as well. For the present invention, the bitstream may originate from any kind ofa setting. The apparatus of Fig.9 may also capture the audio locally, and transmit it to aremote device, where the remote device may perform the rendering similarly to the device of Fig.9. With respect to Fig.10 an example electronic device which may be used as any ofthe apparatus parts of the system as described above. The device may be any suitableelectronics device or apparatus. For example in some embodiments the device 1700 is amobile device, user equipment, tablet computer, computer, audio playback apparatus,etc. The device may for example be configured to implement the encoder / analyser part101 or the decoder / synthesizer part 105 as shown in Fig.1 or any functional block as described above. In some embodiments the device 1700 comprises at least one processor or centralprocessing unit 1707. The processor 1707 can be configured to execute various programcodes such as the methods such as described herein.In some embodiments the device 1700 comprises a memory 1711. In someembodiments the at least one processor 1707 is coupled to the memory 1711. Thememory 1711 can be any suitable storage means. In some embodiments the memory 1711 comprises a program code section for storing program codes implementable upon the processor 1707. Furthermore in some embodiments the memory 1711 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1707 whenever needed via the memory-processor coupling. In some embodiments the device 1700 comprises a user interface 1705. The userinterface 1705 can be coupled in some embodiments to the processor 1707. In some embodiments the processor 1707 can control the operation of the user interface 1705 and receive inputs from the user interface 1705. In some embodiments the user interface 1705 can enable a user to input commands to the device 1700, for example via a keypad. In some embodiments the user interface 1705 can enable the user to obtain informationfrom the device 1700. For example the user interface 1705 may comprise a displayconfigured to display information from the device 1700 to the user. The user interface1705 can in some embodiments comprise a touch screen or touch interface capable ofboth enabling information to be entered to the device 1700 and further displayinginformation to the user of the device 1700. In some embodiments the user interface 1705may be the user interface for communicating. In some embodiments the device 1700 comprises an input / output port 1709. Theinput / output port 1709 in some embodiments comprises a transceiver. The transceiver insuch embodiments can be coupled to the processor 1707 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling. The transceiver can communicate with further apparatus by any suitable known communications protocol. For example in some embodiments the transceiver can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as for example IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or infrared data communication pathway (IRDA). The transceiver input / output port 1709 may be configured to receive the signals. In some embodiments the device 1700 may be employed as at least part of thesynthesis device. The input / output port 1709 may be coupled to headphones (which maybe a headtracked or a non-tracked headphones) or similar.It is also noted herein that while the above describes example embodiments, there are several variations and modifications which may be made to the disclosed solution without departing from the scope of the present invention. The following abbreviations have been used in the description: CLDFB - complex-valued low-delay filterbankIVAS - immersive voice and audio servicesMASA - metadata-assisted spatial audioSTFT - short-time Fourier transformIn general, the various embodiments, the apparatus described above may be implemented in hardware or special purpose circuitry, software, logic or any combination thereof. For example the apparatus can be implemented as a chipset. Some aspects of the disclosure may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the disclosure is not limited thereto. While various aspects of the disclosure may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof. As used in this application, the term “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device. The embodiments of this disclosure may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Computer software or program, also called program product, including software routines, applets and / or macros, may be stored in any apparatus-readable data storage medium and they comprise program instructions to perform particular tasks. A computer program product may comprise one or more computer-executable components which, when the program is run, are configured to carry out embodiments. The one or more computer-executable components may be at least one software code or portions of it. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD. The physical media is a non-transitory media. The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may comprise one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), FPGA, gate level circuits and processors based on multi core processor architecture, as non-limiting examples. Embodiments of the disclosure may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate. The scope of protection sought for various embodiments of the disclosure is set out by the independent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the disclosure. The foregoing description has provided by way of non-limiting examples a full and informative description of the exemplary embodiment of this disclosure. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this disclosure will still fall within the scope of this invention as defined in the appended claims. Indeed, there is a further embodiment comprising a combination of one or more embodiments with any of the other embodiments previously discussed.

Claims

CLAIMS:

1. An apparatus for rendering spatial audio with reverberation, the apparatuscomprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: receive a spatial audio stream, the spatial audio stream comprising at least one audio signal and spatial metadata associated with the at least one audio signal; obtain at least one room effect parameter comprising at least one reverberation time; adjust at least one of the at least one room effect parameter when the at least one reverberation time is less than a determined threshold value; and generate a spatial audio signal using the spatial audio stream and the adjusted at least one room effect parameter, the generated spatial audio signal comprising: a first spatial audio portion generated based on the at least one audio signal and processed dependent on the spatial metadata; and a second spatial audio portion generated based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter.

2. The apparatus as claimed in claim 1, caused to generate the spatial audio signalusing the spatial audio stream and the adjusted at least one room effect parameter is caused to: generate the first spatial audio portion based on the at least one audio signal and processed dependent on the spatial metadata; separately generate the second spatial audio portion based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter; and combine the first spatial audio portion and the second spatial audio portion to generate the spatial audio signal.

3. The apparatus as claimed in any of claim 1 or 2, wherein the first spatial audioportion is at least one of:a direct audio signal portion of the spatial audio signal; andan early reverberation signal portion of the spatial audio signal.

4. The apparatus as claimed in any of claims 1 to 3, wherein the first spatial audioportion is further processed dependent on the adjusted at least one room effect parameter.

5. The apparatus as claimed in any of claims 1 to 4, wherein the second spatial audioportion is a late reverberation portion of the spatial audio signal.

6. The apparatus as claimed in any of claims 1 to 5, wherein the second spatial audioportion is further processed dependent on the spatial metadata.

7. The apparatus as claimed in any of claims 1 to 6, wherein the adjusted at leastone room effect parameter comprises at least one of: at least one adjusted early part energy correction coefficient; at least one adjusted late part energy correction coefficient; at least one adjusted early part energy parameter; at least one adjusted late part energy parameter; at least one adjusted early part energy coefficient; at least one adjusted late part energy coefficient; and at least one adjusted reverberation time.

8. The apparatus as claimed in any of claims 1 to 7, caused to adjust at least one ofthe at least one room effect parameter when the at least one reverberation time is less than a determined threshold value is further caused to: generate a reverberation time modification parameter based on the reverberationtime and the determined threshold value; andgenerate an at least one adjusted reverberation time by processing the at least one reverberation time with the reverberation time modification parameter.

9. The apparatus as claimed in claim 8, caused to adjust at least one of the at leastone room effect parameter when the at least one reverberation time is less than a determined threshold value is further caused to: generate an energy modification value based on the at least one adjustedreverberation time and the at least one reverberation time; andgenerate at least one adjusted early part energy correction coefficient based on anearly part energy correction coefficient and the energy modification value.

10. The apparatus as claimed in claim 9, caused to adjust at least one of the at leastone room effect parameter when the at least one reverberation time is less than a determined threshold value is further caused to generate at least one adjusted late part energy correction coefficient based on a late part energy correction coefficient and the energy modification value.

11. A method for rendering spatial audio with reverberation, the method comprising:receiving a spatial audio stream, the spatial audio stream comprising at least one audio signal and spatial metadata associated with the at least one audio signal; obtaining at least one room effect parameter comprising at least one reverberation time; adjusting at least one of the at least one room effect parameter when the at least one reverberation time is less than a determined threshold value; and generating a spatial audio signal using the spatial audio stream and the adjusted at least one room effect parameter, the generated spatial audio signal comprising: a first spatial audio portion generated based on the at least one audio signal and processed dependent on the spatial metadata; and a second spatial audio portion generated based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter.

12. The method as claimed in claim 11, wherein generating the spatial audio signalusing the spatial audio stream and the adjusted at least one room effect parameter further comprises:generating the first spatial audio portion based on the at least one audio signal and processed dependent on the spatial metadata; separately generating the second spatial audio portion based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter; and combining the first spatial audio portion and the second spatial audio portion to generate the spatial audio signal.

13. The method as claimed in any of claim 11 or 12, wherein the first spatial audioportion is at least one of: adirect audio signal portion of the spatial audio signal; andan early reverberation signal portion of the spatial audio signal.

14. The method as claimed in any of claims 11 to 13, wherein the first spatial audioportion is further processed dependent on the adjusted at least one room effect parameter.

15. The method as claimed in any of claims 11 to 14, wherein the second spatial audioportion is a late reverberation portion of the spatial audio signal.

16. The method as claimed in any of claims 11 to 15, wherein the second spatial audioportion is further processed dependent on the spatial metadata.

17. The method as claimed in any of claims 11 to 16, wherein the adjusted at leastone room effect parameter comprises at least one of: at least one adjusted early part energy correction coefficient; at least one adjusted late part energy correction coefficient; at least one adjusted early part energy parameter; at least one adjusted late part energy parameter; at least one adjusted early part energy coefficient; at least one adjusted late part energy coefficient; and at least one adjusted reverberation time.

18. The method as claimed in any of claims 11 to 17, wherein adjusting at least one ofthe at least one room effect parameter when the at least one reverberation time is less than a determined threshold value further comprises: generating a reverberation time modification parameter based on the reverberationtime and the determined threshold value; andgenerating an at least one adjusted reverberation time by processing the at least one reverberation time with the reverberation time modification parameter.

19. The method as claimed in claim 18, wherein adjusting at least one of the at leastone room effect parameter when the at least one reverberation time is less than a determined threshold value further comprises: generating an energy modification value based on the at least one adjustedreverberation time and the at least one reverberation time; andgenerating at least one adjusted early part energy correction coefficient based onan early part energy correction coefficient and the energy modification value.

20. The method as claimed in claim 19, wherein adjusting at least one of the at leastone room effect parameter when the at least one reverberation time is less than a determined threshold value further comprises generating at least one adjusted late part energy correction coefficient based on a late part energy correction coefficient and the energy modification value.

21. An apparatus comprising means configured to:receive a spatial audio stream, the spatial audio stream comprising at least one audio signal and spatial metadata associated with the at least one audio signal; obtain at least one room effect parameter comprising at least one reverberation time; adjust at least one of the at least one room effect parameter when the at least one reverberation time is less than a determined threshold value; and generate a spatial audio signal using the spatial audio stream and the adjusted at least one room effect parameter, the generated spatial audio signal comprising:a first spatial audio portion generated based on the at least one audio signal and processed dependent on the spatial metadata; and a second spatial audio portion generated based on the at least one audio signal and processed dependent on the adjusted at least one room effect parameter.

Citation Information

Patent Citations

  • Analysis of spatial metadata from multi-microphones having asymmetric geometry in devices

    GB201619573D0

  • Spatial audio representation and rendering

    GB201914712D0

  • Analysis of spatial metadata from multi-microphones having asymmetric geometry in devices

    WO2018091776A1

  • An audio apparatus and method of operation therefor

    EP4174846A1

  • Spatial audio representation and rendering

    GB2588171A

Cited By

  • Latency handling for point-to-point communications

    US20240233742A9